Compare commits

...
Author SHA1 Message Date
osobhandClaude Opus 5.5 4cc3111151 docs: known-issues — cite PRs #30, #31, #32
CI / test-arm64 (pull_request) Successful in 1m39s
CI / test (pull_request) Successful in 21m9s
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 21:33:48 -05:00
osobhandClaude Opus 5.5 7fba2f3996 BENCHMARKS: streaming FileBuilder::write A/B (the 1 MiB buffer regression and the fix)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 21:33:37 -05:00
osobhandClaude Opus 5.5 549e442aff FileBuilder::write: 64 KiB write buffer, not 1 MiB
A 1 MiB BufWriter sits above glibc's 128 KiB mmap threshold; depending on
allocator history it was mapped afresh by every write, and its page faults
(1.6 M vs 13 K over the criterion write_2d_chunked group) made 1 MiB
chunked writes 1.35x-1.83x slower than main. With 64 KiB they match main.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 21:16:01 -05:00
osobhandClaude Opus 5.5 728ffeef16 docs: chunk dimensions of 2^32 or more and copy-free 4 GiB writes
known-issues: two fixed entries (chunk dimensions of 2^32 or more were
refused; a 4 GiB unfiltered chunk was held several times when written),
the open "Chunks of 4 GiB or more: limits" updated (LZ4/Zstd tested end
to end, memory figures of 2026-09-29) and the open-issues table in sync.
CHANGELOG (Unreleased) entry with the public API changes and the peak
RSS measurements; README capability table. The wasm package test lists
the new fixture and expects an Overflow error reading its chunks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:40 -05:00
osobhandClaude Opus 5.5 e4b3a76dc7 Chunk dimensions of 2^32 or more; unfiltered 4 GiB chunks written without copies
- DataLayout::Chunked::chunk_dimensions is Vec<u64> (was Vec<u32>), and
  the chunk index readers, writers and serializers take &[u64]: layout
  messages of version 4/5 store each dimension in up to 8 bytes, and
  libhdf5 2.x writes dimensions of 2^32 or more (layout version 5). Such
  dimensions were refused on read (InvalidChunkDimensions) and write. A
  version-3 layout (4-byte dimensions) is never written for them; a chunk
  whose size overflows 64 bits is refused when opened.
- The file writer lays chunked datasets out as pieces referring to the
  chunks instead of copying them into one buffer per pass, and an
  unfiltered chunk that is a contiguous run of the dataset's data (a
  dataset stored as one chunk of its shape, row blocks) borrows it; a
  filtered one is compressed straight from it. Contiguous datasets are not
  copied either. FileWriter::finish_with streams the file to a callback;
  FileBuilder::write uses it, so the file is never assembled in memory.
  DatasetBuilder::with_u8_data_owned takes the data without a copy.
  Peak RSS writing one unfiltered 1 GiB chunk (with_u8_data_owned +
  write): 5.0 GiB before, 1.0 GiB after; with_u8_data: 6.0 -> 2.0 GiB.
- extract_chunk no longer panics on data shorter than the shape.
- Tests: huge_chunk_dims.h5 fixture (libhdf5 2.0.0 via h5py 3.16, u8
  chunks of 2^32 + 7), and opt-in end-to-end tests of chunk dims >= 2^32
  (filtered and unfiltered, read and written, h5py and h5dump 2.2.0), of
  an unfiltered 4 GiB+ chunk written by clawhdf5, and of LZ4/Zstd chunks
  of that size; example write_one_chunk for memory measurements.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:33 -05:00
osobhandClaude Opus 5.5 e4ba09946f HDF5 2.0 native complex as a first-class type on read; Python libver=
CI / test-arm64 (pull_request) Successful in 1m38s
CI / test (pull_request) Successful in 21m11s
- Datatype::parse returns Datatype::Complex for class 11 (also inside
  compounds, arrays and VL types) instead of the {r, i} compound view.
- Facade: DType::Complex(Box<DType>); read_complex_f32/f64 accept it.
- h5rs dump/ls/diff print native complex as h5dump/h5ls/h5diff 2.2.0 do
  (checked against a fixture written by h5py 3.16 / libhdf5 2.0.0);
  dump --json keeps the {r, i} compound (hdf5-json has no complex class).
- clawhdf5-wasm reads native complex datasets as [re, im] pairs.
- Python: clawhdf5.File(path, 'w', libver=...) with h5py's values,
  mapped to FileBuilder::libver_bounds; 'v108' output opens in HDF5 1.8.23.
- Docs: known-issues entry moved to Fixed (history), CHANGELOG, READMEs.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:33 -05:00
40 changed files with 2339 additions and 359 deletions
+36
View File
@@ -601,6 +601,42 @@ about 150 per dataset here) against a Fixed Array's few blocks. Writing costs th
same; files grow by about 36 bytes per chunk. The default therefore stays
the 1.10 format; the 1.8 format is opt-in.
### Streaming `FileBuilder::write` (2026-09-29, tank)
`FileBuilder::write` now streams the file through a buffered writer
instead of building it in memory first (`feat/huge-chunk-dims`, the change
that removes copies of 4 GiB unfiltered chunks). Measured 2026-09-29 on
tank (AMD Ryzen 7 7800X3D), idle (1-minute load average 1.66 to 1.83 at the
start of each run): `main` at `4260af4` against the stacked branch at
`549e442`, alternating, three full runs each; median (min-max) of
criterion's estimate, ms.
> **Run:** `cargo bench -p clawhdf5-bench --bench h5bench_write -- --noplot --warm-up-time 1 --measurement-time 3`
The first candidate used a 1 MiB write buffer and was 1.35x to 1.83x slower
on every 512 x 512 (1 MiB) chunked case when the full suite ran (e.g.
`write_2d_chunked/512x512` 0.454 -> 0.754 ms): the buffer is above glibc's
128 KiB mmap threshold, and after the smaller cases it was mapped afresh on
each write (1 629 368 minor page faults over the `write_2d_chunked` group
against 12 888 on `main`). Run alone, the case showed no difference. With a
64 KiB buffer (`549e442`):
| benchmark | main | candidate | ratio |
|---|---:|---:|---:|
| write_1d_contiguous/100000 | 0.216 (0.215-0.218) | 0.206 (0.204-0.207) | 0.96x |
| write_2d_chunked/32x32 | 0.062 (0.061-0.063) | 0.061 (0.061-0.064) | 0.99x |
| write_2d_chunked/128x128 | 0.102 (0.100-0.103) | 0.101 (0.099-0.102) | 0.98x |
| write_2d_chunked/512x512 | 0.460 (0.457-0.463) | 0.466 (0.463-0.477) | 1.01x |
| write_2d_chunked_zstd/deflate-6/512x512 | 0.457 (0.450-0.465) | 0.465 (0.458-0.477) | 1.02x |
| write_2d_chunked_zstd/zstd-3/512x512 | 0.376 (0.369-0.381) | 0.387 (0.379-0.390) | 1.03x |
| write_2d_chunked_pcodec/pcodec/512x512 | 0.906 (0.895-0.908) | 0.901 (0.898-0.914) | 0.99x |
| write_multi_dataset/64 | 0.136 (0.135-0.138) | 0.136 (0.135-0.136) | 1.00x |
| write_with_attrs/64 | 0.063 (0.062-0.063) | 0.063 (0.062-0.063) | 1.00x |
The other 18 cases are within 2%. Minor page faults per full run: 59 602 to
64 486 on both sides. The write path is as fast as before; what changes is
peak memory for large unfiltered chunks (see CHANGELOG, 2026-09-29).
### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank)
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load
+122
View File
@@ -56,6 +56,128 @@
and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4).
Affected v2.1.0 to v2.7.0.
### HDF5 2.0 native complex is its own type on read; Python `libver=` (2026-09-29)
- **`Datatype::parse` returns `Datatype::Complex { size, base_type }`** for
a class-11 message, also as a compound member, array base or
variable-length base. It used to return the equivalent `{r, i}`
compound (`docs/known-issues.md`, "HDF5 2.0 native complex numbers read
as a `{r, i}` compound"). **Breaking** for code that matched the
compound view of a native complex type: `Dataset::raw_datatype()` and
attribute datatypes are now `Datatype::Complex`;
`Datatype::complex_as_compound(size, base)` still gives the `{r, i}` view,
and `data_read::read_compound_fields` accepts a `Complex` directly.
Re-serializing a parsed type (e.g. copying it to another file) now writes
class 11 again instead of a compound.
- **Facade:** `DType::Complex(Box<DType>)` (new variant, **breaking** for
exhaustive `match`es over `DType`): `Complex(F32)` is numpy `complex64`,
`Complex(F64)` `complex128`, a binary16 part `Complex(Other("float16"))`
(the same `Other` name small floats already use); `Display` prints
`complex<f64>`. h5py's own complex encoding, the compound `{r, i}`, is
still `DType::Compound`. `read_complex_f64`/`read_complex_f32` read both.
- **`h5rs`:** `dump` prints native complex as h5dump 2.2.0 does:
`H5T_COMPLEX_IEEE_F{16,32,64}{LE,BE}` (else `H5T_COMPLEX { <base> }`), in
arrays and compounds too, and values as `1.5-2i` (`%g%+gi`; for binary16
parts, which h5dump has no C type for, `1+-2i` as it prints them). `ls`
shows `complex64`, `complex128-be`, `complex32` (and `complex<part>` for
non-IEEE parts); `ls -v` shows h5ls 2.2.0's `complex number of` /
`IEEE 64-bit big-endian float`. `diff` compares complex values part by
part and prints the difference as h5diff 2.2.0 does (`1+0i`); a native
complex and an `{r, i}` compound are no longer comparable (different
classes, as in h5diff). `dump --json` keeps the `{r, i}` compound and
`[re, im]` values: hdf5-json (h5json 2.0.0) has no complex class.
- **Browser (`clawhdf5-wasm`):** native complex datasets are readable, as
`[re, im]` pairs: `info().elementShape` ends in `2`, `read()` returns the
parts interleaved in a `Float32Array`/`Float64Array` (binary16 parts
widened to f32), and `dtype` is `complex<f64>` (`complex<f32
(big-endian)>`, `array[2]<complex<f32>>`). The viewer shows each element
as `[re, im]` with no change to the page. h5py's `{r, i}` compound is
still refused, as every compound is.
- **Python:** `clawhdf5.File(path, 'w', libver=...)` takes h5py's values:
`'v108'`, `'v110'`, `'v112'`, `'v114'`, `'v200'`, `'latest'` (the low
bound; high `'latest'`, as h5py) or a `(low, high)` tuple, mapped to
`FileBuilder::libver_bounds`. `'earliest'` as the low bound writes the
1.8 format with a `UserWarning` (clawhdf5 cannot write the pre-1.8
format); as the high bound it is a `ValueError`, as are unknown names and
a low bound above the high one. It is ignored for `'r'` and refused
(`NotImplementedError`) for `'r+'`/`'a'`. Native complex already read as
numpy `complex64`/`complex128` and still does
(`test_read_native_complex_from_h5py`).
- Tests (tank, 2026-09-29): `crates/clawhdf5-tools/tests/native_complex_dump.rs`
compares `h5rs dump` with h5dump 2.2.0's output stored next to a fixture
written by h5py 3.16 / libhdf5 2.0.0
(`crates/clawhdf5/tests/fixtures/gen_native_complex.py`:
`F32LE`/`F64LE`/`F64BE`/`F16LE`, scalar, compound member, array, root
attribute), line for line outside the values and value for value within
them, and `ls`/`diff` with h5ls/h5diff 2.2.0;
`integration_tests::complex_datasets_and_attributes_round_trip`
(`DType::Complex`, `raw_datatype()`); `clawhdf5-wasm` unit tests on the
fixture, `h5py_interop` and `examples/wasm-viewer/test` (Node:
`test.mjs`, 274 + 1324 checks; the page in headless Chromium for
`/native_c128`, `/native_c64_be`, `/pairs`) with two native complex
datasets added to `make_fixture.py` (when h5py's libhdf5 is 2.0+);
`crates/clawhdf5-py/tests/test_libver.py` (every bound; `'v108'` output
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
hard-codes the fixture's length.
### Chunk dimensions of 2^32 or more; 4 GiB chunks written without copies (2026-09-29)
- **Chunk dimensions of 2^32 or more** (libhdf5 2.x writes them with
layout message version 5, up to 8 bytes per dimension; a 4 GiB chunk of
1-byte elements needs one) are read and written. They were refused
(`InvalidChunkDimensions`). **Public API changes:**
`DataLayout::Chunked::chunk_dimensions` is `Vec<u64>` (was `Vec<u32>`),
and the `chunk_dimensions` argument of `chunked_read::{collect_chunk_info_checked,
collect_chunk_info_checked_in, generate_implicit_chunks_in_grid}`,
`fixed_array::{read_fixed_array_chunks, read_fixed_array_chunks_in}`,
`extensible_array::{read_extensible_array_chunks,
read_extensible_array_chunks_in}`, and the `chunk_dims` argument of
`chunked_write::serialize_v4_single_chunk_pub`, are `&[u64]` (were
`&[u32]`). A chunk whose size overflows 64 bits is refused when a dataset is read
(`InvalidChunkDimensions`), and the writer refuses a chunk dimension of
0 up front. A version-3 layout (4-byte dimensions) is never written for
such a chunk: over 4 GiB, it always takes version 5, as in libhdf5.
- **Writing without copies:** `FileWriter` no longer copies chunks into a
buffer per pass and then into the file's buffer. A chunk that is a
contiguous run of the dataset's data (a dataset stored as one chunk of
its shape, blocks of whole rows) is borrowed when unfiltered and
compressed straight from it when filtered, and contiguous datasets are
not copied. New `FileWriter::finish_with(put)` hands the file out piece
by piece; `FileBuilder::write` uses it (through a 1 MiB `BufWriter`,
still to a temporary file renamed into place), so a file is never
assembled in memory. New `DatasetBuilder::with_u8_data_owned(Vec<u8>)`
takes the data without the copy `with_u8_data` makes. Peak resident
memory of one unfiltered 1 GiB chunk of `u8`
(`cargo run --release -p clawhdf5 --example write_one_chunk --
1073741824 OUT owned|slice`, tank, 2026-09-29): 5.0 GiB before (the
writer of `4260af4`), 1.0 GiB after with `owned`; 6.0 and 2.0 GiB with
`slice`; 4 GiB + 7 bytes `owned`, 4.00 GiB. Same bytes written. Write
speed was not re-measured (the machine was not idle): the A/B of
`FileBuilder::write` and `finish` on small and large files is still to
be run.
- `chunked_write::extract_chunk` (behind `split_into_chunks`) no longer
panics when the data is shorter than the shape; the missing part is
zeros, as it was meant to be.
- Tests: `crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5` (31 KB,
libhdf5 2.0.0 via h5py 3.16, `gen_huge_chunks.py dims`: `u8` chunks of
2^32 + 7 deflated twice, Single Chunk and Extensible Array), and in
`huge_chunks_interop`: `huge_chunk_dims_list` (always), and with
`CLAWHDF5_HUGE_CHUNKS=1` `huge_chunk_dims_read`,
`huge_chunk_dims_unfiltered_read` (a sparse file h5py writes, Single
Chunk and Fixed Array, mmap and positioned reads),
`writer_huge_chunk_dims_round_trip` (deflated, read back by clawhdf5,
h5py 3.16 and — layouts — h5dump 2.2.0),
`writer_unfiltered_huge_chunk_round_trip` (an unfiltered 4 GiB + 7 byte
chunk: clawhdf5, h5py, and h5dump 2.2.0 printing values past 2^32) and
`writer_huge_chunks_lz4_zstd_round_trip` (with the `lz4`/`zstd`
features; h5py through hdf5plugin 7.1.0). The whole opt-in suite
(`cargo test --release -p clawhdf5 --features lz4,zstd --test
huge_chunks_interop -- --test-threads=1`, h5py and h5dump 2.2.0
included) passed with a peak of 10.04 GiB resident (tank, 2026-09-29,
commit `0065c6b`). Unit tests: 5-byte dimension encoding, contiguous
chunk detection, `finish_with` against `finish`. The wasm package test
lists the new fixture and gets `FormatError::Overflow` reading it.
`docs/known-issues.md`: two fixed entries; the open "Chunks of 4 GiB or
more: limits" keeps the refused filters, the editor and memory.
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
+5 -4
View File
@@ -102,7 +102,8 @@ writes 85.2 µs vs 877 µs (10.3x); 64 group creates 130 µs vs 1.37 ms
(10.6x); a 512×512 `f32` chunked deflate-6 write 1.44 ms vs 65.0 ms
(re-measured 2026-09-23 with the pure-Rust deflate: 1.46 ms vs 51.4 ms,
35x); a 100K `f32` sequential write is a tie. The writer (`FileBuilder`)
assembles a file in memory and writes it once, which is part of that
lays out the whole file before writing it (when these were measured, in
one buffer in memory), which is part of that
difference; read the caveats in [BENCHMARKS.md](BENCHMARKS.md#caveats)
before quoting these.
@@ -115,12 +116,12 @@ Limits and open issues, with dates, are in
|---|---|---|---|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more and chunk dimensions of 2^32 or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error) |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays with h5py's `libver=`, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`; HDF5 2.0 native complex as `[re, im]` pairs); no Zstd/SZIP/pcodec, no compound (h5py's `{r, i}` complex included), reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`,
`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure
+42 -30
View File
@@ -555,7 +555,7 @@ pub(crate) fn checked_byte_len(elements: u64, elem_size: usize) -> Result<usize,
/// HDF5 2.0 writes them with layout version 5. A zero chunk dimension used
/// to read as all fill values, and a huge one to hang the reader.
pub(crate) fn chunk_geometry(
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
layout_version: u8,
dataspace: &Dataspace,
elem_size: usize,
@@ -576,15 +576,26 @@ pub(crate) fn chunk_geometry(
"chunk size must be > 0, dim = {d}"
)));
}
let bytes = spatial
.iter()
.fold(elem_size as u128, |acc, &c| acc * u128::from(c));
// Saturating: 33 dimensions of up to 64 bits each can overflow u128.
let bytes = spatial.iter().fold(elem_size as u128, |acc, &c| {
acc.saturating_mul(u128::from(c))
});
// So the indexes can compute a chunk's size in 64 bits.
if bytes > u128::from(u64::MAX) {
return Err(FormatError::InvalidChunkDimensions(format!(
"chunk size overflows 64 bits (chunk {spatial:?} of {elem_size}-byte elements)"
)));
}
if layout_version < 4 && bytes > u128::from(u32::MAX) {
return Err(FormatError::InvalidChunkDimensions(format!(
"chunk size must be < 4GB with v1 b-tree index (chunk {spatial:?} of {elem_size}-byte elements)"
)));
}
Ok((rank, spatial.iter().map(|&c| c as usize).collect()))
let spatial = spatial
.iter()
.map(|&c| to_usize(c))
.collect::<Result<Vec<usize>, _>>()?;
Ok((rank, spatial))
}
/// The size of one element of `dt` as stored in the file: a
@@ -625,7 +636,7 @@ pub fn check_chunk_element_size(
return Ok(());
};
let expected = stored_element_size(datatype, offset_size);
if u64::from(stored) != expected {
if stored != expected {
return Err(FormatError::InvalidChunkDimensions(format!(
"stored datatype size in chunk layout does not match datatype description \
(layout {stored} bytes, datatype {expected})"
@@ -759,7 +770,7 @@ pub fn collect_chunk_info_in<S: Storage + ?Sized>(
pub fn collect_chunk_info_checked(
file_data: &[u8],
btree_address: u64,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
@@ -776,7 +787,7 @@ pub fn collect_chunk_info_checked(
pub fn collect_chunk_info_checked_in<S: Storage + ?Sized>(
file_data: &S,
btree_address: u64,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
@@ -808,7 +819,7 @@ pub fn collect_chunk_info_checked_in<S: Storage + ?Sized>(
.iter_mut()
.zip(chunk.offsets.iter().zip(chunk_dimensions))
{
*w = o / u64::from(d);
*w = o / d;
}
wanted[ndims - 1] = 0;
if let Some(i) = root.find(&wanted) {
@@ -849,9 +860,9 @@ impl ChunkNode {
}
}
fn scale_keys(&mut self, dims: &[u32]) {
fn scale_keys(&mut self, dims: &[u64]) {
for (k, &d) in self.keys.iter_mut().zip(dims.iter().cycle()) {
*k /= u64::from(d);
*k /= d;
}
if let ChunkChildren::Nodes(nodes) = &mut self.children {
for n in nodes {
@@ -919,9 +930,9 @@ fn btree_cmp3(lt: &[u64], scaled: &[u64], rt: &[u64]) -> core::cmp::Ordering {
/// Check one v1 B-tree chunk key's offsets (see
/// [`collect_chunk_info_checked`]).
fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u32]) -> Result<(), FormatError> {
fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u64]) -> Result<(), FormatError> {
for (&offset, &dim) in offsets.iter().zip(chunk_dimensions) {
if dim == 0 || offset % u64::from(dim) != 0 {
if dim == 0 || offset % dim != 0 {
return Err(FormatError::ChunkedReadError(format!(
"bad coordinate offset {offsets:?} for chunk dimensions {chunk_dimensions:?}"
)));
@@ -937,7 +948,7 @@ fn read_key_offsets(
file_data: &[u8],
pos: usize,
ndims: usize,
chunk_dimensions: Option<&[u32]>,
chunk_dimensions: Option<&[u64]>,
out: &mut Vec<u64>,
) -> Result<(), FormatError> {
let start = out.len();
@@ -966,7 +977,7 @@ fn parse_chunk_node<S: Storage + ?Sized>(
file_data: &S,
btree_address: u64,
ndims: usize,
chunk_dimensions: Option<&[u32]>,
chunk_dimensions: Option<&[u64]>,
offset_size: u8,
depth: usize,
stored: &mut Vec<ChunkInfo>,
@@ -1075,7 +1086,7 @@ fn parse_chunk_node<S: Storage + ?Sized>(
pub fn generate_implicit_chunks(
base_address: u64,
dataset_dims: &[u64],
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
) -> Vec<ChunkInfo> {
generate_implicit_chunks_in_grid(
@@ -1097,17 +1108,18 @@ pub fn generate_implicit_chunks_in_grid(
base_address: u64,
dataset_dims: &[u64],
max_dims: &[u64],
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
) -> Vec<ChunkInfo> {
let rank = chunk_dimensions.len();
let chunk_byte_size: u64 =
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
let chunk_byte_size: u64 = chunk_dimensions
.iter()
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
let mut num_chunks_per_dim = Vec::with_capacity(rank);
let mut grid_per_dim = Vec::with_capacity(rank);
for d in 0..rank {
let ch = chunk_dimensions[d] as u64;
let ch = chunk_dimensions[d];
let n = dataset_dims[d].div_ceil(ch);
num_chunks_per_dim.push(n);
grid_per_dim.push(max_dims.get(d).map_or(n, |m| m.div_ceil(ch)).max(n));
@@ -1125,7 +1137,7 @@ pub fn generate_implicit_chunks_in_grid(
let nchunks = num_chunks_per_dim[d];
let chunk_idx = remaining % nchunks;
remaining /= nchunks;
offsets[d] = chunk_idx * chunk_dimensions[d] as u64;
offsets[d] = chunk_idx * chunk_dimensions[d];
grid_idx = grid_idx.saturating_add(chunk_idx.saturating_mul(down));
down = down.saturating_mul(grid_per_dim[d]);
}
@@ -1350,7 +1362,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
}
(4, Some(2)) => {
// Implicit index — use spatial chunk dims only
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
generate_implicit_chunks_in_grid(
addr,
&dataspace.dimensions,
@@ -1364,7 +1376,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
}
(4, Some(3)) => {
// Fixed Array — use spatial chunk dims only
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
let header = FixedArrayHeader::parse_in(
file_data,
checked_addr(addr)?,
@@ -1384,7 +1396,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
}
(4, Some(4)) => {
// Extensible Array — use spatial chunk dims only
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
let header = ExtensibleArrayHeader::parse_in(
file_data,
checked_addr(addr)?,
@@ -2435,12 +2447,12 @@ mod tests {
/// A leaf whose final key is the one libhdf5 writes past the last chunk:
/// each coordinate of the last chunk plus its chunk dimension (the
/// element size last).
fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u32], offset_size: u8) -> Vec<u8> {
fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u64], offset_size: u8) -> Vec<u8> {
let last = &chunks.last().expect("a chunk").offsets;
let end: Vec<u64> = dims
.iter()
.enumerate()
.map(|(d, &c)| last.get(d).copied().unwrap_or(0) + u64::from(c))
.map(|(d, &c)| last.get(d).copied().unwrap_or(0) + c)
.collect();
build_chunk_btree_leaf_to(chunks, &end, offset_size)
}
@@ -2771,7 +2783,7 @@ mod tests {
}
// Build B-tree at offset 0x100
let dims = [chunk_size_elems as u32, elem_size as u32];
let dims = [chunk_size_elems as u64, elem_size as u64];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
@@ -3145,7 +3157,7 @@ mod tests {
data_offset += compressed.len() + 16; // some padding
}
let dims = [chunk_elems as u32, elem_size as u32];
let dims = [chunk_elems as u64, elem_size as u64];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
@@ -3228,7 +3240,7 @@ mod tests {
}
}
let dims = [chunk_dims[0] as u32, chunk_dims[1] as u32, elem_size as u32];
let dims = [chunk_dims[0] as u64, chunk_dims[1] as u64, elem_size as u64];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
@@ -3408,7 +3420,7 @@ mod tests {
}
let layout = DataLayout::Chunked {
chunk_dimensions: vec![chunk_elems as u32, elem_size as u32],
chunk_dimensions: vec![chunk_elems as u64, elem_size as u64],
btree_address: Some(data_addr as u64),
version: 4,
chunk_index_type: Some(1),
+350 -101
View File
@@ -5,12 +5,14 @@ extern crate alloc;
use crate::addr::saturating_usize;
#[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec};
use alloc::{borrow::Cow, format, vec, vec::Vec};
#[cfg(feature = "std")]
use std::borrow::Cow;
use crate::btree_v1_write;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::checksum::jenkins_lookup3;
use crate::chunk_cache::{CACHE_LINE_SIZE, align_to_cache_line};
use crate::chunk_cache::CACHE_LINE_SIZE;
use crate::chunk_grid::ChunkGrid;
use crate::ea_writer;
use crate::error::FormatError;
@@ -464,7 +466,9 @@ fn extract_chunk(
let dst: usize = (0..rank).map(|d| idx[d] * chunk_strides[d]).sum::<usize>() * element_size;
// Whole elements only, as far as `raw_data` reaches.
let n = row.min(raw_data.len().saturating_sub(src)) / element_size * element_size;
if n > 0 {
chunk[dst..dst + n].copy_from_slice(&raw_data[src..src + n]);
}
// Next row: advance every dimension but the last.
let mut d = rank - 1;
loop {
@@ -481,6 +485,66 @@ fn extract_chunk(
}
}
/// The bytes of the `linear_idx`-th chunk (see [`extract_chunk`]),
/// borrowed from `raw_data` when the chunk is one contiguous run of it: a
/// chunk inside the dataset (not past its edge) whose dimensions after the
/// first one above 1 equal the dataset's. A dataset stored as one chunk of
/// its own shape is such a run, so its chunk is never copied. Other chunks
/// are copied out, as [`extract_chunk`] does.
fn chunk_bytes_of<'a>(
raw_data: &'a [u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
linear_idx: u64,
) -> Cow<'a, [u8]> {
if shape.is_empty() {
return Cow::Borrowed(raw_data);
}
if let Some(range) = contiguous_chunk(shape, chunk_dims, element_size, linear_idx)
&& let (Ok(start), Ok(end)) = (usize::try_from(range.start), usize::try_from(range.end))
&& end <= raw_data.len()
{
return Cow::Borrowed(&raw_data[start..end]);
}
Cow::Owned(extract_chunk(raw_data, shape, chunk_dims, element_size, linear_idx).1)
}
/// The byte range of the `linear_idx`-th chunk in the row-major dataset,
/// when the chunk is a contiguous run of it (see [`chunk_bytes_of`]).
fn contiguous_chunk(
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
linear_idx: u64,
) -> Option<core::ops::Range<u64>> {
let rank = shape.len();
// Past the first chunk dimension above 1, every chunk dimension must
// span the dataset's.
let first_wide = chunk_dims.iter().position(|&c| c > 1).unwrap_or(rank);
if (first_wide + 1..rank).any(|d| chunk_dims[d] != shape[d]) {
return None;
}
let mut remaining = linear_idx;
let mut start = 0u64;
let mut stride = element_size as u64;
for d in (0..rank).rev() {
let n = shape[d].div_ceil(chunk_dims[d]);
let offset = (remaining % n).checked_mul(chunk_dims[d])?;
remaining /= n;
// A chunk past the dataset's edge is padded: not a run of it.
if offset.checked_add(chunk_dims[d])? > shape[d] {
return None;
}
start = start.checked_add(offset.checked_mul(stride)?)?;
stride = stride.checked_mul(shape[d])?;
}
let len = chunk_dims
.iter()
.try_fold(element_size as u64, |acc, &c| acc.checked_mul(c))?;
Some(start..start.checked_add(len)?)
}
/// Parallel compression threshold: use rayon when chunk count exceeds this.
///
/// Lowered to 2 to enable parallel compression for typical 4-chunk workloads
@@ -508,25 +572,21 @@ const PARALLEL_COMPRESS_MAX_CHUNK_BYTES: u64 = 64 << 20;
/// at most [`PARALLEL_COMPRESS_MAX_CHUNK_BYTES`], compression runs across
/// rayon threads; otherwise it is sequential. Output order matches chunk
/// order, so per-chunk bytes are identical to the sequential path.
fn compress_all_chunks(
raw_data: &[u8],
fn compress_all_chunks<'a>(
raw_data: &'a [u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
chunk_bytes: u64,
pipeline: &Option<FilterPipeline>,
) -> Result<Vec<(u64, Vec<u8>, u32)>, FormatError> {
let one = |i: u64| -> Result<(u64, Vec<u8>, u32), FormatError> {
let (_, raw) = if shape.is_empty() {
(Vec::new(), raw_data.to_vec())
} else {
extract_chunk(raw_data, shape, chunk_dims, element_size, i)
};
) -> Result<Vec<PreparedChunk<'a>>, FormatError> {
let one = |i: u64| -> Result<PreparedChunk<'a>, FormatError> {
let raw = chunk_bytes_of(raw_data, shape, chunk_dims, element_size, i);
let raw_size = raw.len() as u64;
match pipeline {
Some(pl) => {
let (stored, mask) = compress_chunk_masked(&raw, pl, element_size as u32)?;
Ok((raw_size, stored, mask))
Ok((raw_size, Cow::Owned(stored), mask))
}
None => Ok((raw_size, raw, 0)),
}
@@ -557,7 +617,7 @@ fn compress_all_chunks(
/// layout/pipeline messages. `base_address` is where the blob will be placed in the file.
/// Serialize a v4 single chunk layout message (public for OH size estimation).
pub fn serialize_v4_single_chunk_pub(
chunk_dims: &[u32],
chunk_dims: &[u64],
chunk_address: u64,
filtered_size: Option<u64>,
filter_mask: Option<u32>,
@@ -578,7 +638,7 @@ pub fn serialize_v4_single_chunk_pub(
/// Serialize a v4 (or, for a chunk of 4 GiB or more, v5) single chunk
/// layout message.
fn serialize_v4_single_chunk(
chunk_dims: &[u32],
chunk_dims: &[u64],
chunk_address: u64,
filtered_size: Option<u64>,
filter_mask: Option<u32>,
@@ -622,7 +682,7 @@ fn serialize_v4_single_chunk(
/// Serialize a v4 Fixed Array layout message.
fn serialize_v4_fixed_array(
chunk_dims: &[u32],
chunk_dims: &[u64],
fixed_array_address: u64,
offset_size: u8,
element_size: u32,
@@ -653,7 +713,8 @@ fn serialize_v4_fixed_array(
/// dimensions, then the element size). Each takes the fewest bytes that hold
/// the largest, as libhdf5 computes it (`H5D__chunk_set_sizes`:
/// `(log2(dim) + 8) / 8`); HDF5 2.0.0 refuses any other width.
pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u32], element_size: u32) {
pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u64], element_size: u32) {
let element_size = u64::from(element_size);
let max_dim = chunk_dims
.iter()
.copied()
@@ -661,14 +722,14 @@ pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u32], element_
.max()
.unwrap_or(1)
.max(1);
let width = (32 - max_dim.leading_zeros()).div_ceil(8) as usize;
let width = (64 - max_dim.leading_zeros()).div_ceil(8) as usize;
buf.push(width as u8);
for &d in chunk_dims.iter().chain(core::iter::once(&element_size)) {
buf.extend_from_slice(&d.to_le_bytes()[..width]);
}
}
fn layout_v4_chunked_prefix(chunk_dims: &[u32], element_size: u32, version: u8) -> Vec<u8> {
fn layout_v4_chunked_prefix(chunk_dims: &[u64], element_size: u32, version: u8) -> Vec<u8> {
let mut buf = Vec::new();
buf.push(version);
buf.push(2); // class = chunked
@@ -884,7 +945,66 @@ pub fn precompress_chunks(
element_size: usize,
options: &ChunkOptions,
) -> Result<PrecompressedChunks, FormatError> {
let (_, chunk_bytes) = checked_chunk_dims(chunk_dims, element_size)?;
let p = prepare_chunks(raw_data, shape, chunk_dims, element_size, options)?;
Ok(PrecompressedChunks {
chunks: p
.chunks
.into_iter()
.map(|(raw, stored, mask)| (raw, stored.into_owned(), mask))
.collect(),
has_filters: p.has_filters,
element_size: p.element_size,
shape: p.shape,
chunk_dims: p.chunk_dims,
pipeline_message: p.pipeline_message,
})
}
/// One chunk of [`PreparedChunks`]: its raw size, its stored bytes and its
/// filter mask.
pub(crate) type PreparedChunk<'a> = (u64, Cow<'a, [u8]>, u32);
/// [`PrecompressedChunks`] whose stored bytes may borrow the dataset's data:
/// an unfiltered chunk that is a contiguous run of it (a dataset stored as
/// one chunk, say) is not copied. What the file writer lays out.
pub(crate) struct PreparedChunks<'a> {
pub(crate) chunks: Vec<PreparedChunk<'a>>,
pub(crate) has_filters: bool,
pub(crate) element_size: usize,
pub(crate) shape: Vec<u64>,
pub(crate) chunk_dims: Vec<u64>,
pub(crate) pipeline_message: Option<Vec<u8>>,
}
impl PrecompressedChunks {
/// A view of these chunks as [`PreparedChunks`] (their bytes borrowed).
fn prepared(&self) -> PreparedChunks<'_> {
PreparedChunks {
chunks: self
.chunks
.iter()
.map(|(raw, stored, mask)| (*raw, Cow::Borrowed(&stored[..]), *mask))
.collect(),
has_filters: self.has_filters,
element_size: self.element_size,
shape: self.shape.clone(),
chunk_dims: self.chunk_dims.clone(),
pipeline_message: self.pipeline_message.clone(),
}
}
}
/// [`precompress_chunks`], borrowing what it can (see [`PreparedChunks`]).
/// A filtered chunk that is a contiguous run of `raw_data` is compressed
/// straight from it, without a copy of the raw chunk.
pub(crate) fn prepare_chunks<'a>(
raw_data: &'a [u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
options: &ChunkOptions,
) -> Result<PreparedChunks<'a>, FormatError> {
let chunk_bytes = checked_chunk_dims(chunk_dims, element_size)?;
if chunk_bytes > MAX_V4_CHUNK_BYTES {
check_huge_chunk_filters(options, chunk_bytes)?;
}
@@ -902,7 +1022,7 @@ pub fn precompress_chunks(
&pipeline,
)?;
Ok(PrecompressedChunks {
Ok(PreparedChunks {
chunks,
has_filters,
element_size,
@@ -912,26 +1032,87 @@ pub fn precompress_chunks(
})
}
/// The chunk dimensions as the layout message stores them (each below
/// 2^32), and one chunk's size in bytes. A chunk dimension of 2^32 or more
/// (which HDF5 2.0 can store, in wider fields) is refused, as it is when
/// read: clawhdf5 holds chunk dimensions as `u32`. So is a chunk whose size
/// overflows 64 bits, or that this platform cannot hold in memory (a chunk
/// of 4 GiB or more on a 32-bit target).
fn checked_chunk_dims(
chunk_dims: &[u64],
element_size: usize,
) -> Result<(Vec<u32>, u64), FormatError> {
let dims = chunk_dims
.iter()
.map(|&d| {
u32::try_from(d).map_err(|_| {
FormatError::InvalidChunkDimensions(format!(
"chunk dimension {d} is 2^32 or more, which clawhdf5 does not support"
))
})
})
.collect::<Result<Vec<u32>, _>>()?;
/// A piece of a chunked dataset's data blob as laid out by
/// [`place_chunked_data`]: zero padding, a chunk's stored bytes (by index
/// into [`PreparedChunks::chunks`]), or index structures.
#[derive(Debug)]
pub(crate) enum Piece {
Zeros(usize),
Chunk(usize),
Bytes(Vec<u8>),
}
/// A chunked dataset laid out at an address, without its chunk bytes
/// copied: the pieces of its data blob, in order, and its messages.
pub(crate) struct ChunkedPlacement {
pub(crate) pieces: Vec<Piece>,
/// Length of the data blob in bytes.
pub(crate) len: u64,
pub(crate) layout_message: Vec<u8>,
pub(crate) pipeline_message: Option<Vec<u8>>,
}
impl ChunkedPlacement {
/// The data blob as one buffer (the chunks copied in).
fn into_result(self, pre: &PreparedChunks<'_>) -> ChunkedDataResult {
let mut data_bytes = Vec::with_capacity(saturating_usize(self.len));
for piece in &self.pieces {
match piece {
Piece::Zeros(n) => data_bytes.resize(data_bytes.len() + n, 0),
Piece::Chunk(i) => data_bytes.extend_from_slice(&pre.chunks[*i].1),
Piece::Bytes(b) => data_bytes.extend_from_slice(b),
}
}
ChunkedDataResult {
data_bytes,
layout_message: self.layout_message,
pipeline_message: self.pipeline_message,
}
}
}
/// Builds the pieces of a data blob, tracking its length.
#[derive(Default)]
struct Pieces {
pieces: Vec<Piece>,
len: u64,
}
impl Pieces {
/// Pad to the next cache line.
fn align(&mut self) {
let aligned = align_chunk_offset(self.len);
if aligned > self.len {
self.pieces
.push(Piece::Zeros(saturating_usize(aligned - self.len)));
self.len = aligned;
}
}
fn chunk(&mut self, i: usize, len: usize) {
self.pieces.push(Piece::Chunk(i));
self.len += len as u64;
}
fn bytes(&mut self, b: Vec<u8>) {
self.len += b.len() as u64;
self.pieces.push(Piece::Bytes(b));
}
}
/// One chunk's size in bytes. A chunk dimension of 0 is refused, and so is
/// a chunk whose size overflows 64 bits, or that this platform cannot hold
/// in memory (a chunk of 4 GiB or more on a 32-bit target). Chunk
/// dimensions of 2^32 or more are allowed, as in libhdf5 2.x: such a chunk
/// is over 4 GiB, so it takes layout message version 5, whose dimension
/// fields are up to 8 bytes wide (a version-3 layout, with 4-byte fields,
/// is never written for it).
fn checked_chunk_dims(chunk_dims: &[u64], element_size: usize) -> Result<u64, FormatError> {
if chunk_dims.contains(&0) {
return Err(FormatError::InvalidChunkDimensions(format!(
"chunk dimensions {chunk_dims:?} include 0"
)));
}
let bytes = chunk_dims
.iter()
.try_fold(element_size as u64, |acc, &d| acc.checked_mul(d))
@@ -941,7 +1122,7 @@ fn checked_chunk_dims(
"a chunk of {chunk_dims:?} x {element_size} bytes exceeds this platform's address space"
))
})?;
Ok((dims, bytes))
Ok(bytes)
}
/// The filters clawhdf5 can apply to a chunk of more than `u32::MAX` bytes:
@@ -1010,7 +1191,22 @@ pub fn build_chunked_data_from_precompressed_libver(
low: LibVer,
high: LibVer,
) -> Result<ChunkedDataResult, FormatError> {
let (_, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?;
let pre = pre.prepared();
let placement = place_chunked_data(&pre, base_address, maxshape, low, high)?;
Ok(placement.into_result(&pre))
}
/// [`build_chunked_data_from_precompressed_libver`] without copying the
/// chunks into one buffer: the data blob as [`Piece`]s referring to
/// `pre`'s chunks. The file writer writes the pieces out one by one.
pub(crate) fn place_chunked_data(
pre: &PreparedChunks<'_>,
base_address: u64,
maxshape: Option<&[u64]>,
low: LibVer,
high: LibVer,
) -> Result<ChunkedPlacement, FormatError> {
let chunk_bytes = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?;
if chunk_bytes > MAX_V4_CHUNK_BYTES {
if high < LibVer::V200 {
return Err(FormatError::LibverBound {
@@ -1020,7 +1216,7 @@ pub fn build_chunked_data_from_precompressed_libver(
});
}
} else if low < LibVer::V110 {
return build_btree_v1_chunked_data(pre, base_address, maxshape);
return place_btree_v1_chunked_data(pre, base_address, maxshape);
}
let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?;
let offset_size: u8 = 8;
@@ -1028,17 +1224,14 @@ pub fn build_chunked_data_from_precompressed_libver(
let num_chunks = pre.chunks.len();
let element_size = pre.element_size;
let mut data_buf = Vec::new();
let mut blob = Pieces::default();
let mut written_chunks = Vec::with_capacity(num_chunks);
for (raw_size, compressed, filter_mask) in &pre.chunks {
let aligned_offset = align_to_cache_line(data_buf.len());
if aligned_offset > data_buf.len() {
data_buf.resize(aligned_offset, 0u8);
}
let address = base_address + data_buf.len() as u64;
for (i, (raw_size, compressed, filter_mask)) in pre.chunks.iter().enumerate() {
blob.align();
let address = base_address + blob.len;
let compressed_size = compressed.len() as u64;
data_buf.extend_from_slice(compressed);
blob.chunk(i, compressed.len());
written_chunks.push(WrittenChunk {
address,
compressed_size,
@@ -1047,28 +1240,22 @@ pub fn build_chunked_data_from_precompressed_libver(
});
}
let (chunk_dims_u32, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, element_size)?;
let version = layout_version_for(chunk_bytes);
let aligned_idx = align_to_cache_line(data_buf.len());
if aligned_idx > data_buf.len() {
data_buf.resize(aligned_idx, 0u8);
}
blob.align();
let layout_message = match &index {
ChunkIndexPlan::ExtensibleArray(grid) => {
let ea_address = base_address + data_buf.len() as u64;
let ea_address = base_address + blob.len;
let slots = index_slots(grid, &pre.shape, &pre.chunk_dims, &written_chunks, None)?;
let ea_bytes = ea_writer::build_extensible_array_at(
blob.bytes(ea_writer::build_extensible_array_at(
&slots,
offset_size,
length_size,
pre.has_filters,
ea_address,
);
data_buf.extend_from_slice(&ea_bytes);
));
ea_writer::serialize_v4_extensible_array(
&chunk_dims_u32,
&pre.chunk_dims,
ea_address,
offset_size,
element_size as u32,
@@ -1084,7 +1271,7 @@ pub fn build_chunked_data_from_precompressed_libver(
};
let filter_mask = pre.has_filters.then_some(written_chunks[0].filter_mask);
serialize_v4_single_chunk(
&chunk_dims_u32,
&pre.chunk_dims,
chunk_addr,
filtered_size,
filter_mask,
@@ -1094,7 +1281,7 @@ pub fn build_chunked_data_from_precompressed_libver(
)
}
ChunkIndexPlan::FixedArray(grid, nslots) => {
let fa_address = base_address + data_buf.len() as u64;
let fa_address = base_address + blob.len;
let slots = index_slots(
grid,
&pre.shape,
@@ -1102,16 +1289,15 @@ pub fn build_chunked_data_from_precompressed_libver(
&written_chunks,
Some(*nslots),
)?;
let fa_bytes = build_fixed_array_at(
blob.bytes(build_fixed_array_at(
&slots,
offset_size,
length_size,
pre.has_filters,
fa_address,
);
data_buf.extend_from_slice(&fa_bytes);
));
serialize_v4_fixed_array(
&chunk_dims_u32,
&pre.chunk_dims,
fa_address,
offset_size,
element_size as u32,
@@ -1120,7 +1306,7 @@ pub fn build_chunked_data_from_precompressed_libver(
)
}
ChunkIndexPlan::BTreeV2 => {
let bt_address = base_address + data_buf.len() as u64;
let bt_address = base_address + blob.len;
let records: Vec<(Vec<u64>, &WrittenChunk)> = written_chunks
.iter()
.enumerate()
@@ -1134,9 +1320,9 @@ pub fn build_chunked_data_from_precompressed_libver(
pre.has_filters,
bt_address,
)?;
data_buf.extend_from_slice(&bt_bytes);
blob.bytes(bt_bytes);
serialize_v4_btree_v2(
&chunk_dims_u32,
&pre.chunk_dims,
bt_address,
offset_size,
element_size as u32,
@@ -1146,8 +1332,9 @@ pub fn build_chunked_data_from_precompressed_libver(
}
};
Ok(ChunkedDataResult {
data_bytes: data_buf,
Ok(ChunkedPlacement {
pieces: blob.pieces,
len: blob.len,
layout_message,
pipeline_message: pre.pipeline_message.clone(),
})
@@ -1156,11 +1343,11 @@ pub fn build_chunked_data_from_precompressed_libver(
/// Lay out precompressed chunks at `base_address` followed by a version-1
/// B-tree chunk index, with a version-3 layout message: what libhdf5 writes
/// for a chunked dataset under a low bound of 1.8.
fn build_btree_v1_chunked_data(
pre: &PrecompressedChunks,
fn place_btree_v1_chunked_data(
pre: &PreparedChunks<'_>,
base_address: u64,
maxshape: Option<&[u64]>,
) -> Result<ChunkedDataResult, FormatError> {
) -> Result<ChunkedPlacement, FormatError> {
if let Some(ms) = maxshape {
let bad = |what: &str| FormatError::ChunkedReadError(format!("maxshape: {what}"));
if ms.len() != pre.shape.len() {
@@ -1171,20 +1358,17 @@ fn build_btree_v1_chunked_data(
}
}
let offset_size: u8 = 8;
let mut data_buf = Vec::new();
let mut blob = Pieces::default();
let mut entries = Vec::with_capacity(pre.chunks.len());
for (i, (_raw_size, stored, filter_mask)) in pre.chunks.iter().enumerate() {
let aligned_offset = align_to_cache_line(data_buf.len());
if aligned_offset > data_buf.len() {
data_buf.resize(aligned_offset, 0u8);
}
blob.align();
entries.push(btree_v1_write::ChunkEntry {
scaled: scaled_coords(&pre.shape, &pre.chunk_dims, i),
nbytes: stored.len() as u64,
filter_mask: *filter_mask,
address: base_address + data_buf.len() as u64,
address: base_address + blob.len,
});
data_buf.extend_from_slice(stored);
blob.chunk(i, stored.len());
}
let element_size = u32::try_from(pre.element_size)
.map_err(|_| FormatError::Overflow("element size".into()))?;
@@ -1193,11 +1377,8 @@ fn build_btree_v1_chunked_data(
let btree_address = if entries.is_empty() {
u64::MAX
} else {
let aligned_idx = align_to_cache_line(data_buf.len());
if aligned_idx > data_buf.len() {
data_buf.resize(aligned_idx, 0u8);
}
let addr = base_address + data_buf.len() as u64;
blob.align();
let addr = base_address + blob.len;
let tree = btree_v1_write::build_chunk_btree_v1_at(
&entries,
&pre.chunk_dims,
@@ -1205,13 +1386,14 @@ fn build_btree_v1_chunked_data(
addr,
offset_size,
)?;
data_buf.extend_from_slice(&tree);
blob.bytes(tree);
addr
};
let layout_message =
serialize_v3_chunked(&pre.chunk_dims, btree_address, offset_size, element_size)?;
Ok(ChunkedDataResult {
data_bytes: data_buf,
Ok(ChunkedPlacement {
pieces: blob.pieces,
len: blob.len,
layout_message,
pipeline_message: pre.pipeline_message.clone(),
})
@@ -1424,7 +1606,7 @@ fn build_btree_v2_chunk_index_at(
/// Serialize a v4 layout message for a version-2 B-tree chunk index.
fn serialize_v4_btree_v2(
chunk_dims: &[u32],
chunk_dims: &[u64],
btree_address: u64,
offset_size: u8,
element_size: u32,
@@ -2184,7 +2366,7 @@ mod tests {
/// filtered index element stores the chunk's size in 8 bytes.
#[test]
fn huge_chunk_layout_messages_and_index_elements() {
let dims = [HUGE_DIM as u32];
let dims = [HUGE_DIM];
let parsed = |msg: &[u8]| {
assert_eq!(msg[0], 5, "layout message version");
// Class chunked, then (after the flags) 2 dimensions of 4 bytes.
@@ -2197,7 +2379,7 @@ mod tests {
chunk_index_type,
..
} => {
assert_eq!(chunk_dimensions, vec![HUGE_DIM as u32, 8]);
assert_eq!(chunk_dimensions, vec![HUGE_DIM, 8]);
chunk_index_type.unwrap()
}
other => panic!("{other:?}"),
@@ -2245,22 +2427,89 @@ mod tests {
assert_eq!(u16::from_le_bytes([bt[10], bt[11]]), 28);
}
/// A chunk dimension of 2^32 or more (a 1-byte element, 2^32 + 7 per
/// chunk): the dimensions take 5 bytes each, as libhdf5 2.x writes
/// them (`(log2(dim) + 8) / 8`), and parse back whole.
#[test]
fn chunk_dimensions_of_2_pow_32_or_more_take_5_bytes() {
const DIM: u64 = (1 << 32) + 7;
let msg = serialize_v4_single_chunk(&[DIM], 0x800, None, None, 8, 1, 5);
assert_eq!(&msg[..5], &[5, 2, 0, 2, 5]);
assert_eq!(&msg[5..10], &[7, 0, 0, 0, 1]);
assert_eq!(&msg[10..15], &[1, 0, 0, 0, 0]);
assert_eq!(msg[15], 1, "Single Chunk index");
match DataLayout::parse(&msg, 8, 8).unwrap() {
DataLayout::Chunked {
chunk_dimensions, ..
} => assert_eq!(chunk_dimensions, [DIM, 1]),
other => panic!("{other:?}"),
}
// The largest dimension takes all 8 bytes.
let mut buf = Vec::new();
push_v4_chunk_dims(&mut buf, &[u64::MAX, 3], 4);
assert_eq!(buf[0], 8);
assert_eq!(buf.len(), 1 + 3 * 8);
// A version-3 layout has 4-byte dimensions: refused, not truncated.
assert!(serialize_v3_chunked(&[DIM], 0x800, 8, 1).is_err());
}
/// Which chunks are contiguous runs of the dataset's data (borrowed,
/// not copied, when unfiltered).
#[test]
fn contiguous_chunks_are_borrowed() {
// Rows of a 2-D dataset: chunk (2, 5) of shape (5, 5).
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 0), Some(0..40));
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 1), Some(40..80));
// The last row chunk runs past the edge: padded, so copied.
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 2), None);
// Columns split: not contiguous.
assert_eq!(contiguous_chunk(&[4, 6], &[2, 3], 1, 0), None);
// Leading dimensions of 1, then a partial row: contiguous.
assert_eq!(contiguous_chunk(&[2, 3, 8], &[1, 1, 4], 2, 5), Some(40..48));
// The whole dataset as one chunk.
assert_eq!(contiguous_chunk(&[3, 4], &[3, 4], 8, 0), Some(0..96));
let raw: Vec<u8> = (0..100).collect();
let whole = chunk_bytes_of(&raw, &[10, 10], &[10, 10], 1, 0);
assert!(matches!(whole, Cow::Borrowed(b) if b.as_ptr() == raw.as_ptr()));
let rows = chunk_bytes_of(&raw, &[10, 10], &[3, 10], 1, 1);
assert!(matches!(rows, Cow::Borrowed(b) if b == &raw[30..60]));
// Data shorter than the shape is copied (and zero-padded) as before.
let short = chunk_bytes_of(&raw[..50], &[10, 10], &[10, 10], 1, 0);
assert!(matches!(&short, Cow::Owned(v) if v.len() == 100));
// Borrowed or copied, the bytes are what extract_chunk gives.
for (shape, chunk) in [
(&[10u64, 10][..], &[3u64, 10][..]),
(&[10, 10], &[3, 4]),
(&[4, 5, 5], &[1, 5, 5]),
(&[4, 5, 5], &[1, 2, 5]),
] {
for i in 0..chunk_count(shape, chunk) {
let got = chunk_bytes_of(&raw, shape, chunk, 1, i);
assert_eq!(&got[..], &extract_chunk(&raw, shape, chunk, 1, i).1[..]);
}
}
}
#[test]
fn huge_chunks_refused_where_unsupported() {
// A chunk dimension of 2^32 or more.
// A chunk dimension of 0.
assert!(matches!(
checked_chunk_dims(&[1 << 32], 1),
Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32")
checked_chunk_dims(&[4, 0], 1),
Err(FormatError::InvalidChunkDimensions(_))
));
// A chunk dimension of 2^32 or more is fine (on a 64-bit target).
#[cfg(target_pointer_width = "64")]
assert_eq!(
checked_chunk_dims(&[(1 << 32) + 5], 1).unwrap(),
(1 << 32) + 5
);
// A chunk size that overflows u64.
assert!(matches!(
checked_chunk_dims(&[u32::MAX.into(), u32::MAX.into(), 2], 8),
Err(FormatError::Overflow(_))
));
assert_eq!(
checked_chunk_dims(&[HUGE_DIM], 8).unwrap(),
(vec![HUGE_DIM as u32], HUGE_BYTES)
);
assert_eq!(checked_chunk_dims(&[HUGE_DIM], 8).unwrap(), HUGE_BYTES);
// Filters that cannot take a chunk that large.
for plugin in [
PluginFilter::Lzf,
+19 -19
View File
@@ -34,7 +34,7 @@ const MAX_LAYOUT_NDIMS: usize = 33;
/// (`H5O__layout_decode`): at most [`MAX_LAYOUT_NDIMS`], no dimension 0, and
/// before version 4 at least one dataspace dimension plus the element size.
/// A zero chunk dimension used to read the dataset as all fill values.
fn check_chunk_dims(dims: Vec<u32>, layout_version: u8) -> Result<Vec<u32>, FormatError> {
fn check_chunk_dims(dims: Vec<u64>, layout_version: u8) -> Result<Vec<u64>, FormatError> {
if dims.len() > MAX_LAYOUT_NDIMS {
return Err(FormatError::InvalidChunkDimensions(
"dimensionality is too large".into(),
@@ -72,7 +72,7 @@ pub enum DataLayout {
/// Chunked: data stored in chunks via a B-tree.
Chunked {
/// Chunk dimension sizes.
chunk_dimensions: Vec<u32>,
chunk_dimensions: Vec<u64>,
/// B-tree address, or `None` if undefined.
btree_address: Option<u64>,
/// Layout version (3 or 4). Version 1/2 messages (HDF5 1.4/1.6-era)
@@ -404,11 +404,11 @@ impl DataLayout {
_ => return Err(FormatError::InvalidLayoutClass(layout_class)),
};
ensure_len(data, p, dimensionality * 4)?;
let dims: Vec<u32> = data[p..p + dimensionality * 4]
let dims: Vec<u64> = data[p..p + dimensionality * 4]
.as_chunks::<4>()
.0
.iter()
.map(|c| u32::from_le_bytes(*c))
.map(|c| u64::from(u32::from_le_bytes(*c)))
.collect();
p += dimensionality * 4;
match layout_class {
@@ -424,7 +424,7 @@ impl DataLayout {
1 => {
let size = dims
.iter()
.try_fold(1u64, |acc, &d| acc.checked_mul(d as u64))
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
.ok_or_else(|| {
FormatError::Overflow(format!("contiguous layout size {dims:?}"))
})?;
@@ -489,7 +489,12 @@ impl DataLayout {
ensure_len(data, p, dimensionality * 4)?;
let mut chunk_dimensions = Vec::with_capacity(dimensionality);
for _ in 0..dimensionality {
let dim = u32::from_le_bytes([data[p], data[p + 1], data[p + 2], data[p + 3]]);
let dim = u64::from(u32::from_le_bytes([
data[p],
data[p + 1],
data[p + 2],
data[p + 3],
]));
chunk_dimensions.push(dim);
p += 4;
}
@@ -565,14 +570,8 @@ impl DataLayout {
.iter()
.rev()
.fold(0u64, |acc, &b| (acc << 8) | u64::from(b));
// Chunk dimensions are held as u32; HDF5 2.0 can write
// larger ones (layout version 5), which are refused
// rather than truncated.
let val = u32::try_from(val).map_err(|_| {
FormatError::InvalidChunkDimensions(format!(
"chunk dimension {val} is larger than 2^32 - 1, which is not supported"
))
})?;
// HDF5 2.0 writes dimensions of 2^32 or more (layout
// version 5, chunks over 4 GiB) in 5 to 8 bytes.
chunk_dimensions.push(val);
p += dim_size_encoded_length;
}
@@ -858,7 +857,7 @@ mod tests {
.unwrap_or_else(|e| panic!("width {width}: {e:?}"));
assert!(
matches!(&layout, DataLayout::Chunked { chunk_dimensions, .. }
if chunk_dimensions.iter().map(|&d| u64::from(d)).eq(dims)),
if chunk_dimensions[..] == dims),
"width {width}: {layout:?}"
);
}
@@ -871,13 +870,14 @@ mod tests {
)
);
}
// A dimension past u32 cannot be represented and is refused, not
// truncated.
// HDF5 2.0 writes dimensions of 2^32 or more (up to 8 bytes each).
for (width, dim) in [(5u8, (1u64 << 32) + 3), (8, u64::MAX)] {
assert!(matches!(
DataLayout::parse(&v4_chunked_msg(5, &[1 << 32, 8]), 8, 8),
Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32")
DataLayout::parse(&v4_chunked_msg(width, &[dim, 1]), 8, 8),
Ok(DataLayout::Chunked { chunk_dimensions, .. }) if chunk_dimensions == [dim, 1]
));
}
}
#[test]
fn v1v2_rejects_bad_class_dimensionality_and_truncation() {
+35 -29
View File
@@ -145,13 +145,16 @@ pub enum Datatype {
/// imaginary, in rectangular form. `size` is twice the base size and the
/// base is an IEEE float.
///
/// This variant exists for **writing** (see
/// [`Datatype::parse`] returns this variant for every class-11 message,
/// also inside compounds, arrays and variable-length types. Readers that
/// want the member view (`{r, i}`, as h5py writes complex numbers by
/// default) use [`Datatype::complex_as_compound`]; `read_compound_fields`
/// accepts this variant directly.
///
/// Writing it is opt-in (see
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
/// newer can read class 11, so it is opt-in and h5py's compound `{r, i}`
/// stays the default complex encoding. [`Datatype::parse`] still
/// surfaces a class-11 message as that equivalent `{r, i}` compound, so
/// every compound reader handles both encodings; parsing what this
/// variant serializes therefore yields a `Compound`, not a `Complex`.
/// newer can read class 11, so h5py's compound `{r, i}` stays the
/// default complex encoding.
Complex { size: u32, base_type: Box<Datatype> },
}
@@ -857,9 +860,7 @@ impl Datatype {
// Complex number (HDF5 2.0, datatype version 5). The properties
// are a single base floating-point datatype message; an element
// is two consecutive base-type values (real, imaginary). There
// is no member list. Surface it as the equivalent two-member
// compound `{r, i}` — the same shape h5py writes for numpy
// complex dtypes — so downstream compound readers work as-is.
// is no member list.
if version != 5 {
return Err(FormatError::InvalidDatatypeVersion {
class: class_id,
@@ -875,7 +876,13 @@ impl Datatype {
actual: size as usize,
});
}
Ok((Self::complex_as_compound(size, &base_type), pos))
Ok((
Datatype::Complex {
size,
base_type: Box::new(base_type),
},
pos,
))
}
_ => Err(FormatError::InvalidDatatypeClass(class_id)),
}
@@ -1174,9 +1181,8 @@ impl Datatype {
/// The `{r, i}` compound equivalent to a native complex type of `size`
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the
/// shape h5py writes for numpy complex dtypes, and what [`Self::parse`]
/// returns for a class-11 message. Readers that meet a
/// [`Datatype::Complex`] handle it through this view.
/// shape h5py writes for numpy complex dtypes. Readers that want a
/// [`Datatype::Complex`] as members handle it through this view.
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
let base_size = base_type.type_size();
Datatype::Compound {
@@ -1849,21 +1855,24 @@ mod tests {
fn test_complex_v5_from_hdf5_2_0() {
let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap();
assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len());
match dt {
Datatype::Compound { size, members } => {
assert_eq!(size, 16);
assert_eq!(members.len(), 2);
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
for m in &members {
match &dt {
Datatype::Complex { size, base_type } => {
assert_eq!(*size, 16);
assert!(matches!(
m.datatype,
base_type.as_ref(),
Datatype::FloatingPoint { size: 8, .. }
));
}
other => panic!("expected Complex, got {other:?}"),
}
other => panic!("expected Compound, got {other:?}"),
}
// The member view is h5py's `{r, i}` compound.
let Datatype::Compound { members, .. } =
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
else {
unreachable!()
};
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
}
#[test]
@@ -1887,7 +1896,7 @@ mod tests {
assert_eq!(members.len(), 2);
assert!(matches!(
&members[0].datatype,
Datatype::Compound { size: 16, members } if members.len() == 2
Datatype::Complex { size: 16, .. }
));
assert_eq!(
(members[1].name.as_str(), members[1].byte_offset),
@@ -1918,12 +1927,9 @@ mod tests {
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
assert_eq!(dt.type_size(), 16);
dt.check_encodable().unwrap();
// Parsing surfaces class 11 as the equivalent `{r, i}` compound.
// Parsing returns the native complex type unchanged.
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
assert_eq!(
parsed,
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
);
assert_eq!(parsed, dt);
let f32c = make_native_complex_f32_type().serialize();
assert_eq!(
+1 -1
View File
@@ -14,7 +14,7 @@ use crate::chunked_write::{
/// Serialize a v4 Extensible Array layout message.
pub(crate) fn serialize_v4_extensible_array(
chunk_dims: &[u32],
chunk_dims: &[u64],
ea_address: u64,
offset_size: u8,
element_size: u32,
@@ -454,7 +454,7 @@ pub fn read_extensible_array_chunks(
header: &ExtensibleArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
length_size: u8,
@@ -480,7 +480,7 @@ pub fn read_extensible_array_chunks_in<S: Storage + ?Sized>(
header: &ExtensibleArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
_length_size: u8,
@@ -489,12 +489,12 @@ pub fn read_extensible_array_chunks_in<S: Storage + ?Sized>(
// Linear indexes follow the maximum dimensions, with the unlimited
// dimension swizzled to the slowest position (see `chunk_grid`).
let dims_u64: Vec<u64> = chunk_dimensions.iter().map(|&d| d as u64).collect();
let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, &dims_u64)?;
let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, chunk_dimensions)?;
let grid = &grid;
let chunk_byte_size: u64 =
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
let chunk_byte_size: u64 = chunk_dimensions
.iter()
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
// Parse index block (EAIB): signature(4) + version(1) + client_id(1)
// + header address(offset_size), then the inline elements, then the
@@ -930,7 +930,7 @@ mod tests {
let header = ExtensibleArrayHeader::parse(&file_data, aehd_offset, os, ls).unwrap();
let ds_dims = vec![40u64]; // 2 chunks × 20 elements
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_extensible_array_chunks(
&file_data,
&header,
@@ -1057,7 +1057,7 @@ mod tests {
let file_data = build_inline_plus_data_blocks();
let header = ExtensibleArrayHeader::parse(&file_data, 0x100, os, ls).unwrap();
let ds_dims = vec![40u64];
let chunk_dims = vec![10u32];
let chunk_dims = vec![10u64];
let chunks = read_extensible_array_chunks(
&file_data,
&header,
+231 -56
View File
@@ -3,15 +3,14 @@
//! Produces valid HDF5 files with v3 superblock, v2 object headers,
//! link messages, contiguous datasets, inline and dense attributes.
use crate::addr::saturating_usize;
use crate::addr::{saturating_usize, to_usize};
#[cfg(not(feature = "std"))]
use alloc::{format, string::String, vec, vec::Vec};
use crate::attribute::AttributeMessage;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::chunked_write::{
ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed_libver,
precompress_chunks,
ChunkOptions, Piece, PreparedChunks, place_chunked_data, prepare_chunks,
};
use crate::data_layout::VdsMapping;
use crate::dataspace::{Dataspace, DataspaceType};
@@ -1612,7 +1611,30 @@ impl FileWriter {
self
}
/// The file, in memory. Its datasets' data is copied into the buffer
/// once (a dataset stored as one chunk of its own shape, or unfiltered
/// chunks that are contiguous runs of its data, are not copied before);
/// to write a file without holding it all, use [`Self::finish_with`].
pub fn finish(self) -> Result<Vec<u8>, FormatError> {
let mut buf = Vec::new();
self.finish_into(&mut buf)?;
Ok(buf)
}
/// The file, handed to `put` in order, a piece at a time, without
/// assembling it in memory: an unfiltered chunk or contiguous dataset
/// is passed straight from the data the dataset was given. Writing a
/// dataset stored as one unfiltered chunk this way needs little memory
/// beyond the dataset's own data. `put`'s first error stops the write
/// and is returned.
pub fn finish_with<E: From<FormatError>>(
self,
put: impl FnMut(&[u8]) -> Result<(), E>,
) -> Result<(), E> {
self.finish_into(&mut FnSink(put))
}
fn finish_into<S: Sink>(self, sink: &mut S) -> Result<(), S::Error> {
let page_size = self.page_size;
if let Some(ps) = page_size
&& !(MIN_FILE_SPACE_PAGE_SIZE..=MAX_FILE_SPACE_PAGE_SIZE).contains(&ps)
@@ -1620,7 +1642,8 @@ impl FileWriter {
return Err(FormatError::SerializationError(format!(
"file space page size {ps} is outside libhdf5's \
{MIN_FILE_SPACE_PAGE_SIZE}..={MAX_FILE_SPACE_PAGE_SIZE} bytes"
)));
))
.into());
}
let (low, high) = (self.low, self.high);
@@ -1774,15 +1797,16 @@ impl FileWriter {
})
.collect::<Result<_, _>>()?;
struct DataBlob {
struct DataBlob<'a> {
data: Vec<u8>,
oh_bytes: Vec<u8>,
/// Cached compressed chunks for chunked datasets; reused in Pass 2
/// to avoid re-compressing the same data.
precompressed: Option<PrecompressedChunks>,
/// to avoid re-compressing the same data. Unfiltered chunks
/// borrow the dataset's data where they can.
precompressed: Option<PreparedChunks<'a>>,
}
let mut dummy_blobs: Vec<DataBlob> = Vec::new();
let mut dummy_blobs: Vec<DataBlob<'_>> = Vec::new();
let mut dummy_cursor = 0u64;
for (i, d) in all_ds.iter().enumerate() {
let dense_blob = ds_dense[i]
@@ -1820,21 +1844,16 @@ impl FileWriter {
.resolve_chunk_dims_for(&d.ds.dimensions, elem_size);
// Compress once in Pass 1; cache the result so Pass 2 can skip
// re-compression and just rebuild the index with real addresses.
let pre = precompress_chunks(
let pre = prepare_chunks(
&d.raw,
&d.ds.dimensions,
&chunk_dims,
elem_size,
&d.chunk_options,
)?;
let result = build_chunked_data_from_precompressed_libver(
&pre,
dummy_cursor,
d.maxshape.as_deref(),
low,
high,
)?;
dummy_cursor += result.data_bytes.len() as u64;
let result =
place_chunked_data(&pre, dummy_cursor, d.maxshape.as_deref(), low, high)?;
dummy_cursor += result.len;
let oh = build_chunked_dataset_oh(
&d.dt,
&d.ds,
@@ -1849,7 +1868,7 @@ impl FileWriter {
d.refcount,
)?;
dummy_blobs.push(DataBlob {
data: result.data_bytes,
data: Vec::new(),
oh_bytes: oh,
precompressed: Some(pre),
});
@@ -1955,7 +1974,7 @@ impl FileWriter {
})
.collect::<Result<_, FormatError>>()?;
let mut ds_blobs2: Vec<DataBlob> = Vec::new();
let mut ds_blobs2: Vec<OutBlob<'_>> = Vec::new();
let global_align_threshold = self.alignment_threshold;
let global_align_bytes = self.alignment_bytes;
for (i, d) in all_ds.iter().enumerate() {
@@ -1977,26 +1996,21 @@ impl FileWriter {
&d.fill_message,
d.refcount,
)?;
ds_blobs2.push(DataBlob {
data: gcol_bytes.clone(),
ds_blobs2.push(OutBlob {
data: vec![Out::Slice(gcol_bytes)],
oh_bytes: oh,
precompressed: None,
});
} else if is_chunked[i] {
let base_address = cursor2 as u64;
// Reuse precompressed chunks from Pass 1 — avoids re-compressing
// the same data a second time.
let result = build_chunked_data_from_precompressed_libver(
dummy_blobs[i]
let pre = dummy_blobs[i]
.precompressed
.as_ref()
.expect("chunked dataset missing precompressed cache"),
base_address,
d.maxshape.as_deref(),
low,
high,
)?;
cursor2 += result.data_bytes.len();
.expect("chunked dataset missing precompressed cache");
let result =
place_chunked_data(pre, base_address, d.maxshape.as_deref(), low, high)?;
cursor2 += to_usize(result.len)?;
let oh = build_chunked_dataset_oh(
&d.dt,
&d.ds,
@@ -2010,11 +2024,16 @@ impl FileWriter {
&d.fill_message,
d.refcount,
)?;
ds_blobs2.push(DataBlob {
data: result.data_bytes,
oh_bytes: oh,
precompressed: None,
});
let data = result
.pieces
.into_iter()
.map(|p| match p {
Piece::Zeros(n) => Out::Zeros(n),
Piece::Chunk(c) => Out::Slice(&pre.chunks[c].1),
Piece::Bytes(b) => Out::Bytes(b),
})
.collect();
ds_blobs2.push(OutBlob { data, oh_bytes: oh });
} else if is_compact[i] {
// Compact: data is inline in the object header, no external blob
let oh = build_compact_dataset_oh(
@@ -2030,10 +2049,9 @@ impl FileWriter {
d.refcount,
layout_version,
)?;
ds_blobs2.push(DataBlob {
ds_blobs2.push(OutBlob {
data: vec![],
oh_bytes: oh,
precompressed: None,
});
} else {
// Determine alignment: per-dataset overrides global
@@ -2060,13 +2078,10 @@ impl FileWriter {
d.refcount,
layout_version,
)?;
let mut data = vec![0u8; padding];
data.extend_from_slice(&d.raw);
cursor2 += d.raw.len();
ds_blobs2.push(DataBlob {
data,
ds_blobs2.push(OutBlob {
data: vec![Out::Zeros(padding), Out::Slice(&d.raw)],
oh_bytes: oh,
precompressed: None,
});
}
}
@@ -2080,7 +2095,8 @@ impl FileWriter {
cursor2 = cursor2.next_multiple_of(ps as usize);
}
let eof_addr2 = cursor2 as u64;
let mut buf = Vec::with_capacity(cursor2);
sink.reserve(cursor2);
let mut out = Counted { sink, written: 0 };
let sb = Superblock {
version: superblock_version,
@@ -2103,9 +2119,9 @@ impl FileWriter {
checksum: None,
page_size: None,
};
buf.extend_from_slice(&sb.serialize());
out.put(&sb.serialize())?;
if let Some(ref ext) = sb_ext {
buf.extend_from_slice(ext);
out.put(ext)?;
}
// Group OHs + dense blobs (link blob, then attr blob, matching pass 2)
@@ -2132,32 +2148,120 @@ impl FileWriter {
g.refcount,
)?;
debug_assert_eq!(oh.len(), group_oh_sizes[gi]);
debug_assert_eq!(buf.len() as u64, group_addrs2[gi]);
buf.extend_from_slice(&oh);
debug_assert_eq!(out.written as u64, group_addrs2[gi]);
out.put(&oh)?;
if let Some(ref b) = link_blob {
buf.extend_from_slice(&b.blob);
out.put(&b.blob)?;
}
if let Some(ref blob) = group_dense_blobs[gi] {
buf.extend_from_slice(&blob.blob);
out.put(&blob.blob)?;
}
}
// Dataset OHs + dense blobs
for (i, blob) in ds_blobs2.iter().enumerate() {
buf.extend_from_slice(&blob.oh_bytes);
out.put(&blob.oh_bytes)?;
if let Some(ref dense) = ds_dense_blobs[i] {
buf.extend_from_slice(&dense.blob);
out.put(&dense.blob)?;
}
}
// Data
for blob in &ds_blobs2 {
buf.extend_from_slice(&blob.data);
for piece in &blob.data {
match piece {
Out::Bytes(b) => out.put(b)?,
Out::Slice(s) => out.put(s)?,
Out::Zeros(n) => out.zeros(*n)?,
}
}
}
debug_assert_eq!(buf.len(), data_end);
buf.resize(cursor2, 0);
Ok(buf)
debug_assert_eq!(out.written, data_end);
out.zeros(cursor2 - data_end)?;
Ok(())
}
}
/// A piece of a dataset's data as [`FileWriter`] writes it out.
enum Out<'a> {
Bytes(Vec<u8>),
/// Borrowed from the dataset's data or its prepared chunks.
Slice(&'a [u8]),
Zeros(usize),
}
/// A dataset's object header and the pieces of its data.
struct OutBlob<'a> {
oh_bytes: Vec<u8>,
data: Vec<Out<'a>>,
}
/// Where [`FileWriter`] puts a file's bytes, in order.
trait Sink {
type Error: From<FormatError>;
/// The file will be `_len` bytes long.
fn reserve(&mut self, _len: usize) {}
fn put(&mut self, bytes: &[u8]) -> Result<(), Self::Error>;
fn zeros(&mut self, n: usize) -> Result<(), Self::Error> {
const ZEROS: [u8; 4096] = [0; 4096];
let mut left = n;
while left > 0 {
let k = left.min(ZEROS.len());
self.put(&ZEROS[..k])?;
left -= k;
}
Ok(())
}
}
impl Sink for Vec<u8> {
type Error = FormatError;
fn reserve(&mut self, len: usize) {
self.reserve_exact(len.saturating_sub(self.len()));
}
fn put(&mut self, bytes: &[u8]) -> Result<(), FormatError> {
self.extend_from_slice(bytes);
Ok(())
}
fn zeros(&mut self, n: usize) -> Result<(), FormatError> {
self.resize(self.len() + n, 0);
Ok(())
}
}
/// A [`Sink`] calling a function ([`FileWriter::finish_with`]).
struct FnSink<F>(F);
impl<E: From<FormatError>, F: FnMut(&[u8]) -> Result<(), E>> Sink for FnSink<F> {
type Error = E;
fn put(&mut self, bytes: &[u8]) -> Result<(), E> {
(self.0)(bytes)
}
}
/// A [`Sink`] and the number of bytes put into it.
struct Counted<'s, S> {
sink: &'s mut S,
written: usize,
}
impl<S: Sink> Counted<'_, S> {
fn put(&mut self, bytes: &[u8]) -> Result<(), S::Error> {
self.written += bytes.len();
self.sink.put(bytes)
}
fn zeros(&mut self, n: usize) -> Result<(), S::Error> {
self.written += n;
self.sink.zeros(n)
}
}
@@ -3015,4 +3119,75 @@ mod tests {
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 3);
assert_eq!(layout_of(&bytes, "x")[0], 3);
}
/// `finish_with` hands out exactly the bytes `finish` returns, for every
/// storage (contiguous, compact, chunked filtered and not, chunks
/// borrowed and copied, a version-1 B-tree, a paged file), and stops at
/// the sink's first error.
#[test]
fn finish_with_matches_finish() {
let build = |low: LibVer, paged: bool| {
let mut fw = FileWriter::new();
fw.libver_bounds(low, LibVer::Latest);
if paged {
fw.with_page_size(4096);
}
let v: Vec<f64> = (0..300).map(f64::from).collect();
fw.create_dataset("contig").with_f64_data(&v);
fw.create_dataset("compact")
.with_f64_data(&[1.0, 2.0])
.compact();
fw.create_dataset("one_chunk")
.with_f64_data(&v)
.with_shape(&[30, 10])
.with_chunks(&[30, 10]);
fw.create_dataset("rows")
.with_f64_data(&v)
.with_shape(&[30, 10])
.with_chunks(&[7, 10]);
fw.create_dataset("tiles")
.with_f64_data(&v)
.with_shape(&[30, 10])
.with_chunks(&[7, 3]);
fw.create_dataset("deflated")
.with_f64_data(&v)
.with_chunks(&[64])
.with_shuffle()
.with_deflate(4);
fw.create_dataset("grows")
.with_u8_data_owned((0..=255).collect())
.with_chunks(&[100])
.with_maxshape(&[u64::MAX]);
fw
};
for (low, paged) in [
(LibVer::V18, false),
(LibVer::V110, false),
(LibVer::Latest, true),
] {
let bytes = build(low, paged).finish().unwrap();
let mut streamed = Vec::new();
build(low, paged)
.finish_with(|b: &[u8]| {
streamed.extend_from_slice(b);
Ok::<(), FormatError>(())
})
.unwrap();
assert_eq!(streamed, bytes, "{low:?} paged {paged}");
}
let mut calls = 0;
let err = build(LibVer::V110, false)
.finish_with(|_| {
calls += 1;
if calls == 3 {
Err(FormatError::SerializationError("full".into()))
} else {
Ok(())
}
})
.unwrap_err();
assert_eq!(err, FormatError::SerializationError("full".into()));
assert_eq!(calls, 3);
}
}
+9 -9
View File
@@ -155,7 +155,7 @@ pub fn read_fixed_array_chunks(
header: &FixedArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
length_size: u8,
@@ -180,7 +180,7 @@ pub fn read_fixed_array_chunks_in<S: Storage + ?Sized>(
header: &FixedArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
_length_size: u8,
@@ -227,11 +227,11 @@ pub fn read_fixed_array_chunks_in<S: Storage + ?Sized>(
// The index is laid out over the chunk grid of the *maximum* dimensions
// (row-major), so a dataset smaller than its maxshape has gaps.
let dims_u64: Vec<u64> = chunk_dimensions.iter().map(|&d| d as u64).collect();
let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, &dims_u64)?;
let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, chunk_dimensions)?;
let chunk_byte_size: u64 =
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
let chunk_byte_size: u64 = chunk_dimensions
.iter()
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
let mut chunks = Vec::new();
// `rel` is relative to the data block, whose bytes are in `w`.
@@ -670,7 +670,7 @@ mod tests {
let header =
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
let ds_dims = vec![100u64];
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_fixed_array_chunks(
&file_data,
&header,
@@ -747,7 +747,7 @@ mod tests {
let header =
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
let ds_dims = vec![60u64];
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_fixed_array_chunks(
&file_data,
&header,
@@ -848,7 +848,7 @@ mod tests {
assert_eq!(header.num_elements, 11);
let ds_dims = vec![11u64 * 20];
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_fixed_array_chunks(
&file_data,
&header,
+7 -1
View File
@@ -663,11 +663,17 @@ impl DatasetBuilder {
}
pub fn with_u8_data(&mut self, data: &[u8]) -> &mut Self {
self.with_u8_data_owned(data.to_vec())
}
/// [`Self::with_u8_data`] without the copy: the builder takes the
/// vector. For large data this halves the memory a write needs.
pub fn with_u8_data_owned(&mut self, data: Vec<u8>) -> &mut Self {
self.datatype = Some(make_u8_type());
self.data = Some(data.to_vec());
if self.shape.is_none() {
self.shape = Some(vec![data.len() as u64]);
}
self.data = Some(data);
self
}
+16 -1
View File
@@ -104,7 +104,22 @@ writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and
`complex128` arrays; the file is written on `close()`. Complex arrays are
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
Rust API writes that on request).
Rust API writes that on request). Both forms read back as numpy
`complex64`/`complex128`.
`libver=` sets the library version bounds as in h5py: `'v108'`,
`'v110'`, `'v112'`, `'v114'`, `'v200'` or `'latest'` (the low bound; the
high bound is then `'latest'`), or a `(low, high)` tuple. With `'v108'`
the file is written in the HDF5 1.8 format (version-2 superblock,
version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the
HDF5 1.10 format. `'earliest'` as the low bound writes the 1.8 format too,
with a `UserWarning`: clawhdf5 cannot write the pre-1.8 format. The
argument is ignored when reading and refused for `'r+'`/`'a'`.
```python
with clawhdf5.File("old.h5", "w", libver="v108") as f: # HDF5 1.8 reads it
f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip")
```
## Editing a file in place
+90 -2
View File
@@ -13,10 +13,13 @@ use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
use clawhdf5_rs::LibVer;
/// Internal state for write mode.
struct WriteState {
path: PathBuf,
/// `libver=` as (low, high); `None` keeps the writer's default.
libver: Option<(LibVer, LibVer)>,
root_datasets: Vec<DatasetSpec>,
root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
groups: Vec<Arc<Mutex<WriteGroupState>>>,
@@ -80,10 +83,32 @@ impl PyFile {
/// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write
/// libver: library version bounds for mode 'w', as h5py's: one of
/// 'earliest', 'v108', 'v110', 'v112', 'v114', 'v200', 'latest' (the
/// low bound; the high bound is then 'latest') or a (low, high)
/// tuple of them. The low bound is the oldest HDF5 release whose
/// format the file uses ('v108': HDF5 1.8 can read it); the high
/// bound the newest whose features it may use. clawhdf5 cannot write
/// the pre-1.8 format, so a low bound of 'earliest' writes the 1.8
/// format (with a warning) and a high bound of 'earliest' is an
/// error. Default (None): the HDF5 1.10 format clawhdf5 has always
/// written. Ignored for reading; not supported with 'r+' / 'a'.
#[new]
#[pyo3(signature = (path, mode="r"))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
#[pyo3(signature = (path, mode="r", libver=None))]
fn new(
py: Python<'_>,
path: &str,
mode: &str,
libver: Option<&Bound<'_, PyAny>>,
) -> PyResult<Self> {
let filename = path.to_string();
let libver = libver.map(|v| parse_libver(py, v)).transpose()?;
if libver.is_some() && matches!(mode, "r+" | "a") {
return Err(PyNotImplementedError::new_err(format!(
"libver with mode '{mode}': clawhdf5's in-place editor keeps the format \
versions the file already uses"
)));
}
if is_url(path) {
if mode != "r" {
return Err(PyValueError::new_err(format!(
@@ -113,6 +138,7 @@ impl PyFile {
// Absolute now: the file is written at close, possibly
// after the working directory changed.
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
libver,
root_datasets: Vec::new(),
root_attrs: Arc::new(Mutex::new(Vec::new())),
groups: Vec::new(),
@@ -460,10 +486,71 @@ fn parse_compression(
}
}
/// One of h5py's `libver` names as a bound; `high` says which end it is.
/// `Ok(None)` is 'earliest' as the low bound: the pre-1.8 format, which
/// clawhdf5 cannot write.
fn libver_name(name: &str, high: bool) -> PyResult<Option<LibVer>> {
Ok(Some(match name {
"earliest" if high => {
return Err(PyValueError::new_err(
"libver high bound 'earliest' (the pre-1.8 format) cannot be written by \
clawhdf5; the oldest format it writes is 'v108'",
));
}
"earliest" => return Ok(None),
"v108" => LibVer::V18,
"v110" => LibVer::V110,
"v112" => LibVer::V112,
"v114" => LibVer::V114,
"v200" => LibVer::V200,
"latest" => LibVer::Latest,
other => {
return Err(PyValueError::new_err(format!(
"unknown libver '{other}'; expected 'earliest', 'v108', 'v110', 'v112', \
'v114', 'v200' or 'latest'"
)));
}
}))
}
/// h5py's `libver=`: a name (the low bound, high bound 'latest') or a
/// `(low, high)` pair.
fn parse_libver(py: Python<'_>, v: &Bound<'_, PyAny>) -> PyResult<(LibVer, LibVer)> {
let (low, high): (String, String) = match v.extract::<String>() {
Ok(name) => (name, "latest".into()),
Err(_) => v.extract().map_err(|_| {
PyValueError::new_err("libver must be a string or a (low, high) tuple of strings")
})?,
};
let Some(high) = libver_name(&high, true)? else {
unreachable!("a high bound is never None")
};
let low = match libver_name(&low, false)? {
Some(low) => low,
None => {
PyModule::import(py, "warnings")?.getattr("warn")?.call1((
"libver 'earliest': clawhdf5 cannot write the pre-1.8 format; the file \
is written in the HDF5 1.8 format ('v108') instead",
py.get_type::<pyo3::exceptions::PyUserWarning>(),
))?;
LibVer::V18
}
};
if low > high {
return Err(PyValueError::new_err(format!(
"libver low bound {low} is newer than the high bound {high}"
)));
}
Ok((low, high))
}
/// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| {
let mut builder = clawhdf5_rs::FileBuilder::new();
if let Some((low, high)) = state.libver {
builder.libver_bounds(low, high);
}
// Root attributes
let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner());
@@ -520,6 +607,7 @@ mod tests {
let state = WriteState {
path: path.clone(),
libver: None,
root_datasets: vec![DatasetSpec {
name: "data".into(),
data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]),
+1 -6
View File
@@ -189,12 +189,7 @@ pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Optio
{
clawhdf5_format::data_layout::DataLayout::Chunked {
chunk_dimensions, ..
} if chunk_dimensions.len() >= rank => Some(
chunk_dimensions[..rank]
.iter()
.map(|&d| u64::from(d))
.collect(),
),
} if chunk_dimensions.len() >= rank => Some(chunk_dimensions[..rank].to_vec()),
_ => None,
}
}
+143
View File
@@ -0,0 +1,143 @@
"""`clawhdf5.File(path, 'w', libver=...)`: h5py's library version bounds.
A file written with libver='v108' must open in HDF5 1.8. Its h5dump is
found through CLAWHDF5_H5DUMP18 or at ~/.cache/hdf5-1.8.23/bin/h5dump
(scripts/build-hdf5-1.8.sh builds it); without it that check is skipped.
"""
import os
import subprocess
import numpy as np
import pytest
import clawhdf5
def superblock_version(path):
with open(path, "rb") as f:
head = f.read(9)
assert head[:8] == b"\x89HDF\r\n\x1a\n"
return head[8]
def h5dump18():
path = os.environ.get("CLAWHDF5_H5DUMP18") or os.path.expanduser(
"~/.cache/hdf5-1.8.23/bin/h5dump"
)
try:
out = subprocess.run([path, "--version"], capture_output=True, text=True)
except OSError:
return None
return path if "1.8." in out.stdout else None
def write_sample(path, libver):
with clawhdf5.File(path, "w", libver=libver) as f:
f.create_dataset("x", data=np.arange(10, dtype=np.float64))
f.create_dataset(
"chunked",
data=np.arange(100, dtype=np.int32),
chunks=(30,),
compression="gzip",
)
f.create_dataset("z", data=np.array([1 + 2j, -3j], dtype=np.complex128))
g = f.create_group("g")
g.create_dataset("y", data=np.ones(3, dtype=np.float32))
f.attrs["version"] = 1
@pytest.mark.parametrize(
"libver, sb",
[
(None, 3),
("v108", 2),
(("v108", "v108"), 2),
(("v108", "latest"), 2),
("v110", 3),
("v112", 3),
("v114", 3),
("v200", 3),
("latest", 3),
(("v110", "v200"), 3),
],
)
def test_libver_bounds(h5py, tmp_path, libver, sb):
path = str(tmp_path / "libver.h5")
write_sample(path, libver)
# v108 low bound: HDF5 1.8's version-2 superblock; 1.10 and later: 3.
assert superblock_version(path) == sb
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
np.testing.assert_array_equal(f["z"][:], [1 + 2j, -3j])
np.testing.assert_array_equal(f["g/y"][:], np.ones(3))
assert f.attrs["version"] == 1
with clawhdf5.File(path, "r") as f:
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
def test_libver_v108_opens_in_hdf5_1_8(tmp_path):
h5dump = h5dump18()
if h5dump is None:
pytest.skip("no HDF5 1.8 h5dump (CLAWHDF5_H5DUMP18)")
path = str(tmp_path / "v108.h5")
write_sample(path, "v108")
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode == 0, out.stderr
assert "h5dump error" not in out.stderr
assert 'DATASET "y"' in out.stdout
# The chunked, deflated dataset (a version-1 B-tree) reads in full.
out = subprocess.run(
[h5dump, "-w", "0", "-d", "/chunked", path], capture_output=True, text=True
)
assert out.returncode == 0, out.stderr
assert "(0): " + ", ".join(str(k) for k in range(100)) + "\n" in out.stdout
# The default (1.10 format) does not open in 1.8: the bound matters.
path = str(tmp_path / "default.h5")
write_sample(path, None)
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode != 0
def test_libver_earliest_writes_v108_with_a_warning(h5py, tmp_path):
path = str(tmp_path / "earliest.h5")
with pytest.warns(UserWarning, match="pre-1.8"):
write_sample(path, "earliest")
assert superblock_version(path) == 2
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
@pytest.mark.parametrize(
"libver, match",
[
(("v108", "earliest"), "high bound 'earliest'"),
("v109", "unknown libver 'v109'"),
(("latest", "v108"), "newer than the high bound"),
(("v108",), "tuple"),
(3, "tuple"),
],
)
def test_libver_errors(tmp_path, libver, match):
with pytest.raises(ValueError, match=match):
clawhdf5.File(str(tmp_path / "bad.h5"), "w", libver=libver)
def test_libver_v108_high_bound_writes_everything(tmp_path):
# Native complex numbers need HDF5 2.0; the Python writer stores complex
# as h5py's {r, i} compound, which 1.8 reads, so the 1.8 high bound is
# fine for everything clawhdf5.File writes.
path = str(tmp_path / "v18only.h5")
write_sample(path, ("v108", "v108"))
assert superblock_version(path) == 2
def test_libver_not_for_editing(tmp_path):
path = str(tmp_path / "e.h5")
write_sample(path, None)
with pytest.raises(NotImplementedError, match="libver"):
clawhdf5.File(path, "r+", libver="latest")
# Reading ignores it, as the bounds only affect what is written.
with clawhdf5.File(path, "r", libver="v108") as f:
assert f["x"].shape == (10,)
+4 -4
View File
@@ -873,9 +873,9 @@ impl Checker<'_> {
self.counts.chunk_index_checksummed += 1;
}
let filtered = matches!(&info.filters, Ok(Some(p)) if !p.filters.is_empty());
let chunk_bytes = cdims.iter().try_fold(u64::from(dt.type_size()), |a, &d| {
a.checked_mul(u64::from(d))
});
let chunk_bytes = cdims
.iter()
.try_fold(u64::from(dt.type_size()), |a, &d| a.checked_mul(d));
let max = ds.max_dimensions.clone();
let mut seen: HashSet<Vec<u64>> = HashSet::with_capacity(chunks.len().min(1 << 20));
let mut reported = 0usize;
@@ -893,7 +893,7 @@ impl Checker<'_> {
));
} else {
for (i, (&o, &cd)) in c.offsets.iter().zip(cdims).enumerate() {
if o % u64::from(cd) != 0 {
if o % cd != 0 {
bad.push(format!(
"offset {o} in dimension {i} is not a multiple of the chunk size {cd}"
));
+22 -1
View File
@@ -603,7 +603,22 @@ impl Diff {
let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) {
(Some(x), Some(y), ..) => x.abs_diff(y).to_string(),
(_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64),
// Each part's absolute difference, as h5diff 2.x prints
// it (`1+0i`).
_ => match (&va, &vb) {
(
Value::Complex { re: ar, im: ai, .. },
Value::Complex { re: br, im: bi, .. },
) => match (number(ar), number(ai), number(br), number(bi)) {
(Some(ar), Some(ai), Some(br), Some(bi)) => format!(
"{}+{}i",
value::fmt_float((ar - br).abs(), 64),
value::fmt_float((ai - bi).abs(), 64)
),
_ => String::new(),
},
_ => String::new(),
},
};
rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}"));
}
@@ -713,6 +728,9 @@ impl Diff {
(Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => {
p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v))
}
(Value::Complex { re: pr, im: pi, .. }, Value::Complex { re: qr, im: qi, .. }) => {
self.equal(a, b, pr, qr) && self.equal(a, b, pi, qi)
}
(Value::Ref(None), Value::Ref(None)) => true,
(Value::Ref(Some(p)), Value::Ref(Some(q))) => {
// Addresses mean nothing across files: compare the paths the
@@ -846,7 +864,10 @@ fn types_comparable(a: &Datatype, b: &Datatype) -> bool {
(
Datatype::Enumeration { base_type: x, .. },
Datatype::Enumeration { base_type: y, .. },
) => types_comparable(x, y),
)
| (Datatype::Complex { base_type: x, .. }, Datatype::Complex { base_type: y, .. }) => {
types_comparable(x, y)
}
_ => true,
}
}
+34 -6
View File
@@ -135,9 +135,16 @@ pub fn short(dt: &Datatype) -> String {
return f.short.into();
}
match dt {
Datatype::Complex { size, base_type } => {
short(&Datatype::complex_as_compound(*size, base_type))
}
// numpy's names (`complex64` is two `float32`) for IEEE parts,
// `complex<part>` otherwise.
Datatype::Complex { size, base_type } => match base_type.as_ref() {
Datatype::FloatingPoint { byte_order, .. } if is_ieee(base_type) => format!(
"complex{}{}",
u64::from(*size) * 8,
if be(byte_order) { "-be" } else { "" }
),
_ => format!("complex<{}>", short(base_type)),
},
Datatype::FixedPoint {
size,
signed,
@@ -216,8 +223,9 @@ pub fn long(dt: &Datatype) -> String {
return f.long.into();
}
match dt {
Datatype::Complex { size, base_type } => {
long(&Datatype::complex_as_compound(*size, base_type))
// As h5ls 2.x prints a complex type it has no native name for.
Datatype::Complex { base_type, .. } => {
format!("complex number of\n {}", long(base_type))
}
Datatype::FixedPoint {
size,
@@ -361,6 +369,19 @@ fn atomic_ddl(dt: &Datatype) -> Option<String> {
)),
// As h5dump 2.x names them (checked against h5dump 2.2.0).
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
// h5dump 2.x's predefined complex names: IEEE binary16/32/64 parts.
Datatype::Complex { base_type, .. } => match base_type.as_ref() {
Datatype::FloatingPoint {
size, byte_order, ..
} if is_ieee(base_type) && !matches!(byte_order, DatatypeByteOrder::Vax) => {
Some(format!(
"H5T_COMPLEX_IEEE_F{}{}",
u64::from(*size) * 8,
order_suffix(byte_order)
))
}
_ => None,
},
Datatype::BitField {
size, byte_order, ..
} => Some(format!(
@@ -418,6 +439,10 @@ pub fn ddl(dt: &Datatype, ind: usize) -> String {
Datatype::VariableLength { base_type, .. } => {
format!("H5T_VLEN {{ {} }}", ddl(base_type, ind))
}
// Parts h5dump has no predefined complex name for.
Datatype::Complex { base_type, .. } => {
format!("H5T_COMPLEX {{ {} }}", ddl(base_type, ind))
}
Datatype::Opaque { size, tag } => {
let tag = String::from_utf8_lossy(tag);
format!(
@@ -498,6 +523,9 @@ fn string_ddl(
/// hdf5-json type object.
pub fn json(dt: &Datatype) -> J {
match dt {
// hdf5-json (h5json 2.0.0) has no complex class: a native complex
// is written as the `{r, i}` compound h5py uses for complex numbers,
// its values as `[re, im]` pairs.
Datatype::Complex { size, base_type } => {
json(&Datatype::complex_as_compound(*size, base_type))
}
@@ -598,7 +626,7 @@ fn pad_json(p: &StringPadding) -> &'static str {
/// Class name used to decide whether two datatypes can be compared.
pub fn class(dt: &Datatype) -> &'static str {
match dt {
Datatype::Complex { .. } => "compound",
Datatype::Complex { .. } => "complex",
Datatype::FixedPoint { .. } => "integer",
Datatype::FloatingPoint { .. } => "float",
Datatype::Time { .. } => "time",
+1 -1
View File
@@ -292,7 +292,7 @@ impl Ls<'_> {
.unwrap_or(0);
let bytes = dims
.iter()
.try_fold(esize, |a, &d| a.checked_mul(u64::from(d)))
.try_fold(esize, |a, &d| a.checked_mul(d))
.map(|b| b.to_string())
.unwrap_or_else(|| "?".into());
writeln!(
+31 -5
View File
@@ -27,6 +27,15 @@ pub enum Value {
/// An enum member (name, when the value matches one) and its value.
Enum(Option<String>, i128),
Compound(Vec<(String, Value)>),
/// A native complex number (HDF5 2.0 class 11): real and imaginary
/// part. `sign_free` is set when h5dump joins the parts with a bare `+`
/// (parts it has no native C type for: binary16, non-IEEE), giving
/// `1+-2i`; otherwise the imaginary part carries its sign (`1-2i`).
Complex {
re: Box<Value>,
im: Box<Value>,
sign_free: bool,
},
Array(Vec<Value>),
/// A variable-length sequence.
Seq(Vec<Value>),
@@ -183,11 +192,18 @@ impl<'a> Decoder<'a> {
return Value::Error("short element".into());
};
match dt {
Datatype::Complex { size, base_type } => self.decode(
&Datatype::complex_as_compound(*size, base_type),
b,
depth + 1,
),
Datatype::Complex { base_type, .. } => {
let part = base_type.type_size() as usize;
let (re, im) = b.split_at(part.min(b.len()));
Value::Complex {
re: Box::new(self.decode(base_type, re, depth + 1)),
im: Box::new(self.decode(base_type, im, depth + 1)),
// h5dump prints `float`/`double` complex (whatever the
// byte order) with `%g%+gi`, anything else as
// `<re>+<im>i`.
sign_free: !(dtype::is_ieee(base_type) && matches!(part, 4 | 8)),
}
}
Datatype::FixedPoint { .. } => match decode_int(dt, b) {
Some(v) => Value::Int(v),
None => Value::Bytes(b.to_vec()),
@@ -354,6 +370,15 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
.collect::<Vec<_>>()
.join(", ")
),
Value::Complex { re, im, sign_free } => {
let (re, im) = (text(re, h5paths), text(im, h5paths));
let plus = if *sign_free || !im.starts_with('-') {
"+"
} else {
""
};
format!("{re}{plus}{im}i")
}
Value::Array(vs) => format!(
"[ {} ]",
vs.iter()
@@ -408,6 +433,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
Value::Complex { re, im, .. } => J::Array(vec![to_json(re, h5paths), to_json(im, h5paths)]),
Value::Array(vs) | Value::Seq(vs) => {
J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect())
}
@@ -0,0 +1,202 @@
//! `h5rs dump` and `h5rs ls` of HDF5 2.0 native complex numbers (datatype
//! class 11), against the output of h5dump 2.2.0 and h5ls 2.2.0 (the Debian
//! h5dump CI installs, 1.14.x, predates the type). The h5dump output is
//! stored next to the fixture; see
//! `crates/clawhdf5/tests/fixtures/gen_native_complex.py`.
use std::path::PathBuf;
use std::process::Command;
fn fixture(name: &str) -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name)
}
fn h5rs(args: &[&str]) -> String {
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(args)
.output()
.unwrap();
assert!(out.status.success(), "h5rs {args:?}: {out:?}");
String::from_utf8(out.stdout).unwrap()
}
/// Split a dump into its lines, with the data lines (`(i): ...`) replaced
/// by their index, and the data values in order.
fn split(ddl: &str) -> (Vec<String>, Vec<String>) {
let (mut frame, mut data) = (Vec::new(), Vec::new());
for line in ddl.lines() {
match line.split_once("):") {
Some((idx, body)) if idx.trim_start().starts_with('(') => {
frame.push(format!("{idx}):"));
// Values are split at `, `; array elements come bracketed.
data.extend(
body.split(", ")
.map(|s| s.trim().trim_matches(|c| c == '[' || c == ']' || c == ','))
.map(str::trim)
.filter(|s| !s.is_empty() && *s != "{")
.map(String::from),
);
}
_ => {
let t = line.trim().trim_end_matches(',');
if t.ends_with('i') || t.parse::<f64>().is_ok() {
// A compound member's value on its own line.
data.push(t.to_string());
} else {
frame.push(line.to_string());
}
}
}
}
(frame, data)
}
/// `a+bi`, `a-bi` or `a+-bi` as its two parts (a real number as one).
fn parts(v: &str) -> Vec<f64> {
let Some(z) = v.strip_suffix('i') else {
return vec![v.parse().unwrap()];
};
let b = z.as_bytes();
let at = (1..b.len())
.find(|&k| (b[k] == b'+' || b[k] == b'-') && !matches!(b[k - 1], b'e' | b'E' | b'+'))
.unwrap_or_else(|| panic!("not a complex value: {v}"));
let (re, im) = z.split_at(at);
let im = im.strip_prefix('+').unwrap_or(im);
vec![re.parse().unwrap(), im.parse().unwrap()]
}
#[test]
fn dump_matches_h5dump_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
let ours = h5rs(&["dump", path.to_str().unwrap()]);
let reference = std::fs::read_to_string(fixture("native_complex_hdf5_2.ddl")).unwrap();
let (mut our_frame, our_data) = split(&ours);
let (mut ref_frame, ref_data) = split(&reference);
// The first line names the file as it was given.
our_frame.remove(0);
ref_frame.remove(0);
// Everything but the values is h5dump's byte for byte: the types print
// as H5T_COMPLEX_IEEE_F64LE, H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }, ...
assert_eq!(our_frame, ref_frame);
assert_eq!(our_data.len(), ref_data.len(), "{our_data:?}\n{ref_data:?}");
assert_eq!(our_data.len(), 36);
// Values: h5dump prints `%g` (6 digits), inf/nan, and the binary16
// parts it has no C type for as `<re>+<im>i` (`1+-2i`); h5rs prints
// the shortest round-trip string and Inf/NaN, joined the same way.
for (i, (o, r)) in our_data.iter().zip(&ref_data).enumerate() {
assert_eq!(o.contains("+-"), r.contains("+-"), "[{i}] {o} vs {r}");
let (op, rp) = (parts(o), parts(r));
assert_eq!(op.len(), rp.len(), "[{i}] {o} vs {r}");
for (x, y) in op.iter().zip(&rp) {
let same = if y.is_nan() || y.is_infinite() {
x.is_nan() == y.is_nan() && (y.is_nan() || x == y)
} else {
*x as f32 == *y as f32 || (x - y).abs() <= 5e-6 * y.abs()
};
assert!(same, "[{i}] h5rs {o}, h5dump {r}");
assert_eq!(
x.is_sign_negative(),
y.is_sign_negative(),
"[{i}] {o} vs {r}"
);
}
}
}
#[test]
fn ls_describes_complex_like_h5ls_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
// h5ls 2.2.0 -v, `Type:` of the types it has no native C name for (it
// prints `native double _Complex` for the others, as it prints `native
// double` where h5rs prints `IEEE 64-bit little-endian float`).
for (name, long) in [
(
"c128be",
"complex number of\n IEEE 64-bit big-endian float",
),
(
"c128",
"complex number of\n IEEE 64-bit little-endian float",
),
(
"c32",
"complex number of\n IEEE 16-bit little-endian float",
),
(
"array",
"[2] complex number of\n IEEE 32-bit little-endian float",
),
] {
let target = format!("{}/{name}", path.display());
let verbose = h5rs(&["ls", "-v", &target]);
assert!(
verbose.contains(&format!(" Type: {long}\n")),
"{name}:\n{verbose}"
);
}
let listing = h5rs(&["ls", path.to_str().unwrap()]);
for (name, short) in [
("c64", "complex64"),
("c128", "complex128"),
("c128be", "complex128-be"),
("c32", "complex32"),
("scalar", "complex128"),
("array", "array[2]<complex64>"),
("compound", "compound{z: complex128, k: int64}"),
] {
assert!(
listing
.lines()
.any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))),
"{name}:\n{listing}"
);
}
}
#[test]
fn json_dump_uses_the_r_i_compound() {
// hdf5-json (h5json 2.0.0) has no complex class: the type is written as
// the `{r, i}` compound h5py uses, the values as `[re, im]` pairs.
let path = fixture("native_complex_hdf5_2.h5");
let j: serde_json::Value =
serde_json::from_str(&h5rs(&["dump", "--json", path.to_str().unwrap()])).unwrap();
let ds = j["datasets"]
.as_object()
.unwrap()
.values()
.find(|d| d["alias"][0] == "/scalar")
.unwrap();
assert_eq!(ds["type"]["class"], "H5T_COMPOUND");
assert_eq!(ds["type"]["fields"][0]["name"], "r");
assert_eq!(ds["type"]["fields"][1]["type"]["base"], "H5T_IEEE_F64LE");
assert_eq!(ds["value"], serde_json::json!([2.5, -0.5]));
}
#[test]
fn diff_compares_complex_values() {
let path = fixture("native_complex_hdf5_2.h5");
let p = path.to_str().unwrap();
// Same values, different byte order: no differences.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", p, p, "/c128", "/c128be"])
.output()
.unwrap();
assert!(out.status.success(), "{out:?}");
// One element differs in its imaginary part: one difference, exit
// status 1, as h5diff 2.2.0 reports it.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", "-r", p, p, "/c128", "/c128x"])
.output()
.unwrap();
assert_eq!(out.status.code(), Some(1), "{out:?}");
let text = String::from_utf8_lossy(&out.stdout);
assert!(text.contains("1 difference(s) found"), "{text}");
// h5diff 2.2.0: `[ 1 1 ] 0+1i 1+1i 1+0i`.
let row = text.lines().find(|l| l.starts_with("[ 1 1 ]")).unwrap();
assert_eq!(
row.split_whitespace().collect::<Vec<_>>(),
["[", "1", "1", "]", "0+1i", "1+1i", "1+0i"]
);
}
+85 -7
View File
@@ -312,11 +312,14 @@ impl Reader {
let base = array_base(dt);
let is_array = !std::ptr::eq(base, dt);
Ok(match base {
// Read as the part type: an array or complex element is its
// parts stored one after another (a complex number's real, then
// imaginary part).
Datatype::FloatingPoint { size, .. } if *size <= 4 => {
Data::F32(data_read::read_as_f32(raw, dt).map_err(err)?)
Data::F32(data_read::read_as_f32(raw, base).map_err(err)?)
}
Datatype::FloatingPoint { .. } => {
Data::F64(data_read::read_as_f64(raw, dt).map_err(err)?)
Data::F64(data_read::read_as_f64(raw, base).map_err(err)?)
}
Datatype::FixedPoint { size, signed, .. } => {
let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err);
@@ -393,10 +396,13 @@ fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T
.collect()
}
/// Innermost element type of (possibly nested) array datatypes.
/// Innermost element type of (possibly nested) array datatypes; for a
/// native complex type, its part type (each element holds two).
fn array_base(dt: &Datatype) -> &Datatype {
match dt {
Datatype::Array { base_type, .. } => array_base(base_type),
Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => {
array_base(base_type)
}
_ => dt,
}
}
@@ -411,6 +417,12 @@ fn element_shape(dt: &Datatype) -> Vec<u64> {
dims.extend(element_shape(base_type));
dims
}
// `[re, im]`: a complex element reads as its two parts.
Datatype::Complex { base_type, .. } => {
let mut dims = vec![2];
dims.extend(element_shape(base_type));
dims
}
_ => Vec::new(),
}
}
@@ -487,9 +499,7 @@ pub fn describe(dt: &Datatype) -> String {
}
}
match dt {
Datatype::Complex { size, base_type } => {
describe(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::Complex { base_type, .. } => format!("complex<{}>", describe(base_type)),
Datatype::FixedPoint {
size,
signed,
@@ -681,6 +691,74 @@ mod tests {
assert!(e.contains("not supported"), "{e}");
}
#[test]
fn native_complex_reads_as_re_im_pairs() {
let z = [[1.0f64, -2.0], [0.5, 0.0], [-0.0, 3.25], [f64::MAX, 1e-300]];
let mut b = FileBuilder::new();
b.create_dataset("z")
.with_native_complex_f64_data(&z)
.with_shape(&[2, 2]);
b.create_dataset("z32")
.with_native_complex_f32_data(&[[1.5f32, -2.5]]);
let r = Reader::open(b.finish().unwrap()).unwrap();
let i = r.info("z").unwrap();
assert_eq!(
(i.shape, i.dtype, i.element_shape),
(vec![2, 2], "complex<f64>".to_string(), vec![2])
);
let all = r.read("z", None).unwrap();
assert_eq!(all.shape, vec![2, 2, 2]);
assert_eq!(all.data, Data::F64(z.iter().flatten().copied().collect()));
let slab = Hyperslab {
start: vec![1, 0],
count: vec![1, 2],
stride: None,
block: None,
};
let part = r.read("z", Some(&slab)).unwrap();
assert_eq!(part.shape, vec![1, 2, 2]);
assert_eq!(part.data, Data::F64(vec![-0.0, 3.25, f64::MAX, 1e-300]));
let z32 = r.read("z32", None).unwrap();
assert_eq!(
(z32.shape, z32.data),
(vec![1, 2], Data::F32(vec![1.5, -2.5]))
);
}
#[test]
fn native_complex_written_by_libhdf5_2() {
// Written by h5py 3.16 / libhdf5 2.0.0 (gen_native_complex.py).
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5/tests/fixtures/native_complex_hdf5_2.h5"
);
let r = Reader::open(std::fs::read(path).unwrap()).unwrap();
let be = r.read("c128be", None).unwrap();
assert_eq!(be.shape, vec![2, 3, 2]);
assert_eq!(be.data, r.read("c128", None).unwrap().data);
assert_eq!(r.info("c128be").unwrap().dtype, "complex<f64 (big-endian)>");
// binary16 parts widen to f32.
assert_eq!(
r.read("c32", None).unwrap().data,
Data::F32(vec![1.0, -2.0, 0.5, 65504.0, 0.0, -0.0])
);
// An array of complex: array dimensions, then the two parts.
let i = r.info("array").unwrap();
assert_eq!(
(i.dtype.as_str(), i.element_shape),
("array[2]<complex<f32>>", vec![2, 2])
);
assert_eq!(
r.read("array", None).unwrap().data,
Data::F32(vec![1.0, 2.0, 3.0, 4.0, -1.0, -1.0, 0.0, 0.0])
);
let s = r.read("scalar", None).unwrap();
assert_eq!((s.shape, s.data), (vec![2], Data::F64(vec![2.5, -0.5])));
// A compound holding a complex member is still refused.
let e = r.read("compound", None).unwrap_err();
assert!(e.contains("compound{z: complex<f64>, k: i64}"), "{e}");
}
#[test]
fn garbage_is_an_error() {
assert!(Reader::open(vec![0u8; 64]).is_err());
@@ -0,0 +1,46 @@
//! Write one dataset of `u8` stored as a single chunk, for measuring the
//! writer's peak memory (`/usr/bin/time -v`):
//!
//! ```text
//! cargo run --release -p clawhdf5 --example write_one_chunk -- BYTES OUT.h5 [slice|owned] [deflate|lz4|zstd]
//! ```
//!
//! `slice` passes the data by reference (`with_u8_data`, which copies it),
//! `owned` hands the builder the vector (`with_u8_data_owned`). The data is
//! `i % 251`, one byte per element; with no filter the chunk is stored raw.
fn main() {
let args: Vec<String> = std::env::args().collect();
let n: u64 = args[1].parse().expect("BYTES");
let out = &args[2];
let mode = args.get(3).map_or("slice", String::as_str);
let filter = args.get(4).map(String::as_str);
let data: Vec<u8> = (0..n).map(|i| (i % 251) as u8).collect();
let mut b = clawhdf5::FileBuilder::new();
let ds = b.create_dataset("x");
match mode {
"slice" => {
ds.with_u8_data(&data);
}
"owned" => {
ds.with_u8_data_owned(data);
}
m => panic!("mode {m}"),
}
ds.with_chunks(&[n]);
match filter {
None => {}
Some("deflate") => {
ds.with_deflate(1);
}
#[cfg(feature = "lz4")]
Some("lz4") => {
ds.with_lz4();
}
#[cfg(feature = "zstd")]
Some("zstd") => {
ds.with_zstd(1);
}
Some(f) => panic!("filter {f} (not built in?)"),
}
b.write(out).unwrap();
}
+1 -4
View File
@@ -408,10 +408,7 @@ impl<'t> ChunkedEdit<'t> {
"chunk rank differs from the dataspace".into(),
));
}
let cd: Vec<u64> = chunk_dimensions[..rank]
.iter()
.map(|&c| u64::from(c))
.collect();
let cd: Vec<u64> = chunk_dimensions[..rank].to_vec();
// Chunks of 4 GiB or more (HDF5 2.0 writes them with layout message
// version 5) are read, but not rewritten: the editor holds a chunk
// it rewrites in memory, compressed and not. Writing values, and a
+4 -1
View File
@@ -1342,7 +1342,10 @@ impl<'f> Dataset<'f> {
) -> Result<(Vec<T>, Vec<T>), Error> {
let dt = self.datatype()?;
let is_complex = match &dt {
// Class 11 parses to this same `{r, i}` compound.
Datatype::Complex { size, base_type } => {
matches!(base_type.as_ref(), Datatype::FloatingPoint { .. })
&& *size == 2 * base_type.type_size()
}
Datatype::Compound { size, members } => {
matches!(members.as_slice(), [r, i]
if r.name == "r" && i.name == "i"
+11
View File
@@ -41,6 +41,13 @@ pub enum DType {
Array(Box<DType>, Vec<u32>),
/// Variable-length UTF-8 string stored via a global heap reference.
VariableLengthString,
/// HDF5 2.0 native complex number (datatype class 11) over the given
/// floating-point part type: `Complex(F32)` is numpy `complex64`,
/// `Complex(F64)` `complex128`, and a half-precision part is
/// `Complex(Other("float16"))`. h5py's default complex encoding, the
/// compound `{r, i}`, stays a [`DType::Compound`]; `read_complex_f64`
/// and `read_complex_f32` read both.
Complex(Box<DType>),
/// Catch-all for HDF5 datatypes that do not map to a specific variant.
Other(String),
}
@@ -60,6 +67,7 @@ impl fmt::Display for DType {
DType::U64 => write!(f, "u64"),
DType::String => write!(f, "string"),
DType::VariableLengthString => write!(f, "vlen_string"),
DType::Complex(base) => write!(f, "complex<{base}>"),
DType::Compound(fields) => {
write!(f, "compound{{")?;
for (i, (name, dt)) in fields.iter().enumerate() {
@@ -148,6 +156,9 @@ pub(crate) fn classify_datatype(dt: &clawhdf5_format::datatype::Datatype) -> DTy
base_type,
dimensions,
} => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()),
Datatype::Complex { base_type, .. } => {
DType::Complex(Box::new(classify_datatype(base_type)))
}
_ => DType::Other(format!("{dt:?}")),
}
}
+31 -8
View File
@@ -137,10 +137,15 @@ impl FileBuilder {
Ok(self.writer.finish()?)
}
/// Serialize and write the file to the given path.
/// Serialize and write the file to the given path. The file is written
/// as it is serialized, not assembled in memory first: unfiltered
/// chunks and contiguous data go straight from the datasets' data.
pub fn write<P: AsRef<std::path::Path>>(self, path: P) -> Result<(), Error> {
let bytes = self.finish()?;
write_file_atomically(path.as_ref(), &bytes).map_err(Error::Io)
use std::io::Write;
write_file_atomically_with(path.as_ref(), |f| {
self.writer
.finish_with(|bytes| f.write_all(bytes).map_err(Error::Io))
})
}
}
@@ -259,8 +264,19 @@ pub fn create_datasets_parallel(specs: Vec<DatasetSpec>) -> Result<Vec<u8>, Erro
/// file or the complete new one — never a truncated mix. `std::fs::write`
/// truncates the destination first, so dying mid-write used to destroy the
/// existing file.
#[cfg(test)]
fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Result<()> {
use std::io::Write;
write_file_atomically_with(path, |f| f.write_all(bytes))
}
/// [`write_file_atomically`] with the contents written by `write` (into a
/// buffered writer over the temporary file).
fn write_file_atomically_with<E: From<std::io::Error>>(
path: &std::path::Path,
write: impl FnOnce(&mut std::io::BufWriter<std::fs::File>) -> Result<(), E>,
) -> Result<(), E> {
use std::io::Write;
// Same directory as the target, so the rename stays on one filesystem.
let mut tmp_name = path
@@ -272,11 +288,18 @@ fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Resul
tmp_name.push(format!(".tmp-{}", std::process::id()));
let tmp_path = path.with_file_name(tmp_name);
let result = (|| {
let mut f = std::fs::File::create(&tmp_path)?;
f.write_all(bytes)?;
f.sync_all()?;
std::fs::rename(&tmp_path, path)
let result = (|| -> Result<(), E> {
// 64 KiB, below glibc's 128 KiB mmap threshold: a 1 MiB buffer was
// mapped afresh by each write once the threshold had not risen, and
// its page faults made a 1 MiB chunked write 1.6x slower (criterion
// `write_2d_chunked/512x512` after the smaller cases, 2026-09-29).
// Writes larger than the buffer go straight to the file.
let mut f = std::io::BufWriter::with_capacity(64 << 10, std::fs::File::create(&tmp_path)?);
write(&mut f)?;
f.flush()?;
f.get_ref().sync_all()?;
std::fs::rename(&tmp_path, path)?;
Ok(())
})();
if result.is_err() {
let _ = std::fs::remove_file(&tmp_path);
+57 -2
View File
@@ -2,6 +2,12 @@
python gen_huge_chunks.py filtered OUT.h5 # huge_chunks_filtered.h5
python gen_huge_chunks.py unfiltered OUT.h5 # generated at test time
python gen_huge_chunks.py dims OUT.h5 # huge_chunk_dims.h5
python gen_huge_chunks.py dims-unfiltered OUT.h5 # generated at test time
`dims` and `dims-unfiltered` (see `make_dims`) hold `u1` datasets whose
chunk dimension is N8 = 2**32 + 7: a chunk dimension of 2**32 or more,
which a layout message of version 5 stores in 5 bytes.
libhdf5 2.x writes a chunk of more than 0xFFFFFFFF bytes with layout message
version 5 (`H5D__chunk_construct`: "chunk size > 4GB requires
@@ -38,7 +44,8 @@ allocated early (see `make`).
libhdf5 holds a whole filtered chunk in memory while it writes it, so the
`filtered` run needs about 4 GiB of memory and 100 s (tank).
Generated 2026-09-28 with h5py 3.16.0 (libhdf5 2.0.0).
huge_chunks_filtered.h5 was generated 2026-09-28 and huge_chunk_dims.h5
2026-09-29, both with h5py 3.16.0 (libhdf5 2.0.0).
"""
import sys
@@ -92,12 +99,60 @@ def make(fid, name, index, filtered):
dsid.close()
N8 = 2**32 + 7
def make_dims(fid, name, index, filtered):
"""A `u1` dataset whose chunks are N8 elements (a chunk dimension of
2**32 or more). `d[0:10] = 0..9` and `d[N8-10:N8] = 100..109`; where
there is a second chunk (`farray`, `earray`: shape N8 + 10),
`d[N8:N8+10] = 200..209`.
filtered (`dims`, committed): deflate level 9 twice, fill value 7;
`single` (Single Chunk) and `earray` (Extensible Array, maxshape None).
Writing it holds a 4 GiB chunk: 4.1 GiB peak resident memory and about
a minute (tank, 2026-09-29).
unfiltered (`dims-unfiltered`): fill time "never", sparse; `single`
(allocated early, as in `make`) and `farray` (Fixed Array).
"""
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
if index == "single":
shape, maxshape = (N8,), (N8,)
elif index == "farray":
shape, maxshape = (N8 + 10,), (N8 + 10,)
elif index == "earray":
shape, maxshape = (N8 + 10,), (UNLIM,)
dcpl.set_chunk((N8,))
if filtered:
dcpl.set_deflate(9)
dcpl.set_deflate(9)
dcpl.set_fill_value(np.array(7, dtype="u1"))
else:
dcpl.set_fill_time(h5py.h5d.FILL_TIME_NEVER)
if index == "single":
dcpl.set_alloc_time(h5py.h5d.ALLOC_TIME_EARLY)
space = h5py.h5s.create_simple(shape, maxshape)
dsid = h5py.h5d.create(fid, name.encode(), h5py.h5t.STD_U8LE, space, dcpl=dcpl)
ds = h5py.Dataset(dsid)
ds[0:10] = np.arange(10, dtype="u1")
ds[N8 - 10 : N8] = np.arange(100, 110, dtype="u1")
if shape[0] > N8:
ds[N8 : N8 + 10] = np.arange(200, 210, dtype="u1")
dsid.close()
def main():
mode, out = sys.argv[1], sys.argv[2]
filtered = {"filtered": True, "unfiltered": False}[mode]
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
fapl.set_libver_bounds(h5py.h5f.LIBVER_EARLIEST, h5py.h5f.LIBVER_V200)
fid = h5py.h5f.create(out.encode(), h5py.h5f.ACC_TRUNC, fapl=fapl)
if mode in ("dims", "dims-unfiltered"):
filtered = mode == "dims"
for index in ["single", "earray"] if filtered else ["single", "farray"]:
make_dims(fid, index, index, filtered)
fid.close()
return
filtered = {"filtered": True, "unfiltered": False}[mode]
indexes = ["single", "farray", "earray", "btree2"]
if not filtered:
indexes.insert(1, "implicit")
+75
View File
@@ -0,0 +1,75 @@
"""Generate native_complex_hdf5_2.h5: HDF5 2.0 native complex (class 11).
Datasets (every one class 11, or a type holding it):
c64 H5T_COMPLEX_IEEE_F32LE, 4 values incl. -0, inf, nan, subnormal
c128 H5T_COMPLEX_IEEE_F64LE, 2 x 3
c128be H5T_COMPLEX_IEEE_F64BE, the same values
c128x H5T_COMPLEX_IEEE_F64LE, c128 with element (1,1) changed
c32 H5T_COMPLEX_IEEE_F16LE, 3 values
scalar H5T_COMPLEX_IEEE_F64LE, scalar dataspace
compound {z: H5T_COMPLEX_IEEE_F64LE @0, k: H5T_STD_I64LE @16}
array H5T_ARRAY [2] of H5T_COMPLEX_IEEE_F32LE, 2 elements
plus the root attribute `zattr` (H5T_COMPLEX_IEEE_F64LE, 2 values).
Written with h5py 3.16 (libhdf5 2.0.0) through its low-level API, the file
type as the memory type (no conversion). Re-run only if the fixture ever
needs regenerating (generated 2026-09-29):
python gen_native_complex.py native_complex_hdf5_2.h5
The h5dump / h5ls 2.2.0 reference output the tests compare with was made
with the tools of the libhdf5 2.2.0 build described in gen_mx_floats.py:
h5dump native_complex_hdf5_2.h5 > native_complex_hdf5_2.ddl
h5ls -r -v native_complex_hdf5_2.h5 > native_complex_hdf5_2.ls
"""
import sys
import h5py
import numpy as np
from h5py import h5a, h5d, h5s, h5t
out = sys.argv[1]
f = h5py.File(out, "w", libver="latest")
def ds(name, tid, arr, shape=None):
shape = arr.shape if shape is None else shape
sp = h5s.create_simple(shape) if shape != () else h5s.create(h5s.SCALAR)
d = h5d.create(f.id, name.encode(), tid, sp)
d.write(h5s.ALL, h5s.ALL, np.ascontiguousarray(arr), mtype=tid)
nan, inf = float("nan"), float("inf")
c64 = np.array(
[1.5 - 2j, complex(0.0, -0.0), complex(inf, nan), complex(-inf, 1e-40)],
dtype=np.complex64,
)
ds("c64", h5t.COMPLEX_IEEE_F32LE, c64)
c128 = np.array(
[[1 + 2j, -3.5 + 4e300j, 0.1 + 0.2j], [5e-324 - 0j, 1j, -1]], dtype=np.complex128
)
ds("c128", h5t.COMPLEX_IEEE_F64LE, c128)
ds("c128be", h5t.COMPLEX_IEEE_F64BE, c128.astype(">c16"))
c128x = c128.copy()
c128x[1, 1] = 1 + 1j
ds("c128x", h5t.COMPLEX_IEEE_F64LE, c128x)
half = np.array([1.0, -2.0, 0.5, 65504.0, 0.0, -0.0], dtype=np.float16)
ds("c32", h5t.COMPLEX_IEEE_F16LE, half, shape=(3,))
ds("scalar", h5t.COMPLEX_IEEE_F64LE, np.array(2.5 - 0.5j), shape=())
ct = h5t.create(h5t.COMPOUND, 24)
ct.insert(b"z", 0, h5t.COMPLEX_IEEE_F64LE)
ct.insert(b"k", 16, h5t.STD_I64LE)
cd = np.array([(1 + 1j, 7), (-2.5 + 0j, -8)], dtype=[("z", "<c16"), ("k", "<i8")])
ds("compound", ct, cd)
at = h5t.array_create(h5t.COMPLEX_IEEE_F32LE, (2,))
ad = np.array([[1 + 2j, 3 + 4j], [-1 - 1j, 0j]], dtype=np.complex64)
ds("array", at, ad, shape=(2,))
a = h5a.create(f.id, b"zattr", h5t.COMPLEX_IEEE_F64LE, h5s.create_simple((2,)))
a.write(np.array([0.5 - 1.5j, 2j]), mtype=h5t.COMPLEX_IEEE_F64LE)
f.close()
Binary file not shown.
@@ -0,0 +1,80 @@
HDF5 "native_complex_hdf5_2.h5" {
GROUP "/" {
ATTRIBUTE "zattr" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): 0.5-1.5i, 0+2i
}
}
DATASET "array" {
DATATYPE H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): [ 1+2i, 3+4i ], [ -1-1i, 0+0i ]
}
}
DATASET "c128" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128be" {
DATATYPE H5T_COMPLEX_IEEE_F64BE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128x" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 1+1i, -1+0i
}
}
DATASET "c32" {
DATATYPE H5T_COMPLEX_IEEE_F16LE
DATASPACE SIMPLE { ( 3 ) / ( 3 ) }
DATA {
(0): 1+-2i, 0.5+65504i, 0+-0i
}
}
DATASET "c64" {
DATATYPE H5T_COMPLEX_IEEE_F32LE
DATASPACE SIMPLE { ( 4 ) / ( 4 ) }
DATA {
(0): 1.5-2i, 0-0i, inf+nani, -inf+9.99995e-41i
}
}
DATASET "compound" {
DATATYPE H5T_COMPOUND {
H5T_COMPLEX_IEEE_F64LE "z";
H5T_STD_I64LE "k";
}
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): {
1+1i,
7
},
(1): {
-2.5+0i,
-8
}
}
}
DATASET "scalar" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SCALAR
DATA {
(0): 2.5-0.5i
}
}
}
}
Binary file not shown.
+348 -1
View File
@@ -81,8 +81,15 @@ fn layout_of(
let lm = msg(MessageType::DataLayout);
let layout = DataLayout::parse(&lm.data, sb.offset_size, sb.length_size).unwrap();
let space = Dataspace::parse(&msg(MessageType::Dataspace).data, sb.length_size).unwrap();
// The element size is the layout's last dimension.
let elem = match &layout {
DataLayout::Chunked {
chunk_dimensions, ..
} => *chunk_dimensions.last().unwrap() as usize,
_ => 8,
};
let (chunks, _) =
list_chunks(data, &layout, &space, 8, sb.offset_size, sb.length_size).unwrap();
list_chunks(data, &layout, &space, elem, sb.offset_size, sb.length_size).unwrap();
(lm.data[0], layout, space, chunks)
}
@@ -516,3 +523,343 @@ print("ok")
}
}
}
// ---- Chunk dimensions of 2^32 or more ----
/// Elements per chunk of the `u8` datasets below: a chunk dimension of
/// 2^32 or more (4 GiB + 7 bytes), which a version-5 layout message stores
/// in 5 bytes.
const N8: u64 = (1 << 32) + 7;
fn dims_fixture() -> PathBuf {
Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures/huge_chunk_dims.h5")
}
/// `u8` values of `sel` in dataset `name` of `file`.
fn sel_u8(file: &File, name: &str, start: u64, count: u64) -> Vec<u8> {
let ds = file.dataset(name).unwrap();
let s = Selection::Hyperslab {
start: vec![start],
stride: vec![1],
count: vec![count],
block: vec![1],
};
ds.read_selection(&s).unwrap()
}
fn u8s(range: std::ops::Range<u8>) -> Vec<u8> {
range.collect()
}
/// `fixtures/huge_chunk_dims.h5` (libhdf5 2.0.0 through h5py 3.16,
/// `gen_huge_chunks.py dims`): `u8` datasets chunked by N8 = 2^32 + 7, in a
/// Single Chunk and an Extensible Array index. Their layouts parse to the
/// whole dimension (it was refused: chunk dimensions were `u32`) and their
/// chunks list at the right origins.
#[test]
fn huge_chunk_dims_list() {
let data = std::fs::read(dims_fixture()).unwrap();
for (name, index, shape, origins) in [
("single", 1u8, N8, &[0u64][..]),
("earray", 4, N8 + 10, &[0, N8]),
] {
let (version, layout, space, mut chunks) = layout_of(&data, name);
assert_eq!(version, 5, "{name}");
let DataLayout::Chunked {
chunk_dimensions,
chunk_index_type,
..
} = &layout
else {
panic!("{name}: {layout:?}");
};
assert_eq!(chunk_dimensions, &[N8, 1], "{name}");
assert_eq!(*chunk_index_type, Some(index), "{name}");
assert_eq!(space.dimensions, [shape], "{name}");
chunks.sort_by_key(|c| c.offsets[0]);
let got: Vec<u64> = chunks.iter().map(|c| c.offsets[0]).collect();
assert_eq!(got, origins, "{name}");
}
}
/// Decoding the fixture's chunks (4 GiB each, one at a time): written
/// values, the fill value (7) around them, and the second chunk.
#[test]
fn huge_chunk_dims_read() {
if !heavy() {
return;
}
let file = File::open(dims_fixture()).unwrap();
let mut first = u8s(0..10);
first.extend([7, 7]);
let mut last = vec![7, 7];
last.extend(u8s(100..110));
for name in ["single", "earray"] {
assert_eq!(sel_u8(&file, name, 0, 12), first, "{name}");
assert_eq!(sel_u8(&file, name, N8 - 12, 12), last, "{name}");
}
let mut across = u8s(108..110);
across.extend(u8s(200..210));
assert_eq!(sel_u8(&file, "earray", N8 - 2, 12), across);
}
/// Unfiltered chunks of N8 bytes that libhdf5 2.0.0 wrote (a sparse file,
/// `gen_huge_chunks.py dims-unfiltered`), in a Single Chunk and a Fixed
/// Array index: read through a memory map, and through positioned reads
/// that fetch only the rows selected.
#[test]
fn huge_chunk_dims_unfiltered_read() {
if !heavy() {
return;
}
let dir = scratch();
let Some(path) = generate(dir.path(), "dims-unfiltered") else {
return;
};
let check = |file: &File| {
let mut first = u8s(0..10);
first.extend([0, 0]);
let mut last = vec![0, 0];
last.extend(u8s(100..110));
for name in ["single", "farray"] {
assert_eq!(sel_u8(file, name, 0, 12), first, "{name}");
assert_eq!(sel_u8(file, name, N8 - 12, 12), last, "{name}");
}
let mut across = u8s(108..110);
across.extend(u8s(200..210));
assert_eq!(sel_u8(file, "farray", N8 - 2, 12), across);
};
check(&File::open(&path).unwrap());
let file = std::fs::File::open(&path).unwrap();
let len = file.metadata().unwrap().len();
let storage = std::sync::Arc::new(Counting {
file,
len,
read: 0.into(),
});
check(&File::open_storage(storage.clone()).unwrap());
let read = storage.read.load(std::sync::atomic::Ordering::Relaxed);
assert!(read < 1 << 20, "{read} bytes read");
}
/// Run `script` under h5py with `path` as `sys.argv[1]`; skipped without
/// h5py (a failure under `CLAWHDF5_REQUIRE_INTEROP=1`).
fn h5py_check(script: &str, path: &Path) {
if !python_available() {
assert!(!interop_required(), "h5py not available");
eprintln!("SKIP: python3 with h5py not available");
return;
}
let out = Command::new(python())
.arg("-c")
.arg(script)
.arg(path)
.output()
.unwrap();
assert!(
out.status.success(),
"h5py:\n{}",
String::from_utf8_lossy(&out.stderr)
);
}
/// `h5dump` 2.x's output for `args` on `path`, checked to succeed without
/// errors; `None` without `CLAWHDF5_H5DUMP2`.
fn h5dump2_text(args: &[&str], path: &Path) -> Option<String> {
let h5dump = h5dump2()?;
let out = Command::new(&h5dump).args(args).arg(path).output().unwrap();
let text = format!(
"{}{}",
String::from_utf8_lossy(&out.stdout),
String::from_utf8_lossy(&out.stderr)
);
assert!(out.status.success(), "h5dump {args:?}:\n{text}");
assert!(!text.contains("rror"), "h5dump {args:?}:\n{text}");
Some(text)
}
/// clawhdf5 writes deflated `u8` chunks of N8 elements (a chunk dimension
/// of 2^32 or more) in a Single Chunk and an Extensible Array index; it,
/// h5py (libhdf5 2.0.0) and h5dump 2.2.0 (layouts only: it cannot inflate
/// chunks over 4 GiB, see known-issues) read them.
#[test]
fn writer_huge_chunk_dims_round_trip() {
if !heavy() {
return;
}
let dir = scratch();
let path = dir.path().join("dims.h5");
let values = u8s(0..10);
let mut b = clawhdf5::FileBuilder::new();
for (name, max) in [("single", N8), ("earray", u64::MAX)] {
b.create_dataset(name)
.with_u8_data(&values)
.with_maxshape(&[max])
.with_chunks(&[N8])
.with_deflate(1);
}
b.write(&path).unwrap();
let data = std::fs::read(&path).unwrap();
// Deflate level 1 leaves about 40 MiB of each 4 GiB chunk of zeros.
assert!(data.len() < 256 << 20, "{} bytes", data.len());
for (name, index) in [("single", 1), ("earray", 4)] {
let (version, layout, _, chunks) = layout_of(&data, name);
assert_eq!(version, 5, "{name}");
assert!(
matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(t), .. }
if *t == index && chunk_dimensions == &[N8, 1]),
"{name}: {layout:?}"
);
assert_eq!(chunks.len(), 1, "{name}");
}
let file = File::open(&path).unwrap();
for name in ["single", "earray"] {
assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}");
}
drop(file);
h5py_check(
r#"
import sys, h5py
with h5py.File(sys.argv[1], "r") as f:
for name in ("single", "earray"):
d = f[name]
assert d.chunks == (2**32 + 7,), (name, d.chunks)
assert list(d[:]) == list(range(10)), (name, d[:])
"#,
&path,
);
for name in ["single", "earray"] {
if let Some(text) = h5dump2_text(&["-H", "-p", "-d", name], &path) {
assert!(text.contains("CHUNKED ( 4294967303 )"), "{name}:\n{text}");
}
}
}
/// An unfiltered chunk of N8 bytes (a chunk dimension of 2^32 or more)
/// written by clawhdf5 from data it was handed (`with_u8_data_owned`):
/// `FileBuilder::write` streams it to the file without copying it, so the
/// write needs about the data's own 4 GiB. Read back by clawhdf5, h5py
/// (a few elements) and h5dump 2.2.0.
#[test]
fn writer_unfiltered_huge_chunk_round_trip() {
if !heavy() {
return;
}
let dir = scratch();
let path = dir.path().join("raw.h5");
let value = |i: u64| (i % 251) as u8;
let mut b = clawhdf5::FileBuilder::new();
b.create_dataset("single")
.with_u8_data_owned((0..N8).map(value).collect())
.with_chunks(&[N8]);
b.write(&path).unwrap();
let file = File::open(&path).unwrap();
let data = file.as_bytes();
let (version, layout, space, chunks) = layout_of(data, "single");
assert_eq!(version, 5);
assert!(
matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(1), .. }
if chunk_dimensions == &[N8, 1]),
"{layout:?}"
);
assert_eq!(space.dimensions, [N8]);
assert_eq!(chunks.len(), 1);
assert_eq!(chunks[0].chunk_size, N8);
for start in [0, (1 << 32) - 3, N8 - 6] {
let want: Vec<u8> = (start..start + 6).map(value).collect();
assert_eq!(sel_u8(&file, "single", start, 6), want, "at {start}");
}
drop(file);
h5py_check(
r#"
import sys, h5py, numpy as np
n = 2**32 + 7
with h5py.File(sys.argv[1], "r") as f:
d = f["single"]
assert d.chunks == (n,) and d.shape == (n,) and d.compression is None
for start in (0, 2**32 - 3, n - 6):
want = [i % 251 for i in range(start, start + 6)]
assert list(d[start:start + 6]) == want, (start, d[start:start + 6])
"#,
&path,
);
let start = (1u64 << 32) - 3;
if let Some(text) = h5dump2_text(
&["-d", "single", "-s", &start.to_string(), "-c", "6"],
&path,
) {
let want: Vec<String> = (start..start + 6).map(|i| value(i).to_string()).collect();
assert!(text.contains(&want.join(", ")), "{text}");
}
}
/// LZ4 and Zstandard chunks of 4 GiB or more: written by clawhdf5, read
/// back by it and by h5py (through `hdf5plugin`'s LZ4 and Zstandard
/// filters when installed).
#[cfg(any(feature = "lz4", feature = "zstd"))]
#[test]
fn writer_huge_chunks_lz4_zstd_round_trip() {
if !heavy() {
return;
}
let dir = scratch();
let path = dir.path().join("codecs.h5");
let values = u8s(0..10);
let mut b = clawhdf5::FileBuilder::new();
let mut names = Vec::new();
#[cfg(feature = "lz4")]
{
b.create_dataset("lz4")
.with_u8_data(&values)
.with_maxshape(&[N8])
.with_chunks(&[N8])
.with_lz4();
names.push("lz4");
}
#[cfg(feature = "zstd")]
{
b.create_dataset("zstd")
.with_u8_data(&values)
.with_maxshape(&[N8])
.with_chunks(&[N8])
.with_zstd(3);
names.push("zstd");
}
b.write(&path).unwrap();
let data = std::fs::read(&path).unwrap();
assert!(data.len() < 256 << 20, "{} bytes", data.len());
for name in &names {
let (version, _, _, chunks) = layout_of(&data, name);
assert_eq!(version, 5, "{name}");
assert_eq!(chunks.len(), 1, "{name}");
}
drop(data);
let file = File::open(&path).unwrap();
for name in &names {
assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}");
}
drop(file);
h5py_check(
&format!(
r#"
import sys, h5py
try:
import hdf5plugin
except ImportError:
print("SKIP: no hdf5plugin")
sys.exit(0)
with h5py.File(sys.argv[1], "r") as f:
for name in {names:?}:
d = f[name]
assert d.chunks == (2**32 + 7,), (name, d.chunks)
assert list(d[0:10]) == list(range(10)), (name, d[0:10])
"#,
names = names
),
&path,
);
}
+19 -5
View File
@@ -1064,12 +1064,26 @@ fn complex_datasets_and_attributes_round_trip() {
let ds = file.dataset(name).unwrap();
assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}");
assert_eq!(ds.shape().unwrap(), vec![3]);
// Both encodings read as the same {r, i} compound.
assert_eq!(
ds.dtype().unwrap(),
// h5py's encoding is a compound, the native one a complex type.
let want = if name == "native64" {
DType::Complex(Box::new(DType::F32))
} else {
DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)])
);
};
assert_eq!(ds.dtype().unwrap(), want, "{name}");
}
assert_eq!(
file.dataset("native128").unwrap().dtype().unwrap(),
DType::Complex(Box::new(DType::F64))
);
assert_eq!(
file.dataset("native128").unwrap().raw_datatype().unwrap(),
clawhdf5::make_native_complex_f64_type()
);
assert_eq!(
DType::Complex(Box::new(DType::F64)).to_string(),
"complex<f64>"
);
for name in ["compound128", "native128"] {
same(
file.dataset(name).unwrap().read_complex_f64().unwrap(),
@@ -1094,7 +1108,7 @@ fn complex_datasets_and_attributes_round_trip() {
data,
}) => {
assert!(shape.is_empty());
assert_eq!(datatype.type_size(), 16);
assert_eq!(datatype, clawhdf5::make_native_complex_f64_type());
let fields =
clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap();
let part = |k: usize| {
+119 -31
View File
@@ -17,6 +17,11 @@ Checked against `main` at `9b5803f` on 2026-09-28.
|---|---|---|
| [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
@@ -241,7 +246,10 @@ Values are correct in every case; this is cost only. Selections other than
## Chunks of 4 GiB or more: limits
**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more
**Status:** open (documented 2026-09-28; updated 2026-09-29, when
[chunk dimensions of 2^32 or more](#chunk-dimensions-of-232-or-more-were-refused)
and [unfiltered writes](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)
were fixed). HDF5 2.0 writes a chunk of more
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
also makes a filtered chunk index element store the chunk's size in "size
@@ -253,34 +261,40 @@ The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
What remains:
- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under
`H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`)
and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec<u32>`.
A 4 GiB chunk of 1-byte elements needs one.
- **Filters the writer refuses for such a chunk** (`FilterError`, before
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
filters record sizes or block lengths in 32 bits, or cannot take such a
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
written; only shuffle + deflate is tested end to end (LZ4's framing and
the 256 MiB limit it had are unit-tested; Zstd is not tested at this
size).
- **Unfiltered chunks this size are written but not tested end to end:**
`FileBuilder` assembles the whole file in memory, several copies of the
chunk (well over the 12 GiB the tests may use). Their layout messages
are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`).
written and tested end to end (`writer_huge_chunks_round_trip`,
`writer_huge_chunk_dims_round_trip`,
`writer_huge_chunks_lz4_zstd_round_trip`: read back by clawhdf5 and
h5py 3.16, LZ4 and Zstd through hdf5plugin 7.1.0).
- **`FileEditor`** refuses to rewrite such chunks (see
[its limits](#in-place-modification-fileeditor-limits)).
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
twice when it is shuffled (the inflated and the unshuffled copy, as
libhdf5's filters do too); writing one holds the chunk and its shuffled
copy. A selection decodes only the chunks it touches; one of an
unfiltered chunk reads only the rows it selects, through a memory map or
libhdf5's filters do too). Writing a filtered chunk holds its compressed
copy and, where it is not a contiguous run of the dataset's data (a
chunk past the dataset's edge is zero-padded), the raw chunk; shuffled,
the shuffled copy too. An unfiltered chunk is not copied when it is such
a run (a dataset stored as one chunk of its own shape): writing one
unfiltered chunk of 2^32 + 7 bytes with `with_u8_data_owned` and
`FileBuilder::write` peaked at 4.00 GiB resident, the data itself
(`write_one_chunk 4294967303 OUT owned`, tank, 2026-09-29, commit
`0065c6b`; see [the fix](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)).
`FileWriter::finish` (a `Vec`) holds the data and the file.
A selection decodes only the chunks it touches; one of an unfiltered
chunk reads only the rows it selects, through a memory map or
positioned reads (`File::open_storage`): every selection of
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop
-- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py
included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection
reads of the double-deflated fixture alone, 4.0 GiB.
memory of `cargo test --release -p clawhdf5 --features lz4,zstd --test
huge_chunks_interop -- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1`
(all twelve tests, h5py and h5dump included) was 10.04 GiB (tank,
2026-09-29, commit `0065c6b`), in `writer_huge_chunks_lz4_zstd_round_trip`
(the other tests, 8.2 GiB or less; the first run of the suite, six
tests, 8.06 GiB on 2026-09-28). Which step of that test holds 10 GiB
(clawhdf5's LZ4 or Zstd write or read, or hdf5plugin's) was not
measured separately.
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
reading or writing one is `FormatError::Overflow` ("exceeds the
addressable size" / "address space"); the file still opens and lists
@@ -315,15 +329,12 @@ wrong data.
decoded, and external references are an error (object references
decode). A multi-dimensional numeric attribute is returned as a flat
array (its shape is not reported; `AttrValue::Raw` carries the shape).
HDF5 2.0's native complex type (class 11) is read as the equivalent
`{r, i}` compound: values are right (`Dataset::read_complex_f64`, and
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs
ls`/`dump` and the browser reader show a compound where h5dump 2.x
prints `H5T_COMPLEX_IEEE_F64LE`, so `h5rs dump` of such a file does not
match h5dump 2.x (h5dump 1.14 cannot read it at all). Writing class 11
(added 2026-09-28) is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`); the Python bindings write complex
arrays as h5py's compound only.
HDF5 2.0's native complex type (class 11) is read as its own type since
2026-09-29 ([fixed](#hdf5-20-native-complex-numbers-read-as-a-r-i-compound));
writing it is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`), and the Python bindings write complex
arrays as h5py's compound only. `h5rs dump --json` writes it as the
`{r, i}` compound: hdf5-json (h5json 2.0.0) has no complex class.
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
that: libhdf5 fails only the first metadata read of an image it cannot
load and then reads the file's own (possibly stale) metadata, where we
@@ -533,8 +544,10 @@ browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27.
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included; not in Firefox or Safari.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type (HDF5 2.0 native complex datasets
read, as `[re, im]` pairs); attributes of those types come back
as `value: null` with their `dtype`. (VL strings read, with h5py's
values, through the same `VlResolver` as `File` and `h5rs`.)
- No Zstd or SZIP (both link C): such datasets fail with
@@ -616,7 +629,7 @@ given.
## NetCDF-4: files without dimension scales, types and order differed from netCDF-C
**Status:** fixed 2026-09-29 (branch `fix/netcdf-phony-dims`). Affected
**Status:** fixed 2026-09-29 (#30). Affected
every release (v2.1.0 to v2.7.0) and `main` after #27. Wrong metadata only
(names, order and types of dimensions and variables); values were read
right. Users: the dimensions of a file without dimension scales are now
@@ -669,6 +682,81 @@ netCDF4-python-written files with netCDF-C itself, and
`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in
[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)).
## HDF5 2.0 native complex numbers read as a `{r, i}` compound
**Status:** fixed 2026-09-29 (#31). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0)
and `main` until then. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
A class-11 datatype (`H5T_COMPLEX_IEEE_F64LE`, ...) parsed into the
equivalent compound `{r, i}`. Values were right (`read_complex_f64`,
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs ls`/`dump`
and the browser reader showed a compound, so `h5rs dump` did not match
h5dump 2.x, the browser refused the dataset, and a parsed type written
back out became a compound. `Datatype::parse` now returns
`Datatype::Complex` (in compounds, arrays and VL types too),
`Dataset::dtype()` is `DType::Complex(..)`, `h5rs` prints what h5dump,
h5ls and h5diff 2.2.0 print (checked on tank, 2026-09-29, against a file
written by h5py 3.16 / libhdf5 2.0.0:
`cargo test -p clawhdf5-tools --test native_complex_dump`), and the
browser reads `[re, im]` pairs.
**What users must do:** code that matched a native complex type as a
`Datatype::Compound` (or a `DType::Compound` of `r`/`i`) must match
`Datatype::Complex` / `DType::Complex`, or take the compound view with
`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a
`Complex` arm. Files are unchanged; nothing needs rewriting.
## Chunk dimensions of 2^32 or more were refused
**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). An error, never wrong
data. Nothing for users to do but upgrade; code that matches on
`DataLayout::Chunked::chunk_dimensions` gets `u64`s now.
libhdf5 2.x writes a chunk larger than 4 GiB with layout message version
5, whose dimensions take up to 8 bytes each, so a chunk dimension can be
2^32 or more (a 4 GiB chunk of 1-byte elements needs one).
`DataLayout::Chunked::chunk_dimensions` was `Vec<u32>`: such a file was
refused when opened for reading (`InvalidChunkDimensions`, "larger than
2^32 - 1"), and the writer refused such a chunk. Chunk dimensions are
`u64` throughout (format crate, facade, editor, `h5rs`, Python bindings);
a version-3 layout, whose dimensions are 4 bytes, is never written for
them (a chunk that large always takes version 5, as in libhdf5). Tests:
`huge_chunks_interop` (`huge_chunk_dims_list` always, over
`fixtures/huge_chunk_dims.h5`, which libhdf5 2.0.0 wrote: `u8` chunks of
2^32 + 7 in a Single Chunk and an Extensible Array index; with
`CLAWHDF5_HUGE_CHUNKS=1`, reads of it and of an unfiltered sparse file
h5py writes, and clawhdf5's own such chunks read back by h5py 3.16 and
h5dump 2.2.0), unit tests of the 5-byte encoding, and the wasm package
test (the file lists; reading such a chunk on wasm32 is
`FormatError::Overflow`).
## Writing a 4 GiB unfiltered chunk held several copies of it
**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). Memory only. To write
large data with the least memory, hand the builder the data
(`DatasetBuilder::with_u8_data_owned`) and write with
`FileBuilder::write` (or `FileWriter::finish_with`).
`FileWriter::finish` copied every chunk into a per-chunk buffer, laid the
chunks out into one buffer per pass (two passes, the first kept), then
copied everything into the file's buffer, and `FileBuilder::write` wrote
that buffer out: about five times a dataset's size for one unfiltered
chunk, six with `with_u8_data` (which copies its argument). Measured
before and after on tank, 2026-09-29, peak resident memory of
`write_one_chunk 1073741824 OUT <mode>` (`crates/clawhdf5/examples`, one
unfiltered 1 GiB chunk of `u8`): `owned` 5.0 GiB before (the writer at
`4260af4`), 1.0 GiB after (commit `0065c6b`); `slice` 6.0 GiB before,
2.0 GiB after. With 4 GiB + 7 bytes and `owned`: 4.00 GiB after. The
files written are byte for byte the same. Now a chunk that is a contiguous
run of the dataset's data is borrowed (and a filtered one compressed
straight from it), the layout refers to chunks instead of copying them,
and `FileBuilder::write` streams the file to disk
(`FileWriter::finish_with`); contiguous datasets are not copied either.
Test: `writer_unfiltered_huge_chunk_round_trip` (with
`CLAWHDF5_HUGE_CHUNKS=1`: 4 GiB + 7 bytes written and read back by
clawhdf5, h5py 3.16 and h5dump 2.2.0; the test process peaked at 4.00 GiB).
## NetCDF-4: variables' dimensions are guessed from sizes
+6 -2
View File
@@ -123,8 +123,12 @@ else throws an `Error` naming the type.
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in
`docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with
- HDF5 2.0 native complex datasets read as `[re, im]` pairs: `dtype`
`complex<f64>`, `elementShape` ending in `2`, the parts interleaved in
the typed array.
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque and variable-length-sequence datasets are refused with
an error. Attributes of those types are listed with
`value: null` and their `dtype`.
- No Zstd or SZIP filters (they link C): such a dataset fails with
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
+6
View File
@@ -93,6 +93,12 @@ expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
# 3-D: leading dimension held at 0, window over the last two.
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
# HDF5 2.0 native complex (written when h5py's libhdf5 is 2.0 or later):
# each cell is the [re, im] pair.
if "$PY" -c 'import sys, h5py; sys.exit(not getattr(h5py.get_config(), "has_native_complex", False))'; then
expect $H5 "/native_c128" '<dd>complex&lt;f64&gt;</dd>' '<dd>(3, 2)</dd>' '<td>[1, -1.5]</td>' \
'<td>[5, -7.5]</td>'
fi
# Unsupported type: an error, not values.
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
+22 -3
View File
@@ -80,6 +80,15 @@ with h5py.File(h5, "w") as f:
**hdf5plugin.Zstd())
comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")])
f.create_dataset("table", data=comp)
if getattr(h5py.get_config(), "has_native_complex", False):
# HDF5 2.0 native complex (class 11): read as [re, im] pairs.
from h5py import h5d, h5s, h5t
for name, t, dt in [(b"native_c128", h5t.COMPLEX_IEEE_F64LE, "<c16"),
(b"native_c64_be", h5t.COMPLEX_IEEE_F32BE, ">c8")]:
z = (np.arange(6) - 1.5j * np.arange(6)).reshape(3, 2).astype(dt)
d = h5d.create(f.id, name, t, h5s.create_simple(z.shape))
d.write(h5s.ALL, h5s.ALL, z, mtype=t)
g = f.create_group("sensors")
g.attrs["location"] = "lab"
g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4"))
@@ -105,6 +114,8 @@ def kind(dt):
return "strings"
if dt.subdtype is not None:
return kind(dt.subdtype[0])
if dt.kind == "c":
return "f64" if dt.itemsize == 16 else "f32"
if dt.kind == "f":
return "f64" if dt.itemsize == 8 else "f32"
if dt.kind in "iu":
@@ -120,20 +131,25 @@ def flat(a, k):
return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()]
if k.startswith(("i", "u")):
return [str(int(x)) for x in a.ravel()]
if a.dtype.kind == "c":
# [re, im] pairs, in order.
a = np.stack([a.real, a.imag], axis=-1)
return [float(x) for x in a.ravel()]
def entry(ds, slab=None):
k = kind(ds.dtype)
data = ds[()]
e = {"kind": k, "shape": list(np.shape(data)), "values": flat(data, k)}
# A complex element reads as its two parts: one more dimension.
pair = [2] if ds.dtype.kind == "c" else []
e = {"kind": k, "shape": list(np.shape(data)) + pair, "values": flat(data, k)}
if slab:
start, count, stride = slab
idx = tuple(slice(s, s + (c - 1) * st + 1, st)
for s, c, st in zip(start, count, stride))
part = ds[idx]
e["slab"] = {"start": start, "count": count, "stride": stride,
"shape": list(part.shape), "values": flat(part, k)}
"shape": list(part.shape) + pair, "values": flat(part, k)}
return e
@@ -176,7 +192,10 @@ def describe(path):
"datasets": sorted(sets)}
for n, o in members.items():
walk(key.rstrip("/") + "/" + n, o)
elif obj.dtype.names:
elif obj.dtype.names or (
obj.dtype.kind == "c"
and obj.id.get_type().get_class() != getattr(h5py.h5t, "COMPLEX", None)):
# h5py's own complex numbers are a compound {r, i}: refused.
expected["errors"][key] = "compound"
else:
expected["datasets"][key] = entry(obj, slab_for(obj))
+11 -2
View File
@@ -399,7 +399,7 @@ async function remoteTests() {
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": `bytes 0-511/${statSync(join(fixDir, "fixture.h5")).size}` })),
});
await f.read("/grid");
}, /the server sent 0-511/, "wrong range");
@@ -514,6 +514,15 @@ async function limitTests() {
`huge chunk: ${name}`);
}
hc.free();
// Likewise chunks whose dimension is 2^32 or more (u8 chunks of 2^32 + 7).
const dims = pkg.open(new Uint8Array(readFileSync(join(import.meta.dirname, "..", "..", "..",
"crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5"))));
eq(dims.list("/").map((e) => e.name), ["earray", "single"], "huge chunk dims: list");
for (const name of ["single", "earray"]) {
await fails(() => dims.readHyperslab(`/${name}`, [0], [4]), /exceeds this platform.s address space/,
`huge chunk dims: ${name}`);
}
dims.free();
}
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
@@ -550,7 +559,7 @@ async function floodTests() {
const range = new Headers(init.headers).get("Range");
if (range === "bytes=0-511") return fetch(url, init);
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/${statSync(join(fixDir, "fixture.h5")).size}` } });
},
}).catch((e) => e);
// The open itself may need a second range: flooded either way.