docs: chunk dimensions of 2^32 or more and copy-free 4 GiB writes

known-issues: two fixed entries (chunk dimensions of 2^32 or more were
refused; a 4 GiB unfiltered chunk was held several times when written),
the open "Chunks of 4 GiB or more: limits" updated (LZ4/Zstd tested end
to end, memory figures of 2026-09-29) and the open-issues table in sync.
CHANGELOG (Unreleased) entry with the public API changes and the peak
RSS measurements; README capability table. The wasm package test lists
the new fixture and expects an Overflow error reading its chunks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-29 20:47:40 -05:00
co-authored by Claude Opus 5.5
parent e4b3a76dc7
commit 728ffeef16
4 changed files with 156 additions and 21 deletions
+85 -19
View File
@@ -17,6 +17,11 @@ Checked against `main` at `9b5803f` on 2026-09-28.
|---|---|---|
| [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
@@ -241,7 +246,10 @@ Values are correct in every case; this is cost only. Selections other than
## Chunks of 4 GiB or more: limits
**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more
**Status:** open (documented 2026-09-28; updated 2026-09-29, when
[chunk dimensions of 2^32 or more](#chunk-dimensions-of-232-or-more-were-refused)
and [unfiltered writes](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)
were fixed). HDF5 2.0 writes a chunk of more
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
also makes a filtered chunk index element store the chunk's size in "size
@@ -253,34 +261,40 @@ The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
What remains:
- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under
`H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`)
and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec<u32>`.
A 4 GiB chunk of 1-byte elements needs one.
- **Filters the writer refuses for such a chunk** (`FilterError`, before
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
filters record sizes or block lengths in 32 bits, or cannot take such a
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
written; only shuffle + deflate is tested end to end (LZ4's framing and
the 256 MiB limit it had are unit-tested; Zstd is not tested at this
size).
- **Unfiltered chunks this size are written but not tested end to end:**
`FileBuilder` assembles the whole file in memory, several copies of the
chunk (well over the 12 GiB the tests may use). Their layout messages
are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`).
written and tested end to end (`writer_huge_chunks_round_trip`,
`writer_huge_chunk_dims_round_trip`,
`writer_huge_chunks_lz4_zstd_round_trip`: read back by clawhdf5 and
h5py 3.16, LZ4 and Zstd through hdf5plugin 7.1.0).
- **`FileEditor`** refuses to rewrite such chunks (see
[its limits](#in-place-modification-fileeditor-limits)).
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
twice when it is shuffled (the inflated and the unshuffled copy, as
libhdf5's filters do too); writing one holds the chunk and its shuffled
copy. A selection decodes only the chunks it touches; one of an
unfiltered chunk reads only the rows it selects, through a memory map or
libhdf5's filters do too). Writing a filtered chunk holds its compressed
copy and, where it is not a contiguous run of the dataset's data (a
chunk past the dataset's edge is zero-padded), the raw chunk; shuffled,
the shuffled copy too. An unfiltered chunk is not copied when it is such
a run (a dataset stored as one chunk of its own shape): writing one
unfiltered chunk of 2^32 + 7 bytes with `with_u8_data_owned` and
`FileBuilder::write` peaked at 4.00 GiB resident, the data itself
(`write_one_chunk 4294967303 OUT owned`, tank, 2026-09-29, commit
`0065c6b`; see [the fix](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)).
`FileWriter::finish` (a `Vec`) holds the data and the file.
A selection decodes only the chunks it touches; one of an unfiltered
chunk reads only the rows it selects, through a memory map or
positioned reads (`File::open_storage`): every selection of
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop
-- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py
included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection
reads of the double-deflated fixture alone, 4.0 GiB.
memory of `cargo test --release -p clawhdf5 --features lz4,zstd --test
huge_chunks_interop -- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1`
(all twelve tests, h5py and h5dump included) was 10.04 GiB (tank,
2026-09-29, commit `0065c6b`), in `writer_huge_chunks_lz4_zstd_round_trip`
(the other tests, 8.2 GiB or less; the first run of the suite, six
tests, 8.06 GiB on 2026-09-28). Which step of that test holds 10 GiB
(clawhdf5's LZ4 or Zstd write or read, or hdf5plugin's) was not
measured separately.
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
reading or writing one is `FormatError::Overflow` ("exceeds the
addressable size" / "address space"); the file still opens and lists
@@ -694,6 +708,58 @@ browser reads `[re, im]` pairs.
`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a
`Complex` arm. Files are unchanged; nothing needs rewriting.
## Chunk dimensions of 2^32 or more were refused
**Status:** fixed 2026-09-29 (branch `feat/huge-chunk-dims`, PR not yet
opened); affected every release (v2.1.0 to v2.7.0). An error, never wrong
data. Nothing for users to do but upgrade; code that matches on
`DataLayout::Chunked::chunk_dimensions` gets `u64`s now.
libhdf5 2.x writes a chunk larger than 4 GiB with layout message version
5, whose dimensions take up to 8 bytes each, so a chunk dimension can be
2^32 or more (a 4 GiB chunk of 1-byte elements needs one).
`DataLayout::Chunked::chunk_dimensions` was `Vec<u32>`: such a file was
refused when opened for reading (`InvalidChunkDimensions`, "larger than
2^32 - 1"), and the writer refused such a chunk. Chunk dimensions are
`u64` throughout (format crate, facade, editor, `h5rs`, Python bindings);
a version-3 layout, whose dimensions are 4 bytes, is never written for
them (a chunk that large always takes version 5, as in libhdf5). Tests:
`huge_chunks_interop` (`huge_chunk_dims_list` always, over
`fixtures/huge_chunk_dims.h5`, which libhdf5 2.0.0 wrote: `u8` chunks of
2^32 + 7 in a Single Chunk and an Extensible Array index; with
`CLAWHDF5_HUGE_CHUNKS=1`, reads of it and of an unfiltered sparse file
h5py writes, and clawhdf5's own such chunks read back by h5py 3.16 and
h5dump 2.2.0), unit tests of the 5-byte encoding, and the wasm package
test (the file lists; reading such a chunk on wasm32 is
`FormatError::Overflow`).
## Writing a 4 GiB unfiltered chunk held several copies of it
**Status:** fixed 2026-09-29 (branch `feat/huge-chunk-dims`, PR not yet
opened); affected every release (v2.1.0 to v2.7.0). Memory only. To write
large data with the least memory, hand the builder the data
(`DatasetBuilder::with_u8_data_owned`) and write with
`FileBuilder::write` (or `FileWriter::finish_with`).
`FileWriter::finish` copied every chunk into a per-chunk buffer, laid the
chunks out into one buffer per pass (two passes, the first kept), then
copied everything into the file's buffer, and `FileBuilder::write` wrote
that buffer out: about five times a dataset's size for one unfiltered
chunk, six with `with_u8_data` (which copies its argument). Measured
before and after on tank, 2026-09-29, peak resident memory of
`write_one_chunk 1073741824 OUT <mode>` (`crates/clawhdf5/examples`, one
unfiltered 1 GiB chunk of `u8`): `owned` 5.0 GiB before (the writer at
`4260af4`), 1.0 GiB after (commit `0065c6b`); `slice` 6.0 GiB before,
2.0 GiB after. With 4 GiB + 7 bytes and `owned`: 4.00 GiB after. The
files written are byte for byte the same. Now a chunk that is a contiguous
run of the dataset's data is borrowed (and a filtered one compressed
straight from it), the layout refers to chunks instead of copying them,
and `FileBuilder::write` streams the file to disk
(`FileWriter::finish_with`); contiguous datasets are not copied either.
Test: `writer_unfiltered_huge_chunk_round_trip` (with
`CLAWHDF5_HUGE_CHUNKS=1`: 4 GiB + 7 bytes written and read back by
clawhdf5, h5py 3.16 and h5dump 2.2.0; the test process peaked at 4.00 GiB).
## NetCDF-4: variables' dimensions are guessed from sizes