docs: chunk dimensions of 2^32 or more and copy-free 4 GiB writes

known-issues: two fixed entries (chunk dimensions of 2^32 or more were
refused; a 4 GiB unfiltered chunk was held several times when written),
the open "Chunks of 4 GiB or more: limits" updated (LZ4/Zstd tested end
to end, memory figures of 2026-09-29) and the open-issues table in sync.
CHANGELOG (Unreleased) entry with the public API changes and the peak
RSS measurements; README capability table. The wasm package test lists
the new fixture and expects an Overflow error reading its chunks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-29 20:47:40 -05:00
co-authored by Claude Opus 5.5
parent e4b3a76dc7
commit 728ffeef16
4 changed files with 156 additions and 21 deletions
+59
View File
@@ -119,6 +119,65 @@
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
hard-codes the fixture's length.
### Chunk dimensions of 2^32 or more; 4 GiB chunks written without copies (2026-09-29)
- **Chunk dimensions of 2^32 or more** (libhdf5 2.x writes them with
layout message version 5, up to 8 bytes per dimension; a 4 GiB chunk of
1-byte elements needs one) are read and written. They were refused
(`InvalidChunkDimensions`). **Public API changes:**
`DataLayout::Chunked::chunk_dimensions` is `Vec<u64>` (was `Vec<u32>`),
and the `chunk_dimensions` argument of `chunked_read::{collect_chunk_info_checked,
collect_chunk_info_checked_in, generate_implicit_chunks_in_grid}`,
`fixed_array::{read_fixed_array_chunks, read_fixed_array_chunks_in}`,
`extensible_array::{read_extensible_array_chunks,
read_extensible_array_chunks_in}`, and the `chunk_dims` argument of
`chunked_write::serialize_v4_single_chunk_pub`, are `&[u64]` (were
`&[u32]`). A chunk whose size overflows 64 bits is refused when a dataset is read
(`InvalidChunkDimensions`), and the writer refuses a chunk dimension of
0 up front. A version-3 layout (4-byte dimensions) is never written for
such a chunk: over 4 GiB, it always takes version 5, as in libhdf5.
- **Writing without copies:** `FileWriter` no longer copies chunks into a
buffer per pass and then into the file's buffer. A chunk that is a
contiguous run of the dataset's data (a dataset stored as one chunk of
its shape, blocks of whole rows) is borrowed when unfiltered and
compressed straight from it when filtered, and contiguous datasets are
not copied. New `FileWriter::finish_with(put)` hands the file out piece
by piece; `FileBuilder::write` uses it (through a 1 MiB `BufWriter`,
still to a temporary file renamed into place), so a file is never
assembled in memory. New `DatasetBuilder::with_u8_data_owned(Vec<u8>)`
takes the data without the copy `with_u8_data` makes. Peak resident
memory of one unfiltered 1 GiB chunk of `u8`
(`cargo run --release -p clawhdf5 --example write_one_chunk --
1073741824 OUT owned|slice`, tank, 2026-09-29): 5.0 GiB before (the
writer of `4260af4`), 1.0 GiB after with `owned`; 6.0 and 2.0 GiB with
`slice`; 4 GiB + 7 bytes `owned`, 4.00 GiB. Same bytes written. Write
speed was not re-measured (the machine was not idle): the A/B of
`FileBuilder::write` and `finish` on small and large files is still to
be run.
- `chunked_write::extract_chunk` (behind `split_into_chunks`) no longer
panics when the data is shorter than the shape; the missing part is
zeros, as it was meant to be.
- Tests: `crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5` (31 KB,
libhdf5 2.0.0 via h5py 3.16, `gen_huge_chunks.py dims`: `u8` chunks of
2^32 + 7 deflated twice, Single Chunk and Extensible Array), and in
`huge_chunks_interop`: `huge_chunk_dims_list` (always), and with
`CLAWHDF5_HUGE_CHUNKS=1` `huge_chunk_dims_read`,
`huge_chunk_dims_unfiltered_read` (a sparse file h5py writes, Single
Chunk and Fixed Array, mmap and positioned reads),
`writer_huge_chunk_dims_round_trip` (deflated, read back by clawhdf5,
h5py 3.16 and — layouts — h5dump 2.2.0),
`writer_unfiltered_huge_chunk_round_trip` (an unfiltered 4 GiB + 7 byte
chunk: clawhdf5, h5py, and h5dump 2.2.0 printing values past 2^32) and
`writer_huge_chunks_lz4_zstd_round_trip` (with the `lz4`/`zstd`
features; h5py through hdf5plugin 7.1.0). The whole opt-in suite
(`cargo test --release -p clawhdf5 --features lz4,zstd --test
huge_chunks_interop -- --test-threads=1`, h5py and h5dump 2.2.0
included) passed with a peak of 10.04 GiB resident (tank, 2026-09-29,
commit `0065c6b`). Unit tests: 5-byte dimension encoding, contiguous
chunk detection, `finish_with` against `finish`. The wasm package test
lists the new fixture and gets `FormatError::Overflow` reading it.
`docs/known-issues.md`: two fixed entries; the open "Chunks of 4 GiB or
more: limits" keeps the refused filters, the editor and memory.
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited