docs: chunk dimensions of 2^32 or more and copy-free 4 GiB writes
known-issues: two fixed entries (chunk dimensions of 2^32 or more were refused; a 4 GiB unfiltered chunk was held several times when written), the open "Chunks of 4 GiB or more: limits" updated (LZ4/Zstd tested end to end, memory figures of 2026-09-29) and the open-issues table in sync. CHANGELOG (Unreleased) entry with the public API changes and the peak RSS measurements; README capability table. The wasm package test lists the new fixture and expects an Overflow error reading its chunks. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -119,6 +119,65 @@
|
||||
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
|
||||
hard-codes the fixture's length.
|
||||
|
||||
### Chunk dimensions of 2^32 or more; 4 GiB chunks written without copies (2026-09-29)
|
||||
- **Chunk dimensions of 2^32 or more** (libhdf5 2.x writes them with
|
||||
layout message version 5, up to 8 bytes per dimension; a 4 GiB chunk of
|
||||
1-byte elements needs one) are read and written. They were refused
|
||||
(`InvalidChunkDimensions`). **Public API changes:**
|
||||
`DataLayout::Chunked::chunk_dimensions` is `Vec<u64>` (was `Vec<u32>`),
|
||||
and the `chunk_dimensions` argument of `chunked_read::{collect_chunk_info_checked,
|
||||
collect_chunk_info_checked_in, generate_implicit_chunks_in_grid}`,
|
||||
`fixed_array::{read_fixed_array_chunks, read_fixed_array_chunks_in}`,
|
||||
`extensible_array::{read_extensible_array_chunks,
|
||||
read_extensible_array_chunks_in}`, and the `chunk_dims` argument of
|
||||
`chunked_write::serialize_v4_single_chunk_pub`, are `&[u64]` (were
|
||||
`&[u32]`). A chunk whose size overflows 64 bits is refused when a dataset is read
|
||||
(`InvalidChunkDimensions`), and the writer refuses a chunk dimension of
|
||||
0 up front. A version-3 layout (4-byte dimensions) is never written for
|
||||
such a chunk: over 4 GiB, it always takes version 5, as in libhdf5.
|
||||
- **Writing without copies:** `FileWriter` no longer copies chunks into a
|
||||
buffer per pass and then into the file's buffer. A chunk that is a
|
||||
contiguous run of the dataset's data (a dataset stored as one chunk of
|
||||
its shape, blocks of whole rows) is borrowed when unfiltered and
|
||||
compressed straight from it when filtered, and contiguous datasets are
|
||||
not copied. New `FileWriter::finish_with(put)` hands the file out piece
|
||||
by piece; `FileBuilder::write` uses it (through a 1 MiB `BufWriter`,
|
||||
still to a temporary file renamed into place), so a file is never
|
||||
assembled in memory. New `DatasetBuilder::with_u8_data_owned(Vec<u8>)`
|
||||
takes the data without the copy `with_u8_data` makes. Peak resident
|
||||
memory of one unfiltered 1 GiB chunk of `u8`
|
||||
(`cargo run --release -p clawhdf5 --example write_one_chunk --
|
||||
1073741824 OUT owned|slice`, tank, 2026-09-29): 5.0 GiB before (the
|
||||
writer of `4260af4`), 1.0 GiB after with `owned`; 6.0 and 2.0 GiB with
|
||||
`slice`; 4 GiB + 7 bytes `owned`, 4.00 GiB. Same bytes written. Write
|
||||
speed was not re-measured (the machine was not idle): the A/B of
|
||||
`FileBuilder::write` and `finish` on small and large files is still to
|
||||
be run.
|
||||
- `chunked_write::extract_chunk` (behind `split_into_chunks`) no longer
|
||||
panics when the data is shorter than the shape; the missing part is
|
||||
zeros, as it was meant to be.
|
||||
- Tests: `crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5` (31 KB,
|
||||
libhdf5 2.0.0 via h5py 3.16, `gen_huge_chunks.py dims`: `u8` chunks of
|
||||
2^32 + 7 deflated twice, Single Chunk and Extensible Array), and in
|
||||
`huge_chunks_interop`: `huge_chunk_dims_list` (always), and with
|
||||
`CLAWHDF5_HUGE_CHUNKS=1` `huge_chunk_dims_read`,
|
||||
`huge_chunk_dims_unfiltered_read` (a sparse file h5py writes, Single
|
||||
Chunk and Fixed Array, mmap and positioned reads),
|
||||
`writer_huge_chunk_dims_round_trip` (deflated, read back by clawhdf5,
|
||||
h5py 3.16 and — layouts — h5dump 2.2.0),
|
||||
`writer_unfiltered_huge_chunk_round_trip` (an unfiltered 4 GiB + 7 byte
|
||||
chunk: clawhdf5, h5py, and h5dump 2.2.0 printing values past 2^32) and
|
||||
`writer_huge_chunks_lz4_zstd_round_trip` (with the `lz4`/`zstd`
|
||||
features; h5py through hdf5plugin 7.1.0). The whole opt-in suite
|
||||
(`cargo test --release -p clawhdf5 --features lz4,zstd --test
|
||||
huge_chunks_interop -- --test-threads=1`, h5py and h5dump 2.2.0
|
||||
included) passed with a peak of 10.04 GiB resident (tank, 2026-09-29,
|
||||
commit `0065c6b`). Unit tests: 5-byte dimension encoding, contiguous
|
||||
chunk detection, `finish_with` against `finish`. The wasm package test
|
||||
lists the new fixture and gets `FormatError::Overflow` reading it.
|
||||
`docs/known-issues.md`: two fixed entries; the open "Chunks of 4 GiB or
|
||||
more: limits" keeps the refused filters, the editor and memory.
|
||||
|
||||
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
|
||||
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
|
||||
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
|
||||
|
||||
Reference in New Issue
Block a user