docs: chunks of 4 GiB or more — CHANGELOG, known-issues, README

Three fixed entries (reading, writing, LZ4 over 256 MiB) and the open
limits entry; the selection-read entry loses its fill-value case; the
FileEditor limits gain the refused rewrites; the README tables name what
is supported and what is not.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 23:44:55 -05:00
co-authored by Claude Opus 5.5
parent 7a2e61ebd1
commit 740d1124a8
4 changed files with 217 additions and 9 deletions
+74
View File
@@ -111,6 +111,80 @@
the pre-1.8 format (version-0 superblock, symbol-table groups) cannot be
written. `docs/known-issues.md`.
### Chunks of 4 GiB or more (HDF5 2.0's layout message version 5) (2026-09-28)
- **What libhdf5 does** (read in libhdf5 2.2.0's `H5Dchunk.c`, `H5Dfarray.c`,
`H5Dearray.c`, `H5Dbtree2.c`, `H5Dbtree.c`, `H5Olayout.c`): a chunk of
more than 0xFFFFFFFF bytes requires layout message version 5, so a high
bound of `H5F_LIBVER_V200` ("chunk size > 4GB requires H5F_LIBVER_V200"),
whatever the low bound, and version 5 always uses the newer indexes.
Under version 5 a filtered Fixed Array, Extensible Array or v2 B-tree
element stores the chunk's size in "size of lengths" bytes (8); under
version 4 in one byte more than the chunk needs. A filtered Single Chunk
stores it in "size of lengths" bytes under both. Unfiltered elements
store no size, and the Implicit index none at all. A v1 B-tree key keeps
a 32-bit size: libhdf5 never writes a larger chunk there and refuses to
open one ("chunk size must be < 4GB with v1 b-tree index").
- **Reading** such chunks works in every index: `ChunkInfo::chunk_size` and
`ChunkMapping::file_size` are now `u64` (**public API change**). The
Single Chunk, Implicit, Fixed and Extensible Array readers truncated an
unfiltered chunk's size (the read failed), and the v2 B-tree reader
refused any chunk over 4 GiB. Checked against a fixture libhdf5 2.0.0
wrote (`tests/fixtures/huge_chunks_filtered.h5`, 188 KB: a 4 GiB + 8
byte chunk in each filtered index, deflated twice) and a 44 GiB sparse
file h5py writes at test time (every unfiltered index).
- **Selections read what they touch:** a selection of a chunked dataset with
a non-default fill value is read over a box of fill values from the
chunks it touches (new `partial_read::read_selection_filled_in`); it
used to decode the whole dataset. An unfiltered chunk of a file that is
not in memory (`File::open_storage`) is read row by row instead of whole.
A filtered chunk is still decoded whole.
- **Writing:** a chunk of more than `u32::MAX` bytes gets layout message
version 5 and libhdf5's index element widths; the Fixed and Extensible
Array structures match libhdf5 2.0.0's byte for byte. New public helpers
`chunked_write::{layout_version_for, chunk_size_len, MAX_V4_CHUNK_BYTES}`.
Refused before anything is written: chunk dimensions of 2^32 or more
(they were cut to 32 bits) and, for such chunks, LZF, bitshuffle, bzip2,
Blosc and pcodec. Chunks are extracted row by row and one at a time
(compressed in parallel only up to 64 MiB), and a filter pipeline no
longer copies its input first.
- **Deflate** compressed only the first 4 GiB - 1 bytes of a larger chunk
(zlib takes at most that much per call, and `Finish` ended the stream
there; since the one-pass deflate of 2026-09-23, in `clawhdf5-format` and
`clawhdf5-filters`), and it reserved the compressor's worst-case bound,
which flate2's Rust backends zero on every call: 4 GiB kept per 4 GiB
chunk. Inputs over 64 MiB now start at 1/16 of the bound and grow, and a
result more than 64 MiB too large is shrunk. On the read side, an
intermediate deflate stage (one followed by a filter other than shuffle
or Fletcher-32) no longer reserves the chunk's whole bound: reading the
double-deflated fixture peaked at 8.5 GiB, now 4.0 GiB.
- **LZ4:** a chunk decoding to more than 256 MiB was refused ("lz4:
declared size exceeds limit") though the chunk size bounded it; the
ceiling now applies only to a decode of unknown size. A chunk of 4 GiB
or more is read as the registered HDF5 framing, whose 64-bit size no
longer starts with four zero bytes.
- **`FileEditor`** refuses, before writing anything, to write values into
or to prune, fill or allocate chunks of 4 GiB or more; growing the extent
(late allocation) and setting attributes work. It sizes a version-5
filtered index element from the file's size of lengths (it assumed 8).
- **32-bit targets:** reading or writing such a chunk is
`FormatError::Overflow`; `scripts/check-32bit-casts.sh` passes, and the
wasm package test reads the fixture and gets that error in every index.
- Tests: `crates/clawhdf5/tests/huge_chunks_interop.rs` (index listing,
byte-for-byte array indexes and the editor always; with
`CLAWHDF5_HUGE_CHUNKS=1`, reads of every index filtered and unfiltered,
and a write of every index read back by clawhdf5, h5py 3.16 and — layout
and storage size — h5dump 2.2.0 via `CLAWHDF5_H5DUMP2`),
`partial_read_equivalence::fill_value_selections_match_full_reads`, unit
tests of the encodings (`chunked_write::tests`) and of LZ4. The opt-in
run (`cargo test --release -p clawhdf5 --test huge_chunks_interop --
--test-threads=1`) peaked at 8.06 GiB resident, h5py included (tank,
2026-09-28, commit `9143737`). libhdf5 2.2.0's h5dump cannot print these
chunks' values (its deflate filter fails on any chunk over 4 GiB,
libhdf5's own included); h5py 3.16 reads them. `docs/known-issues.md`: three fixed
entries and the open "Chunks of 4 GiB or more: limits" (chunk dimensions
of 2^32 or more, the refused filters, memory, the untested unfiltered
write).
### A dropped `FileEditor` releases its lock at once (2026-09-28)
- `FileEditor`'s `flock` could outlive the editor for a moment when
another thread forked to spawn a process: the child shared the locked