docs: chunks of 4 GiB or more — CHANGELOG, known-issues, README

Three fixed entries (reading, writing, LZ4 over 256 MiB) and the open
limits entry; the selection-read entry loses its fill-value case; the
FileEditor limits gain the refused rewrites; the README tables name what
is supported and what is not.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 23:44:55 -05:00
co-authored by Claude Opus 5.5
parent 7a2e61ebd1
commit 740d1124a8
4 changed files with 217 additions and 9 deletions
+138 -5
View File
@@ -15,6 +15,10 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since |
|---|---|---|
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
@@ -57,6 +61,10 @@ writing anything:
message (they have no dense storage);
- partial edge chunks stored unfiltered (`H5Pset_chunk_opts`), external
raw data files, virtual datasets;
- values of a dataset whose chunks are 4 GiB or more (HDF5 2.0), and a
resize of one that prunes, fills or allocates chunks (added 2026-09-28;
growing its extent under late allocation, and its attributes, work):
see [Chunks of 4 GiB or more](#chunks-of-4-gib-or-more-limits);
- files with a metadata cache image, paged or persistent free-space
management, a driver info block, or version-3 consistency flags set.
@@ -141,18 +149,79 @@ before anything is written. On top of them:
read (and so the Python `ds[...]`) of contiguous data copies just the
selected runs; of chunked data it materialises only the selection's
bounding box when that box covers at most half the dataset
(`partial_read`). It decodes the whole dataset and extracts the selection
instead when:
(`partial_read`), with the dataset's fill value where no chunk was
written. It decodes the whole dataset and extracts the selection instead
when:
- the bounding box covers more than half the dataset — which a strided
selection across a chunked dataset (`ds[::100]`) always does, although
it may touch few chunks;
- the dataset is compact or virtual, or has no storage;
- it is chunked with a non-default fill value (the box path does not fill
unallocated chunks, so the fill-aware full read is used).
- the dataset is compact or virtual, or has no storage.
A chunked dataset with a non-default fill value was read whole for any
selection until 2026-09-28 (`partial_read::read_selection_filled_in` fills
the box now; `partial_read_equivalence::fill_value_selections_match_full_reads`).
Values are correct in every case; this is cost only. Selections other than
`Selection::All` also bypass the file's chunk cache.
## Chunks of 4 GiB or more: limits
**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
also makes a filtered chunk index element store the chunk's size in "size
of lengths" bytes (8) instead of one byte more than the chunk needs.
clawhdf5 reads and writes such chunks in every index libhdf5 gives them
(Single Chunk, Implicit, Fixed Array, Extensible Array, v2 B-tree; libhdf5
never puts one in a v1 B-tree and refuses to open one, as clawhdf5 does).
The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
What remains:
- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under
`H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`)
and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec<u32>`.
A 4 GiB chunk of 1-byte elements needs one.
- **Filters the writer refuses for such a chunk** (`FilterError`, before
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
filters record sizes or block lengths in 32 bits, or cannot take such a
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
written; only shuffle + deflate is tested end to end (LZ4's framing and
the 256 MiB limit it had are unit-tested; Zstd is not tested at this
size).
- **Unfiltered chunks this size are written but not tested end to end:**
`FileBuilder` assembles the whole file in memory, several copies of the
chunk (well over the 12 GiB the tests may use). Their layout messages
are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`).
- **`FileEditor`** refuses to rewrite such chunks (see
[its limits](#in-place-modification-fileeditor-limits)).
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
twice when it is shuffled (the inflated and the unshuffled copy, as
libhdf5's filters do too); writing one holds the chunk and its shuffled
copy. A selection decodes only the chunks it touches; one of an
unfiltered chunk reads only the rows it selects, through a memory map or
positioned reads (`File::open_storage`): every selection of
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop
-- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py
included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection
reads of the double-deflated fixture alone, 4.0 GiB.
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
reading or writing one is `FormatError::Overflow` ("exceeds the
addressable size" / "address space"); the file still opens and lists
(`examples/wasm-viewer/test/test.mjs`, "huge chunk").
- **libhdf5 2.2.0's h5dump cannot print these values:** built here from tag
`2.2.0`, it fails to inflate any deflated chunk over 4 GiB, those
libhdf5 2.0.0 writes included ("memory allocation failed for deflate
uncompression", `H5Zdeflate.c`), while h5py 3.16 (libhdf5 2.0.0) reads
them. The tests check h5dump 2.2.0 on layouts and storage sizes only
(`CLAWHDF5_H5DUMP2`).
- libhdf5 2.0.0 itself (h5py 3.16) drops a write to an unallocated,
unfiltered Single Chunk this large when the fill time is "never" (the
dataset stays unallocated); `fixtures/gen_huge_chunks.py` allocates it
early instead. Not a clawhdf5 issue; recorded because the generator
depends on it.
## HDF5 features still unsupported
**Status:** open. What remains of the gaps the
@@ -564,6 +633,70 @@ under load from other builds) a freshly opened file with 8192 chunks
per dataset reads small selections 1.2x to 2.3x slower through a version-1
B-tree (see `BENCHMARKS.md`, "HDF5 1.8 format").
## Chunks of 4 GiB or more could not be read
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`); affected every
release (v2.1.0 to v2.7.0). Errors, and cost; no wrong values were
returned. Nothing for users to do but upgrade.
HDF5 2.0 writes chunks of more than 4 GiB - 1 bytes (layout message
version 5). `ChunkInfo::chunk_size` was a `u32`, so the Single Chunk,
Implicit, Fixed Array and Extensible Array readers truncated an
unfiltered chunk's size and the read failed ("incorrect chunk size
returned from index for unfiltered chunk"); a v2 B-tree index refused any
such chunk ("chunk larger than 4 GiB"), filtered or not. Filtered chunks
in the other indexes read, because their stored sizes were small. A
selection of a chunked dataset with a non-default fill value decoded the
whole dataset (8 GiB of output for the fixture's 2-D dataset), and a
selection of an unfiltered chunk fetched the whole chunk from a file that
is not in memory. `ChunkInfo::chunk_size` and `ChunkMapping::file_size` are now
`u64`; the selection path fills its box with the fill value and reads an
unfiltered chunk's rows only. Tests: `huge_chunks_interop`
(`filtered_huge_chunk_indexes_list` always; with `CLAWHDF5_HUGE_CHUNKS=1`
`filtered_huge_chunks_read`, over a fixture libhdf5 2.0.0 wrote, and
`unfiltered_huge_chunks_read`, over a 44 GiB sparse file h5py writes at
test time). Limits that remain:
[Chunks of 4 GiB or more](#chunks-of-4-gib-or-more-limits).
## Chunks of 4 GiB or more were written unreadable
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`). The deflate
truncation was before any release (the one-pass deflate dates from
2026-09-23); the rest affected every release (v2.1.0 to v2.7.0). Files
clawhdf5 wrote with a chunk of 4 GiB or more should be written again.
The writer gave such a chunk layout message version 4 (which libhdf5
before 2.0 cannot read, and which libhdf5 2.x never writes for it; whether
2.x reads what clawhdf5 wrote was not checked), cut a chunk dimension of
2^32 or more to 32 bits, and, since 2026-09-23, deflated only the first
4 GiB - 1 bytes of the chunk (zlib takes at most that much per call and
`Finish` ended the stream there): the chunk failed to decode ("decoded to
4294967295 bytes, expected 4294967304"). It now writes layout version 5
with libhdf5's index element widths (the Fixed and Extensible Array
structures match libhdf5 2.0.0's byte for byte:
`huge_chunk_array_indexes_match_libhdf5`), refuses chunk dimensions of
2^32 or more and the filters that cannot take such a chunk, deflates the
whole chunk, and no longer keeps the compressor's worst-case bound (4 GiB
of zeroed memory for a chunk that deflates to 4 MiB): writing the four
datasets of `writer_huge_chunks_round_trip` went past the tests' 12 GiB
cap, and compressing one shuffled 4 GiB chunk now peaks at 4.3 GiB
(tank, 2026-09-28). h5py 3.16 reads what it writes.
## LZ4 chunks larger than 256 MiB were refused
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`); affected every
release (v2.1.0 to v2.7.0). An error, never wrong data. Nothing for users
to do but upgrade.
The LZ4 decoder refused a chunk that decodes to more than 256 MiB
("lz4: declared size exceeds limit") even when the dataset's chunk size
bounded it; that ceiling is meant for a decode whose size is unknown, and
now applies only then (deflate never had it). A chunk of 4 GiB or more is
always read as the registered HDF5 framing, whose 64-bit size then no
longer starts with four zero bytes. Tests:
`filters::tests::lz4_chunks_over_256_mib_decode`,
`lz4_chunks_of_4_gib_use_the_registered_framing`.
## A dropped `FileEditor` could keep its file locked for a moment
**Status:** fixed 2026-09-28 (#23), before any release