docs: chunks of 4 GiB or more — CHANGELOG, known-issues, README
Three fixed entries (reading, writing, LZ4 over 256 MiB) and the open limits entry; the selection-read entry loses its fill-value case; the FileEditor limits gain the refused rewrites; the README tables name what is supported and what is not. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
+138
-5
@@ -15,6 +15,10 @@ Checked against `main` at `9b5803f` on 2026-09-28.
|
||||
|
||||
| Issue | Kind | Since |
|
||||
|---|---|---|
|
||||
|
||||
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
|
||||
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
|
||||
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
|
||||
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
|
||||
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
|
||||
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
|
||||
@@ -57,6 +61,10 @@ writing anything:
|
||||
message (they have no dense storage);
|
||||
- partial edge chunks stored unfiltered (`H5Pset_chunk_opts`), external
|
||||
raw data files, virtual datasets;
|
||||
- values of a dataset whose chunks are 4 GiB or more (HDF5 2.0), and a
|
||||
resize of one that prunes, fills or allocates chunks (added 2026-09-28;
|
||||
growing its extent under late allocation, and its attributes, work):
|
||||
see [Chunks of 4 GiB or more](#chunks-of-4-gib-or-more-limits);
|
||||
- files with a metadata cache image, paged or persistent free-space
|
||||
management, a driver info block, or version-3 consistency flags set.
|
||||
|
||||
@@ -141,18 +149,79 @@ before anything is written. On top of them:
|
||||
read (and so the Python `ds[...]`) of contiguous data copies just the
|
||||
selected runs; of chunked data it materialises only the selection's
|
||||
bounding box when that box covers at most half the dataset
|
||||
(`partial_read`). It decodes the whole dataset and extracts the selection
|
||||
instead when:
|
||||
(`partial_read`), with the dataset's fill value where no chunk was
|
||||
written. It decodes the whole dataset and extracts the selection instead
|
||||
when:
|
||||
- the bounding box covers more than half the dataset — which a strided
|
||||
selection across a chunked dataset (`ds[::100]`) always does, although
|
||||
it may touch few chunks;
|
||||
- the dataset is compact or virtual, or has no storage;
|
||||
- it is chunked with a non-default fill value (the box path does not fill
|
||||
unallocated chunks, so the fill-aware full read is used).
|
||||
- the dataset is compact or virtual, or has no storage.
|
||||
|
||||
A chunked dataset with a non-default fill value was read whole for any
|
||||
selection until 2026-09-28 (`partial_read::read_selection_filled_in` fills
|
||||
the box now; `partial_read_equivalence::fill_value_selections_match_full_reads`).
|
||||
|
||||
Values are correct in every case; this is cost only. Selections other than
|
||||
`Selection::All` also bypass the file's chunk cache.
|
||||
|
||||
## Chunks of 4 GiB or more: limits
|
||||
|
||||
**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more
|
||||
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
|
||||
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
|
||||
also makes a filtered chunk index element store the chunk's size in "size
|
||||
of lengths" bytes (8) instead of one byte more than the chunk needs.
|
||||
clawhdf5 reads and writes such chunks in every index libhdf5 gives them
|
||||
(Single Chunk, Implicit, Fixed Array, Extensible Array, v2 B-tree; libhdf5
|
||||
never puts one in a v1 B-tree and refuses to open one, as clawhdf5 does).
|
||||
The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by
|
||||
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
|
||||
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
|
||||
What remains:
|
||||
- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under
|
||||
`H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`)
|
||||
and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec<u32>`.
|
||||
A 4 GiB chunk of 1-byte elements needs one.
|
||||
- **Filters the writer refuses for such a chunk** (`FilterError`, before
|
||||
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
|
||||
filters record sizes or block lengths in 32 bits, or cannot take such a
|
||||
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
|
||||
written; only shuffle + deflate is tested end to end (LZ4's framing and
|
||||
the 256 MiB limit it had are unit-tested; Zstd is not tested at this
|
||||
size).
|
||||
- **Unfiltered chunks this size are written but not tested end to end:**
|
||||
`FileBuilder` assembles the whole file in memory, several copies of the
|
||||
chunk (well over the 12 GiB the tests may use). Their layout messages
|
||||
are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`).
|
||||
- **`FileEditor`** refuses to rewrite such chunks (see
|
||||
[its limits](#in-place-modification-fileeditor-limits)).
|
||||
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
|
||||
twice when it is shuffled (the inflated and the unshuffled copy, as
|
||||
libhdf5's filters do too); writing one holds the chunk and its shuffled
|
||||
copy. A selection decodes only the chunks it touches; one of an
|
||||
unfiltered chunk reads only the rows it selects, through a memory map or
|
||||
positioned reads (`File::open_storage`): every selection of
|
||||
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
|
||||
memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop
|
||||
-- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py
|
||||
included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection
|
||||
reads of the double-deflated fixture alone, 4.0 GiB.
|
||||
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
|
||||
reading or writing one is `FormatError::Overflow` ("exceeds the
|
||||
addressable size" / "address space"); the file still opens and lists
|
||||
(`examples/wasm-viewer/test/test.mjs`, "huge chunk").
|
||||
- **libhdf5 2.2.0's h5dump cannot print these values:** built here from tag
|
||||
`2.2.0`, it fails to inflate any deflated chunk over 4 GiB, those
|
||||
libhdf5 2.0.0 writes included ("memory allocation failed for deflate
|
||||
uncompression", `H5Zdeflate.c`), while h5py 3.16 (libhdf5 2.0.0) reads
|
||||
them. The tests check h5dump 2.2.0 on layouts and storage sizes only
|
||||
(`CLAWHDF5_H5DUMP2`).
|
||||
- libhdf5 2.0.0 itself (h5py 3.16) drops a write to an unallocated,
|
||||
unfiltered Single Chunk this large when the fill time is "never" (the
|
||||
dataset stays unallocated); `fixtures/gen_huge_chunks.py` allocates it
|
||||
early instead. Not a clawhdf5 issue; recorded because the generator
|
||||
depends on it.
|
||||
|
||||
## HDF5 features still unsupported
|
||||
|
||||
**Status:** open. What remains of the gaps the
|
||||
@@ -564,6 +633,70 @@ under load from other builds) a freshly opened file with 8192 chunks
|
||||
per dataset reads small selections 1.2x to 2.3x slower through a version-1
|
||||
B-tree (see `BENCHMARKS.md`, "HDF5 1.8 format").
|
||||
|
||||
## Chunks of 4 GiB or more could not be read
|
||||
|
||||
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`); affected every
|
||||
release (v2.1.0 to v2.7.0). Errors, and cost; no wrong values were
|
||||
returned. Nothing for users to do but upgrade.
|
||||
|
||||
HDF5 2.0 writes chunks of more than 4 GiB - 1 bytes (layout message
|
||||
version 5). `ChunkInfo::chunk_size` was a `u32`, so the Single Chunk,
|
||||
Implicit, Fixed Array and Extensible Array readers truncated an
|
||||
unfiltered chunk's size and the read failed ("incorrect chunk size
|
||||
returned from index for unfiltered chunk"); a v2 B-tree index refused any
|
||||
such chunk ("chunk larger than 4 GiB"), filtered or not. Filtered chunks
|
||||
in the other indexes read, because their stored sizes were small. A
|
||||
selection of a chunked dataset with a non-default fill value decoded the
|
||||
whole dataset (8 GiB of output for the fixture's 2-D dataset), and a
|
||||
selection of an unfiltered chunk fetched the whole chunk from a file that
|
||||
is not in memory. `ChunkInfo::chunk_size` and `ChunkMapping::file_size` are now
|
||||
`u64`; the selection path fills its box with the fill value and reads an
|
||||
unfiltered chunk's rows only. Tests: `huge_chunks_interop`
|
||||
(`filtered_huge_chunk_indexes_list` always; with `CLAWHDF5_HUGE_CHUNKS=1`
|
||||
`filtered_huge_chunks_read`, over a fixture libhdf5 2.0.0 wrote, and
|
||||
`unfiltered_huge_chunks_read`, over a 44 GiB sparse file h5py writes at
|
||||
test time). Limits that remain:
|
||||
[Chunks of 4 GiB or more](#chunks-of-4-gib-or-more-limits).
|
||||
|
||||
## Chunks of 4 GiB or more were written unreadable
|
||||
|
||||
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`). The deflate
|
||||
truncation was before any release (the one-pass deflate dates from
|
||||
2026-09-23); the rest affected every release (v2.1.0 to v2.7.0). Files
|
||||
clawhdf5 wrote with a chunk of 4 GiB or more should be written again.
|
||||
|
||||
The writer gave such a chunk layout message version 4 (which libhdf5
|
||||
before 2.0 cannot read, and which libhdf5 2.x never writes for it; whether
|
||||
2.x reads what clawhdf5 wrote was not checked), cut a chunk dimension of
|
||||
2^32 or more to 32 bits, and, since 2026-09-23, deflated only the first
|
||||
4 GiB - 1 bytes of the chunk (zlib takes at most that much per call and
|
||||
`Finish` ended the stream there): the chunk failed to decode ("decoded to
|
||||
4294967295 bytes, expected 4294967304"). It now writes layout version 5
|
||||
with libhdf5's index element widths (the Fixed and Extensible Array
|
||||
structures match libhdf5 2.0.0's byte for byte:
|
||||
`huge_chunk_array_indexes_match_libhdf5`), refuses chunk dimensions of
|
||||
2^32 or more and the filters that cannot take such a chunk, deflates the
|
||||
whole chunk, and no longer keeps the compressor's worst-case bound (4 GiB
|
||||
of zeroed memory for a chunk that deflates to 4 MiB): writing the four
|
||||
datasets of `writer_huge_chunks_round_trip` went past the tests' 12 GiB
|
||||
cap, and compressing one shuffled 4 GiB chunk now peaks at 4.3 GiB
|
||||
(tank, 2026-09-28). h5py 3.16 reads what it writes.
|
||||
|
||||
## LZ4 chunks larger than 256 MiB were refused
|
||||
|
||||
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`); affected every
|
||||
release (v2.1.0 to v2.7.0). An error, never wrong data. Nothing for users
|
||||
to do but upgrade.
|
||||
|
||||
The LZ4 decoder refused a chunk that decodes to more than 256 MiB
|
||||
("lz4: declared size exceeds limit") even when the dataset's chunk size
|
||||
bounded it; that ceiling is meant for a decode whose size is unknown, and
|
||||
now applies only then (deflate never had it). A chunk of 4 GiB or more is
|
||||
always read as the registered HDF5 framing, whose 64-bit size then no
|
||||
longer starts with four zero bytes. Tests:
|
||||
`filters::tests::lz4_chunks_over_256_mib_decode`,
|
||||
`lz4_chunks_of_4_gib_use_the_registered_framing`.
|
||||
|
||||
## A dropped `FileEditor` could keep its file locked for a moment
|
||||
|
||||
**Status:** fixed 2026-09-28 (#23), before any release
|
||||
|
||||
Reference in New Issue
Block a user