Chunks of 4 GiB or more: read in every index, write as libhdf5 2.x does #29

Merged
osobh merged 7 commits from feat/huge-chunks into main 2026-09-29 11:57:48 +00:00
4 changed files with 217 additions and 9 deletions
Showing only changes of commit 740d1124a8 - Show all commits
+74
View File
@@ -111,6 +111,80 @@
the pre-1.8 format (version-0 superblock, symbol-table groups) cannot be the pre-1.8 format (version-0 superblock, symbol-table groups) cannot be
written. `docs/known-issues.md`. written. `docs/known-issues.md`.
### Chunks of 4 GiB or more (HDF5 2.0's layout message version 5) (2026-09-28)
- **What libhdf5 does** (read in libhdf5 2.2.0's `H5Dchunk.c`, `H5Dfarray.c`,
`H5Dearray.c`, `H5Dbtree2.c`, `H5Dbtree.c`, `H5Olayout.c`): a chunk of
more than 0xFFFFFFFF bytes requires layout message version 5, so a high
bound of `H5F_LIBVER_V200` ("chunk size > 4GB requires H5F_LIBVER_V200"),
whatever the low bound, and version 5 always uses the newer indexes.
Under version 5 a filtered Fixed Array, Extensible Array or v2 B-tree
element stores the chunk's size in "size of lengths" bytes (8); under
version 4 in one byte more than the chunk needs. A filtered Single Chunk
stores it in "size of lengths" bytes under both. Unfiltered elements
store no size, and the Implicit index none at all. A v1 B-tree key keeps
a 32-bit size: libhdf5 never writes a larger chunk there and refuses to
open one ("chunk size must be < 4GB with v1 b-tree index").
- **Reading** such chunks works in every index: `ChunkInfo::chunk_size` and
`ChunkMapping::file_size` are now `u64` (**public API change**). The
Single Chunk, Implicit, Fixed and Extensible Array readers truncated an
unfiltered chunk's size (the read failed), and the v2 B-tree reader
refused any chunk over 4 GiB. Checked against a fixture libhdf5 2.0.0
wrote (`tests/fixtures/huge_chunks_filtered.h5`, 188 KB: a 4 GiB + 8
byte chunk in each filtered index, deflated twice) and a 44 GiB sparse
file h5py writes at test time (every unfiltered index).
- **Selections read what they touch:** a selection of a chunked dataset with
a non-default fill value is read over a box of fill values from the
chunks it touches (new `partial_read::read_selection_filled_in`); it
used to decode the whole dataset. An unfiltered chunk of a file that is
not in memory (`File::open_storage`) is read row by row instead of whole.
A filtered chunk is still decoded whole.
- **Writing:** a chunk of more than `u32::MAX` bytes gets layout message
version 5 and libhdf5's index element widths; the Fixed and Extensible
Array structures match libhdf5 2.0.0's byte for byte. New public helpers
`chunked_write::{layout_version_for, chunk_size_len, MAX_V4_CHUNK_BYTES}`.
Refused before anything is written: chunk dimensions of 2^32 or more
(they were cut to 32 bits) and, for such chunks, LZF, bitshuffle, bzip2,
Blosc and pcodec. Chunks are extracted row by row and one at a time
(compressed in parallel only up to 64 MiB), and a filter pipeline no
longer copies its input first.
- **Deflate** compressed only the first 4 GiB - 1 bytes of a larger chunk
(zlib takes at most that much per call, and `Finish` ended the stream
there; since the one-pass deflate of 2026-09-23, in `clawhdf5-format` and
`clawhdf5-filters`), and it reserved the compressor's worst-case bound,
which flate2's Rust backends zero on every call: 4 GiB kept per 4 GiB
chunk. Inputs over 64 MiB now start at 1/16 of the bound and grow, and a
result more than 64 MiB too large is shrunk. On the read side, an
intermediate deflate stage (one followed by a filter other than shuffle
or Fletcher-32) no longer reserves the chunk's whole bound: reading the
double-deflated fixture peaked at 8.5 GiB, now 4.0 GiB.
- **LZ4:** a chunk decoding to more than 256 MiB was refused ("lz4:
declared size exceeds limit") though the chunk size bounded it; the
ceiling now applies only to a decode of unknown size. A chunk of 4 GiB
or more is read as the registered HDF5 framing, whose 64-bit size no
longer starts with four zero bytes.
- **`FileEditor`** refuses, before writing anything, to write values into
or to prune, fill or allocate chunks of 4 GiB or more; growing the extent
(late allocation) and setting attributes work. It sizes a version-5
filtered index element from the file's size of lengths (it assumed 8).
- **32-bit targets:** reading or writing such a chunk is
`FormatError::Overflow`; `scripts/check-32bit-casts.sh` passes, and the
wasm package test reads the fixture and gets that error in every index.
- Tests: `crates/clawhdf5/tests/huge_chunks_interop.rs` (index listing,
byte-for-byte array indexes and the editor always; with
`CLAWHDF5_HUGE_CHUNKS=1`, reads of every index filtered and unfiltered,
and a write of every index read back by clawhdf5, h5py 3.16 and — layout
and storage size — h5dump 2.2.0 via `CLAWHDF5_H5DUMP2`),
`partial_read_equivalence::fill_value_selections_match_full_reads`, unit
tests of the encodings (`chunked_write::tests`) and of LZ4. The opt-in
run (`cargo test --release -p clawhdf5 --test huge_chunks_interop --
--test-threads=1`) peaked at 8.06 GiB resident, h5py included (tank,
2026-09-28, commit `9143737`). libhdf5 2.2.0's h5dump cannot print these
chunks' values (its deflate filter fails on any chunk over 4 GiB,
libhdf5's own included); h5py 3.16 reads them. `docs/known-issues.md`: three fixed
entries and the open "Chunks of 4 GiB or more: limits" (chunk dimensions
of 2^32 or more, the refused filters, memory, the untested unfiltered
write).
### A dropped `FileEditor` releases its lock at once (2026-09-28) ### A dropped `FileEditor` releases its lock at once (2026-09-28)
- `FileEditor`'s `flock` could outlive the editor for a moment when - `FileEditor`'s `flock` could outlive the editor for a moment when
another thread forked to spawn a process: the child shared the locked another thread forked to spawn a process: the child shared the locked
+2 -2
View File
@@ -116,9 +116,9 @@ Limits and open issues, with dates, are in
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) | | **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | | **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 | | **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error) | | **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | | **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; filters this build cannot encode (refused before any write) | | **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | | **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) | | **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
+3 -2
View File
@@ -32,10 +32,11 @@ chunk still inflates to 4 GiB + 8 bytes.
writes only the elements written (an unfiltered chunk larger than the chunk writes only the elements written (an unfiltered chunk larger than the chunk
cache is written in place) and the file is sparse: tens of GiB long but a few cache is written in place) and the file is sparse: tens of GiB long but a few
blocks on disk. Unwritten elements read as whatever the file holds there: blocks on disk. Unwritten elements read as whatever the file holds there:
zeros. Do not put it on tmpfs, which is memory. zeros. Do not put it on tmpfs, which is memory. Its Single Chunk is
allocated early (see `make`).
libhdf5 holds a whole filtered chunk in memory while it writes it, so the libhdf5 holds a whole filtered chunk in memory while it writes it, so the
`filtered` run needs about 4.5 GiB of memory (and about 3 minutes). `filtered` run needs about 4 GiB of memory and 100 s (tank).
Generated 2026-09-28 with h5py 3.16.0 (libhdf5 2.0.0). Generated 2026-09-28 with h5py 3.16.0 (libhdf5 2.0.0).
""" """
+138 -5
View File
@@ -15,6 +15,10 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since | | Issue | Kind | Since |
|---|---|---| |---|---|---|
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 | | [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | | [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | | [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
@@ -57,6 +61,10 @@ writing anything:
message (they have no dense storage); message (they have no dense storage);
- partial edge chunks stored unfiltered (`H5Pset_chunk_opts`), external - partial edge chunks stored unfiltered (`H5Pset_chunk_opts`), external
raw data files, virtual datasets; raw data files, virtual datasets;
- values of a dataset whose chunks are 4 GiB or more (HDF5 2.0), and a
resize of one that prunes, fills or allocates chunks (added 2026-09-28;
growing its extent under late allocation, and its attributes, work):
see [Chunks of 4 GiB or more](#chunks-of-4-gib-or-more-limits);
- files with a metadata cache image, paged or persistent free-space - files with a metadata cache image, paged or persistent free-space
management, a driver info block, or version-3 consistency flags set. management, a driver info block, or version-3 consistency flags set.
@@ -141,18 +149,79 @@ before anything is written. On top of them:
read (and so the Python `ds[...]`) of contiguous data copies just the read (and so the Python `ds[...]`) of contiguous data copies just the
selected runs; of chunked data it materialises only the selection's selected runs; of chunked data it materialises only the selection's
bounding box when that box covers at most half the dataset bounding box when that box covers at most half the dataset
(`partial_read`). It decodes the whole dataset and extracts the selection (`partial_read`), with the dataset's fill value where no chunk was
instead when: written. It decodes the whole dataset and extracts the selection instead
when:
- the bounding box covers more than half the dataset — which a strided - the bounding box covers more than half the dataset — which a strided
selection across a chunked dataset (`ds[::100]`) always does, although selection across a chunked dataset (`ds[::100]`) always does, although
it may touch few chunks; it may touch few chunks;
- the dataset is compact or virtual, or has no storage; - the dataset is compact or virtual, or has no storage.
- it is chunked with a non-default fill value (the box path does not fill
unallocated chunks, so the fill-aware full read is used). A chunked dataset with a non-default fill value was read whole for any
selection until 2026-09-28 (`partial_read::read_selection_filled_in` fills
the box now; `partial_read_equivalence::fill_value_selections_match_full_reads`).
Values are correct in every case; this is cost only. Selections other than Values are correct in every case; this is cost only. Selections other than
`Selection::All` also bypass the file's chunk cache. `Selection::All` also bypass the file's chunk cache.
## Chunks of 4 GiB or more: limits
**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
also makes a filtered chunk index element store the chunk's size in "size
of lengths" bytes (8) instead of one byte more than the chunk needs.
clawhdf5 reads and writes such chunks in every index libhdf5 gives them
(Single Chunk, Implicit, Fixed Array, Extensible Array, v2 B-tree; libhdf5
never puts one in a v1 B-tree and refuses to open one, as clawhdf5 does).
The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
What remains:
- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under
`H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`)
and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec<u32>`.
A 4 GiB chunk of 1-byte elements needs one.
- **Filters the writer refuses for such a chunk** (`FilterError`, before
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
filters record sizes or block lengths in 32 bits, or cannot take such a
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
written; only shuffle + deflate is tested end to end (LZ4's framing and
the 256 MiB limit it had are unit-tested; Zstd is not tested at this
size).
- **Unfiltered chunks this size are written but not tested end to end:**
`FileBuilder` assembles the whole file in memory, several copies of the
chunk (well over the 12 GiB the tests may use). Their layout messages
are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`).
- **`FileEditor`** refuses to rewrite such chunks (see
[its limits](#in-place-modification-fileeditor-limits)).
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
twice when it is shuffled (the inflated and the unshuffled copy, as
libhdf5's filters do too); writing one holds the chunk and its shuffled
copy. A selection decodes only the chunks it touches; one of an
unfiltered chunk reads only the rows it selects, through a memory map or
positioned reads (`File::open_storage`): every selection of
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop
-- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py
included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection
reads of the double-deflated fixture alone, 4.0 GiB.
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
reading or writing one is `FormatError::Overflow` ("exceeds the
addressable size" / "address space"); the file still opens and lists
(`examples/wasm-viewer/test/test.mjs`, "huge chunk").
- **libhdf5 2.2.0's h5dump cannot print these values:** built here from tag
`2.2.0`, it fails to inflate any deflated chunk over 4 GiB, those
libhdf5 2.0.0 writes included ("memory allocation failed for deflate
uncompression", `H5Zdeflate.c`), while h5py 3.16 (libhdf5 2.0.0) reads
them. The tests check h5dump 2.2.0 on layouts and storage sizes only
(`CLAWHDF5_H5DUMP2`).
- libhdf5 2.0.0 itself (h5py 3.16) drops a write to an unallocated,
unfiltered Single Chunk this large when the fill time is "never" (the
dataset stays unallocated); `fixtures/gen_huge_chunks.py` allocates it
early instead. Not a clawhdf5 issue; recorded because the generator
depends on it.
## HDF5 features still unsupported ## HDF5 features still unsupported
**Status:** open. What remains of the gaps the **Status:** open. What remains of the gaps the
@@ -564,6 +633,70 @@ under load from other builds) a freshly opened file with 8192 chunks
per dataset reads small selections 1.2x to 2.3x slower through a version-1 per dataset reads small selections 1.2x to 2.3x slower through a version-1
B-tree (see `BENCHMARKS.md`, "HDF5 1.8 format"). B-tree (see `BENCHMARKS.md`, "HDF5 1.8 format").
## Chunks of 4 GiB or more could not be read
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`); affected every
release (v2.1.0 to v2.7.0). Errors, and cost; no wrong values were
returned. Nothing for users to do but upgrade.
HDF5 2.0 writes chunks of more than 4 GiB - 1 bytes (layout message
version 5). `ChunkInfo::chunk_size` was a `u32`, so the Single Chunk,
Implicit, Fixed Array and Extensible Array readers truncated an
unfiltered chunk's size and the read failed ("incorrect chunk size
returned from index for unfiltered chunk"); a v2 B-tree index refused any
such chunk ("chunk larger than 4 GiB"), filtered or not. Filtered chunks
in the other indexes read, because their stored sizes were small. A
selection of a chunked dataset with a non-default fill value decoded the
whole dataset (8 GiB of output for the fixture's 2-D dataset), and a
selection of an unfiltered chunk fetched the whole chunk from a file that
is not in memory. `ChunkInfo::chunk_size` and `ChunkMapping::file_size` are now
`u64`; the selection path fills its box with the fill value and reads an
unfiltered chunk's rows only. Tests: `huge_chunks_interop`
(`filtered_huge_chunk_indexes_list` always; with `CLAWHDF5_HUGE_CHUNKS=1`
`filtered_huge_chunks_read`, over a fixture libhdf5 2.0.0 wrote, and
`unfiltered_huge_chunks_read`, over a 44 GiB sparse file h5py writes at
test time). Limits that remain:
[Chunks of 4 GiB or more](#chunks-of-4-gib-or-more-limits).
## Chunks of 4 GiB or more were written unreadable
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`). The deflate
truncation was before any release (the one-pass deflate dates from
2026-09-23); the rest affected every release (v2.1.0 to v2.7.0). Files
clawhdf5 wrote with a chunk of 4 GiB or more should be written again.
The writer gave such a chunk layout message version 4 (which libhdf5
before 2.0 cannot read, and which libhdf5 2.x never writes for it; whether
2.x reads what clawhdf5 wrote was not checked), cut a chunk dimension of
2^32 or more to 32 bits, and, since 2026-09-23, deflated only the first
4 GiB - 1 bytes of the chunk (zlib takes at most that much per call and
`Finish` ended the stream there): the chunk failed to decode ("decoded to
4294967295 bytes, expected 4294967304"). It now writes layout version 5
with libhdf5's index element widths (the Fixed and Extensible Array
structures match libhdf5 2.0.0's byte for byte:
`huge_chunk_array_indexes_match_libhdf5`), refuses chunk dimensions of
2^32 or more and the filters that cannot take such a chunk, deflates the
whole chunk, and no longer keeps the compressor's worst-case bound (4 GiB
of zeroed memory for a chunk that deflates to 4 MiB): writing the four
datasets of `writer_huge_chunks_round_trip` went past the tests' 12 GiB
cap, and compressing one shuffled 4 GiB chunk now peaks at 4.3 GiB
(tank, 2026-09-28). h5py 3.16 reads what it writes.
## LZ4 chunks larger than 256 MiB were refused
**Status:** fixed 2026-09-28 (branch `feat/huge-chunks`); affected every
release (v2.1.0 to v2.7.0). An error, never wrong data. Nothing for users
to do but upgrade.
The LZ4 decoder refused a chunk that decodes to more than 256 MiB
("lz4: declared size exceeds limit") even when the dataset's chunk size
bounded it; that ceiling is meant for a decode whose size is unknown, and
now applies only then (deflate never had it). A chunk of 4 GiB or more is
always read as the registered HDF5 framing, whose 64-bit size then no
longer starts with four zero bytes. Tests:
`filters::tests::lz4_chunks_over_256_mib_decode`,
`lz4_chunks_of_4_gib_use_the_registered_framing`.
## A dropped `FileEditor` could keep its file locked for a moment ## A dropped `FileEditor` could keep its file locked for a moment
**Status:** fixed 2026-09-28 (#23), before any release **Status:** fixed 2026-09-28 (#23), before any release