Before, an unfiltered chunk this size had its size cut to 32 bits in the single, implicit, fixed array and extensible array indexes. A v2 B-tree index refused any such chunk.
A selection on a chunked dataset with a non-default fill value now reads only the chunks it touches, instead of the whole dataset.
Writing:
These chunks use layout version 5 with libhdf5 2.x's size-field widths. The fixed and extensible array structures are byte-identical to libhdf5 2.0.0's.
Chunk dimensions of 2^32 or more were silently cut to 32 bits, and the LZF, bitshuffle, bzip2, Blosc and pcodec filters produced bad data on such chunks. Both are now refused before anything is written.
Bugs fixed along the way:
The one-pass deflate compressed only the first 4 GiB − 1 bytes of a chunk. Not in any release.
LZ4 chunks that decode to more than 256 MiB were refused. This affected v2.1.0–v2.7.0.
Memory: writing peaks at 4.3 GiB per 4 GiB shuffled chunk, and reading a double-deflated chunk dropped from 8.5 to 4.0 GiB.
FileEditor: rewriting such chunks is refused. Growing the extent and setting attributes still work.
32-bit targets: a clean FormatError::Overflow, checked under Node with the wasm package.
Integration with #28, added while stacking: a chunk of 4 GiB or more never goes into a v1 B-tree, whose key holds a 32-bit size. Under a 1.8 low bound it still takes layout version 5, as in libhdf5, and a high bound below 2.0 is LibverBound. Unit test: huge_chunks_ignore_the_v18_low_bound_and_need_v200.
BENCHMARKS: an idle re-run of #28's A/B (load under 2, five alternating runs each) supersedes that PR's loaded-machine numbers. It confirms the 8192-chunk slowdown and shows that the loaded run's odd 128-chunk windows were noise.
API break:ChunkInfo::chunk_size and ChunkMapping::file_size are now u64. ClawBrainHub doesn't use them.
Tests:
huge_chunks_interop.rs uses a committed 188 KB fixture from libhdf5 2.0.0 with filtered 4 GiB+ chunks in each index. Gated heavy tests (CLAWHDF5_HUGE_CHUNKS=1) cover a 44 GiB sparse unfiltered file; 6/6 pass under a 12 GiB memory cap.
h5py reads our huge-chunk files.
Not tested end to end: unfiltered 4 GiB+ writes (FileBuilder needs over 12 GiB of memory) and LZ4/Zstd at this size.
Verification:scripts/ci-test.sh on the whole stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204.
**Merge order: 3 of 3.** This PR is stacked on #28.
This adds chunks of 4 GiB or more (HDF5 2.0).
- **Reading, now in every chunk index:**
- Before, an unfiltered chunk this size had its size cut to 32 bits in the single, implicit, fixed array and extensible array indexes. A v2 B-tree index refused any such chunk.
- A selection on a chunked dataset with a non-default fill value now reads only the chunks it touches, instead of the whole dataset.
- **Writing:**
- These chunks use layout version 5 with libhdf5 2.x's size-field widths. The fixed and extensible array structures are byte-identical to libhdf5 2.0.0's.
- Chunk dimensions of 2^32 or more were silently cut to 32 bits, and the LZF, bitshuffle, bzip2, Blosc and pcodec filters produced bad data on such chunks. Both are now refused before anything is written.
- **Bugs fixed along the way:**
- The one-pass deflate compressed only the first 4 GiB − 1 bytes of a chunk. Not in any release.
- LZ4 chunks that decode to more than 256 MiB were refused. This affected v2.1.0–v2.7.0.
- Memory: writing peaks at 4.3 GiB per 4 GiB shuffled chunk, and reading a double-deflated chunk dropped from 8.5 to 4.0 GiB.
- **`FileEditor`:** rewriting such chunks is refused. Growing the extent and setting attributes still work.
- **32-bit targets:** a clean `FormatError::Overflow`, checked under Node with the wasm package.
**Integration with #28, added while stacking:** a chunk of 4 GiB or more never goes into a v1 B-tree, whose key holds a 32-bit size. Under a 1.8 low bound it still takes layout version 5, as in libhdf5, and a high bound below 2.0 is `LibverBound`. Unit test: `huge_chunks_ignore_the_v18_low_bound_and_need_v200`.
**BENCHMARKS:** an idle re-run of #28's A/B (load under 2, five alternating runs each) supersedes that PR's loaded-machine numbers. It confirms the 8192-chunk slowdown and shows that the loaded run's odd 128-chunk windows were noise.
**API break:** `ChunkInfo::chunk_size` and `ChunkMapping::file_size` are now `u64`. ClawBrainHub doesn't use them.
**Tests:**
- `huge_chunks_interop.rs` uses a committed 188 KB fixture from libhdf5 2.0.0 with filtered 4 GiB+ chunks in each index. Gated heavy tests (`CLAWHDF5_HUGE_CHUNKS=1`) cover a 44 GiB sparse unfiltered file; 6/6 pass under a 12 GiB memory cap.
- h5py reads our huge-chunk files.
**Not tested end to end:** unfiltered 4 GiB+ writes (`FileBuilder` needs over 12 GiB of memory) and LZ4/Zstd at this size.
**Verification:** `scripts/ci-test.sh` on the whole stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Variables got the first unused dimension of equal size, so a variable on
an unlimited dimension with fewer records got an anonymous dim_<n>, and
dimensions of one size could be swapped. Resolve them as netCDF-C does
(libhdf5/hdf5open.c): _Netcdf4Coordinates ids, else the scales
DIMENSION_LIST references (the last one attached to an axis), searched in
the variable's group and its parents; a coordinate variable is on its own
scale. Size matching remains only for axes the file names nothing for.
variables()/variable_names() leave out dimension scales that are only
dimensions, and _nc4_non_coord_<name> is the variable <name>.
Variable::shape is the netCDF shape (an unlimited dimension's length) and
the reads pad unwritten records with the fill value (_FillValue, else
NC_FILL_*; NaN from read_f64); Variable::stored_shape is the HDF5 extent.
New NetCDF4File::variable_names.
Tests compare with netCDF4-python variable by variable: the known-issues
reproducer, equal sizes, (p, p), scalars, inherited dimensions, unwritten
records, h5py dimension scales, h5netcdf and xarray files. CI installs
h5netcdf. known-issues entry moved to Fixed (history); stale open-table
row for the unlimited-size fix removed.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
New `LibVer` (V18, V110, V112, V114, V200, Latest) and
`libver_bounds(low, high)` on the format crate's `FileWriter` and the
facade's `FileBuilder`, as libhdf5's H5Pset_libver_bounds / h5py's
libver=(low, high). The default stays (V110, Latest), byte for byte what
was written before.
With a low bound of 1.8: superblock version 2, layout message version 3
(contiguous, compact, chunked) and a version-1 B-tree chunk index for
every chunked dataset, resizable ones included -- what libhdf5 2.x writes
under libver=('v108', 'latest'). The new chunk B-tree writer
(btree_v1_write.rs) replays H5B_insert with the H5Dbtree.c callbacks for
row-major insertion (split ratios 0.1/0.5/0.9, right keys moved as
H5D__btree_cmp3 moves them, root kept in place): its trees equal
libhdf5's node for node for 1-D/2-D/3-D, 2- and 3-level, filtered and
unfiltered datasets (libhdf5 writing without a chunk cache).
The high bound refuses, with FormatError::LibverBound before anything is
written, what needs a newer format: virtual datasets and the paged
file-space strategy (1.10), the 1.12 reference types (datatype v4),
native complex (datatype v5, HDF5 2.0), and a low bound above the high.
Tests: tools/tests/libver_v18.rs writes every writer feature under
(V18, V18), and HDF5 1.8.23's h5dump (scripts/build-hdf5-1.8.sh; skipped
when absent) dumps it exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset; h5py, clawhdf5 and h5rs check --data agree;
then FileEditor grows/appends/annotates it and h5py appends, and every
reader checks again. read_harness gains --v18 and --chunk N.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks;
known-issues moves "the writer does not produce output HDF5 1.8 can
read" to Fixed (history) as an opt-in fix and keeps what is still open
(Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs
libhdf5's chunk cache); README's capability table lists 1.8 output and
v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the
two formats on the read harness (tank, 2026-09-28, under load), which is
why the default stays the 1.10 format.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ChunkInfo::chunk_size (and ChunkMapping::file_size) are u64: sizes past
u32 were truncated for Single Chunk, Implicit, Fixed and Extensible Array
indexes, and a v2 B-tree index refused them. A selection of a chunked
dataset with a non-default fill value is read over a box of fill values
instead of a full read, an unfiltered chunk of a file that is not in memory
is read row by row, and an intermediate deflate stage no longer reserves the
chunk's whole bound.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The writer gives a chunk of more than u32::MAX bytes layout message
version 5 and, filtered, index elements whose stored size takes the file's
size of lengths, as libhdf5 2.x does (the Fixed and Extensible Array
structures match libhdf5's byte for byte). Chunk dimensions of 2^32 or more
and filters that cannot take such a chunk (LZF, bitshuffle, bzip2, Blosc,
pcodec) are refused instead of truncated. Chunks are extracted row by row
and one at a time; deflate no longer cuts input at 4 GiB - 1 bytes, nor
holds the worst-case bound of a large chunk; an LZ4 chunk of 4 GiB or more
is read as the registered framing.
FileEditor refuses writing values into, or pruning/allocating, chunks of
4 GiB or more before anything is written; growing the extent and
attributes still work.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A selection of a chunked dataset with a fill value is compared with the
full read in every index, through a map and through positioned reads. An
LZ4 chunk larger than 256 MiB is bounded by the chunk size, not refused.
The wasm package test reads the 4 GiB-chunk fixture and gets a clean
error in every index.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Three fixed entries (reading, writing, LZ4 over 256 MiB) and the open
limits entry; the selection-read entry loses its fill-value case; the
FileEditor limits gain the refused rewrites; the README tables name what
is supported and what is not.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Stacking feat/libver-v18 under feat/huge-chunks sent every chunked dataset
of a file with a 1.8 low bound to a version-1 B-tree, whose key holds a
32-bit chunk size. As in libhdf5, such a chunk now takes layout version 5
whatever the low bound, and a high bound below 2.0 is FormatError::LibverBound.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Merge order: 3 of 3. This PR is stacked on #28.
This adds chunks of 4 GiB or more (HDF5 2.0).
FileEditor: rewriting such chunks is refused. Growing the extent and setting attributes still work.FormatError::Overflow, checked under Node with the wasm package.Integration with #28, added while stacking: a chunk of 4 GiB or more never goes into a v1 B-tree, whose key holds a 32-bit size. Under a 1.8 low bound it still takes layout version 5, as in libhdf5, and a high bound below 2.0 is
LibverBound. Unit test:huge_chunks_ignore_the_v18_low_bound_and_need_v200.BENCHMARKS: an idle re-run of #28's A/B (load under 2, five alternating runs each) supersedes that PR's loaded-machine numbers. It confirms the 8192-chunk slowdown and shows that the loaded run's odd 128-chunk windows were noise.
API break:
ChunkInfo::chunk_sizeandChunkMapping::file_sizeare nowu64. ClawBrainHub doesn't use them.Tests:
huge_chunks_interop.rsuses a committed 188 KB fixture from libhdf5 2.0.0 with filtered 4 GiB+ chunks in each index. Gated heavy tests (CLAWHDF5_HUGE_CHUNKS=1) cover a 44 GiB sparse unfiltered file; 6/6 pass under a 12 GiB memory cap.Not tested end to end: unfiltered 4 GiB+ writes (
FileBuilderneeds over 12 GiB of memory) and LZ4/Zstd at this size.Verification:
scripts/ci-test.shon the whole stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204.🤖 Generated with Claude Code
New `LibVer` (V18, V110, V112, V114, V200, Latest) and `libver_bounds(low, high)` on the format crate's `FileWriter` and the facade's `FileBuilder`, as libhdf5's H5Pset_libver_bounds / h5py's libver=(low, high). The default stays (V110, Latest), byte for byte what was written before. With a low bound of 1.8: superblock version 2, layout message version 3 (contiguous, compact, chunked) and a version-1 B-tree chunk index for every chunked dataset, resizable ones included -- what libhdf5 2.x writes under libver=('v108', 'latest'). The new chunk B-tree writer (btree_v1_write.rs) replays H5B_insert with the H5Dbtree.c callbacks for row-major insertion (split ratios 0.1/0.5/0.9, right keys moved as H5D__btree_cmp3 moves them, root kept in place): its trees equal libhdf5's node for node for 1-D/2-D/3-D, 2- and 3-level, filtered and unfiltered datasets (libhdf5 writing without a chunk cache). The high bound refuses, with FormatError::LibverBound before anything is written, what needs a newer format: virtual datasets and the paged file-space strategy (1.10), the 1.12 reference types (datatype v4), native complex (datatype v5, HDF5 2.0), and a low bound above the high. Tests: tools/tests/libver_v18.rs writes every writer feature under (V18, V18), and HDF5 1.8.23's h5dump (scripts/build-hdf5-1.8.sh; skipped when absent) dumps it exactly as h5dump 1.14 does and returns our bytes for every numeric dataset; h5py, clawhdf5 and h5rs check --data agree; then FileEditor grows/appends/annotates it and h5py appends, and every reader checks again. read_harness gains --v18 and --chunk N. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>View command line instructions
Checkout
From your project repository, check out a new branch and test the changes.