Chunks of 4 GiB or more: read in every index, write as libhdf5 2.x does #29

Open
osobh wants to merge 10 commits from feat/huge-chunks into main
Owner

Merge order: 3 of 3. This PR is stacked on #28.

This adds chunks of 4 GiB or more (HDF5 2.0).

  • Reading, now in every chunk index:
    • Before, an unfiltered chunk this size had its size cut to 32 bits in the single, implicit, fixed array and extensible array indexes. A v2 B-tree index refused any such chunk.
    • A selection on a chunked dataset with a non-default fill value now reads only the chunks it touches, instead of the whole dataset.
  • Writing:
    • These chunks use layout version 5 with libhdf5 2.x's size-field widths. The fixed and extensible array structures are byte-identical to libhdf5 2.0.0's.
    • Chunk dimensions of 2^32 or more were silently cut to 32 bits, and the LZF, bitshuffle, bzip2, Blosc and pcodec filters produced bad data on such chunks. Both are now refused before anything is written.
  • Bugs fixed along the way:
    • The one-pass deflate compressed only the first 4 GiB − 1 bytes of a chunk. Not in any release.
    • LZ4 chunks that decode to more than 256 MiB were refused. This affected v2.1.0–v2.7.0.
    • Memory: writing peaks at 4.3 GiB per 4 GiB shuffled chunk, and reading a double-deflated chunk dropped from 8.5 to 4.0 GiB.
  • FileEditor: rewriting such chunks is refused. Growing the extent and setting attributes still work.
  • 32-bit targets: a clean FormatError::Overflow, checked under Node with the wasm package.

Integration with #28, added while stacking: a chunk of 4 GiB or more never goes into a v1 B-tree, whose key holds a 32-bit size. Under a 1.8 low bound it still takes layout version 5, as in libhdf5, and a high bound below 2.0 is LibverBound. Unit test: huge_chunks_ignore_the_v18_low_bound_and_need_v200.

BENCHMARKS: an idle re-run of #28's A/B (load under 2, five alternating runs each) supersedes that PR's loaded-machine numbers. It confirms the 8192-chunk slowdown and shows that the loaded run's odd 128-chunk windows were noise.

API break: ChunkInfo::chunk_size and ChunkMapping::file_size are now u64. ClawBrainHub doesn't use them.

Tests:

  • huge_chunks_interop.rs uses a committed 188 KB fixture from libhdf5 2.0.0 with filtered 4 GiB+ chunks in each index. Gated heavy tests (CLAWHDF5_HUGE_CHUNKS=1) cover a 44 GiB sparse unfiltered file; 6/6 pass under a 12 GiB memory cap.
  • h5py reads our huge-chunk files.

Not tested end to end: unfiltered 4 GiB+ writes (FileBuilder needs over 12 GiB of memory) and LZ4/Zstd at this size.

Verification: scripts/ci-test.sh on the whole stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204.

🤖 Generated with Claude Code

**Merge order: 3 of 3.** This PR is stacked on #28. This adds chunks of 4 GiB or more (HDF5 2.0). - **Reading, now in every chunk index:** - Before, an unfiltered chunk this size had its size cut to 32 bits in the single, implicit, fixed array and extensible array indexes. A v2 B-tree index refused any such chunk. - A selection on a chunked dataset with a non-default fill value now reads only the chunks it touches, instead of the whole dataset. - **Writing:** - These chunks use layout version 5 with libhdf5 2.x's size-field widths. The fixed and extensible array structures are byte-identical to libhdf5 2.0.0's. - Chunk dimensions of 2^32 or more were silently cut to 32 bits, and the LZF, bitshuffle, bzip2, Blosc and pcodec filters produced bad data on such chunks. Both are now refused before anything is written. - **Bugs fixed along the way:** - The one-pass deflate compressed only the first 4 GiB − 1 bytes of a chunk. Not in any release. - LZ4 chunks that decode to more than 256 MiB were refused. This affected v2.1.0–v2.7.0. - Memory: writing peaks at 4.3 GiB per 4 GiB shuffled chunk, and reading a double-deflated chunk dropped from 8.5 to 4.0 GiB. - **`FileEditor`:** rewriting such chunks is refused. Growing the extent and setting attributes still work. - **32-bit targets:** a clean `FormatError::Overflow`, checked under Node with the wasm package. **Integration with #28, added while stacking:** a chunk of 4 GiB or more never goes into a v1 B-tree, whose key holds a 32-bit size. Under a 1.8 low bound it still takes layout version 5, as in libhdf5, and a high bound below 2.0 is `LibverBound`. Unit test: `huge_chunks_ignore_the_v18_low_bound_and_need_v200`. **BENCHMARKS:** an idle re-run of #28's A/B (load under 2, five alternating runs each) supersedes that PR's loaded-machine numbers. It confirms the 8192-chunk slowdown and shows that the loaded run's odd 128-chunk windows were noise. **API break:** `ChunkInfo::chunk_size` and `ChunkMapping::file_size` are now `u64`. ClawBrainHub doesn't use them. **Tests:** - `huge_chunks_interop.rs` uses a committed 188 KB fixture from libhdf5 2.0.0 with filtered 4 GiB+ chunks in each index. Gated heavy tests (`CLAWHDF5_HUGE_CHUNKS=1`) cover a 44 GiB sparse unfiltered file; 6/6 pass under a 12 GiB memory cap. - h5py reads our huge-chunk files. **Not tested end to end:** unfiltered 4 GiB+ writes (`FileBuilder` needs over 12 GiB of memory) and LZ4/Zstd at this size. **Verification:** `scripts/ci-test.sh` on the whole stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
osobh added 10 commits 2026-09-29 05:01:01 +00:00
clawhdf5-netcdf4: variables' dimensions come from the file
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m24s
00b6f76ee0
Variables got the first unused dimension of equal size, so a variable on
an unlimited dimension with fewer records got an anonymous dim_<n>, and
dimensions of one size could be swapped. Resolve them as netCDF-C does
(libhdf5/hdf5open.c): _Netcdf4Coordinates ids, else the scales
DIMENSION_LIST references (the last one attached to an axis), searched in
the variable's group and its parents; a coordinate variable is on its own
scale. Size matching remains only for axes the file names nothing for.

variables()/variable_names() leave out dimension scales that are only
dimensions, and _nc4_non_coord_<name> is the variable <name>.
Variable::shape is the netCDF shape (an unlimited dimension's length) and
the reads pad unwritten records with the fill value (_FillValue, else
NC_FILL_*; NaN from read_f64); Variable::stored_shape is the HDF5 extent.
New NetCDF4File::variable_names.

Tests compare with netCDF4-python variable by variable: the known-issues
reproducer, equal sizes, (p, p), scalars, inherited dimensions, unwritten
records, h5py dimension scales, h5netcdf and xarray files. CI installs
h5netcdf. known-issues entry moved to Fixed (history); stale open-table
row for the unlimited-size fix removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
New `LibVer` (V18, V110, V112, V114, V200, Latest) and
`libver_bounds(low, high)` on the format crate's `FileWriter` and the
facade's `FileBuilder`, as libhdf5's H5Pset_libver_bounds / h5py's
libver=(low, high). The default stays (V110, Latest), byte for byte what
was written before.

With a low bound of 1.8: superblock version 2, layout message version 3
(contiguous, compact, chunked) and a version-1 B-tree chunk index for
every chunked dataset, resizable ones included -- what libhdf5 2.x writes
under libver=('v108', 'latest'). The new chunk B-tree writer
(btree_v1_write.rs) replays H5B_insert with the H5Dbtree.c callbacks for
row-major insertion (split ratios 0.1/0.5/0.9, right keys moved as
H5D__btree_cmp3 moves them, root kept in place): its trees equal
libhdf5's node for node for 1-D/2-D/3-D, 2- and 3-level, filtered and
unfiltered datasets (libhdf5 writing without a chunk cache).

The high bound refuses, with FormatError::LibverBound before anything is
written, what needs a newer format: virtual datasets and the paged
file-space strategy (1.10), the 1.12 reference types (datatype v4),
native complex (datatype v5, HDF5 2.0), and a low bound above the high.

Tests: tools/tests/libver_v18.rs writes every writer feature under
(V18, V18), and HDF5 1.8.23's h5dump (scripts/build-hdf5-1.8.sh; skipped
when absent) dumps it exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset; h5py, clawhdf5 and h5rs check --data agree;
then FileEditor grows/appends/annotates it and h5py appends, and every
reader checks again. read_harness gains --v18 and --chunk N.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
docs: HDF5 1.8 output (libver_bounds), its cost, and why the default stays
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m27s
01d2a5dc5d
CHANGELOG (Unreleased) records libver_bounds and the 1.8 checks;
known-issues moves "the writer does not produce output HDF5 1.8 can
read" to Fixed (history) as an opt-in fix and keeps what is still open
(Python 'w' has no libver, no pre-1.8 format, B-tree node filling vs
libhdf5's chunk cache); README's capability table lists 1.8 output and
v1 B-tree chunk indexes as written; BENCHMARKS gains the dated A/B of the
two formats on the read harness (tank, 2026-09-28, under load), which is
why the default stays the 1.10 format.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ChunkInfo::chunk_size (and ChunkMapping::file_size) are u64: sizes past
u32 were truncated for Single Chunk, Implicit, Fixed and Extensible Array
indexes, and a v2 B-tree index refused them. A selection of a chunked
dataset with a non-default fill value is read over a box of fill values
instead of a full read, an unfiltered chunk of a file that is not in memory
is read row by row, and an intermediate deflate stage no longer reserves the
chunk's whole bound.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The writer gives a chunk of more than u32::MAX bytes layout message
version 5 and, filtered, index elements whose stored size takes the file's
size of lengths, as libhdf5 2.x does (the Fixed and Extensible Array
structures match libhdf5's byte for byte). Chunk dimensions of 2^32 or more
and filters that cannot take such a chunk (LZF, bitshuffle, bzip2, Blosc,
pcodec) are refused instead of truncated. Chunks are extracted row by row
and one at a time; deflate no longer cuts input at 4 GiB - 1 bytes, nor
holds the worst-case bound of a large chunk; an LZ4 chunk of 4 GiB or more
is read as the registered framing.

FileEditor refuses writing values into, or pruning/allocating, chunks of
4 GiB or more before anything is written; growing the extent and
attributes still work.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A selection of a chunked dataset with a fill value is compared with the
full read in every index, through a map and through positioned reads. An
LZ4 chunk larger than 256 MiB is bounded by the chunk size, not refused.
The wasm package test reads the 4 GiB-chunk fixture and gets a clean
error in every index.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Three fixed entries (reading, writing, LZ4 over 256 MiB) and the open
limits entry; the selection-read entry loses its fill-value case; the
FileEditor limits gain the refused rewrites; the README tables name what
is supported and what is not.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Stacking feat/libver-v18 under feat/huge-chunks sent every chunked dataset
of a file with a 1.8 low bound to a version-1 B-tree, whose key holds a
32-bit chunk size. As in libhdf5, such a chunk now takes layout version 5
whatever the low bound, and a high bound below 2.0 is FormatError::LibverBound.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
docs: known-issues — cite PRs #27, #28, #29
CI / test-arm64 (pull_request) Successful in 1m36s
CI / test (pull_request) Successful in 16m57s
4b075656e7
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
All checks were successful
CI / test-arm64 (pull_request) Successful in 1m36s
CI / test (pull_request) Successful in 16m57s
You are not authorized to merge this pull request.
This pull request can be merged automatically.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin feat/huge-chunks:feat/huge-chunks
git checkout feat/huge-chunks
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: quantumclaw/clawhdf5#29