This closes the remaining limits of 4 GiB+ chunks from #29.
Chunk dimensions of 2^32 or more now read and write. DataLayout::Chunked::chunk_dimensions is Vec<u64>, and layout v4/v5 encode each dimension in 1 to 8 bytes. Layout v3 still refuses them, as libhdf5 does.
New 31 KB fixture from libhdf5 2.0.0, with u8 chunks of 2^32 + 7 in the single-chunk and extensible-array indexes.
Our output is read by h5py and h5dump 2.2.
Copy-free writes:
A chunk that is a contiguous run of the data is borrowed, not copied: the dataset's own shape, or blocks of whole rows. A filtered chunk of that kind is compressed straight from the data.
FileBuilder::write streams to a temporary file through FileWriter::finish_with.
Peak memory for one unfiltered 1 GiB chunk dropped from 5.0 to 1.0 GiB. A 4 GiB + 7 byte chunk now peaks at 4.00 GiB (/usr/bin/time -v, tank).
Output files are byte-identical to before.
New end-to-end tests (gated, CLAWHDF5_HUGE_CHUNKS=1): unfiltered 4 GiB+ chunks, and LZ4/Zstd at that size (h5py reads them through hdf5plugin). All 12 pass under a 12 GB memory cap, peaking at 10.04 GiB.
Also fixed:extract_chunk panicked on data shorter than the shape. It now zero-fills.
Speed, found and fixed while stacking: the streaming writer's first version used a 1 MiB write buffer.
That is above glibc's mmap threshold, so after the smaller benchmarks it was mapped afresh on every write: 1.6 M page faults against 13 K on main.
The result was 1 MiB chunked writes 1.35x to 1.83x slower in the full h5bench_write suite. Run alone, the case showed nothing.
With a 64 KiB buffer, all 27 write benchmarks are within 3% of main. The idle A/B, three alternating runs each, is in BENCHMARKS.md.
API break:chunk_dimensions is now Vec<u64>, and so are several &[u64] parameters (listed in the CHANGELOG). New: FileWriter::finish_with and DatasetBuilder::with_u8_data_owned.
Verification:scripts/ci-test.sh on the stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204.
**Merge order: 3 of 3.** This PR is stacked on #31.
This closes the remaining limits of 4 GiB+ chunks from #29.
- **Chunk dimensions of 2^32 or more** now read and write. `DataLayout::Chunked::chunk_dimensions` is `Vec<u64>`, and layout v4/v5 encode each dimension in 1 to 8 bytes. Layout v3 still refuses them, as libhdf5 does.
- New 31 KB fixture from libhdf5 2.0.0, with `u8` chunks of 2^32 + 7 in the single-chunk and extensible-array indexes.
- Our output is read by h5py and h5dump 2.2.
- **Copy-free writes:**
- A chunk that is a contiguous run of the data is borrowed, not copied: the dataset's own shape, or blocks of whole rows. A filtered chunk of that kind is compressed straight from the data.
- `FileBuilder::write` streams to a temporary file through `FileWriter::finish_with`.
- Peak memory for one unfiltered 1 GiB chunk dropped from 5.0 to 1.0 GiB. A 4 GiB + 7 byte chunk now peaks at 4.00 GiB (`/usr/bin/time -v`, tank).
- Output files are byte-identical to before.
- **New end-to-end tests** (gated, `CLAWHDF5_HUGE_CHUNKS=1`): unfiltered 4 GiB+ chunks, and LZ4/Zstd at that size (h5py reads them through hdf5plugin). All 12 pass under a 12 GB memory cap, peaking at 10.04 GiB.
- **Also fixed:** `extract_chunk` panicked on data shorter than the shape. It now zero-fills.
**Speed, found and fixed while stacking:** the streaming writer's first version used a 1 MiB write buffer.
- That is above glibc's mmap threshold, so after the smaller benchmarks it was mapped afresh on every write: 1.6 M page faults against 13 K on main.
- The result was 1 MiB chunked writes 1.35x to 1.83x slower in the full `h5bench_write` suite. Run alone, the case showed nothing.
- With a 64 KiB buffer, all 27 write benchmarks are within 3% of main. The idle A/B, three alternating runs each, is in BENCHMARKS.md.
**API break:** `chunk_dimensions` is now `Vec<u64>`, and so are several `&[u64]` parameters (listed in the CHANGELOG). New: `FileWriter::finish_with` and `DatasetBuilder::with_u8_data_owned`.
**Verification:** `scripts/ci-test.sh` on the stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204.
**Still open:**
- Five filters are refused for 4 GiB+ chunks.
- `FileEditor` doesn't rewrite such chunks.
- Filtered chunks are decoded whole.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Read a file's metadata the way netCDF-C 4.9.3 does (libhdf5/hdf5open.c),
for the whole file on first use (src/model.rs, replacing src/scope.rs):
- links in creation order when the group tracks it, else name order;
a group's datasets before its subgroups; dimension ids file-wide;
- variables' dimensions from _Netcdf4Coordinates (file-wide ids), else
the scales DIMENSION_LIST attaches when the first axis has one, else
netCDF-C's phony dimensions phony_dim_<id> (create_phony_dims: shared
by length and unlimitedness within a group, not between two axes of
one variable, numbered subgroups first, a zero length unlimited);
- datasets of types netCDF-C cannot represent are not variables
(references, bit fields, time, arrays, compounds/enums/VLENs over
them), replaying netCDF-C's file-wide type list, failed types
included;
- unlimited lengths as nc4_find_dim_len (its group and below).
NcType gains Enum, Compound, VLen, Opaque and is #[non_exhaustive];
Variable::nc_type is netCDF-C's type (1-byte strings NC_CHAR). New
clawhdf5_format::group_v2::links_in_creation_order_in.
Tests compare with netCDF-C itself (tests/netcdf_c_view.py calls the
libnetcdf netCDF4-python bundles through ctypes): new interop cases for
h5py files without dimension scales, every type class, link order; and
the gated corpus_vs_netcdf_c (CLAWHDF5_NETCDF_CORPUS): 420 of the 429
conformance-corpus files netCDF-C opens match (main: 68); the other 9
are explained in tests/corpus_known_differences.txt and known-issues.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Datatype::parse returns Datatype::Complex for class 11 (also inside
compounds, arrays and VL types) instead of the {r, i} compound view.
- Facade: DType::Complex(Box<DType>); read_complex_f32/f64 accept it.
- h5rs dump/ls/diff print native complex as h5dump/h5ls/h5diff 2.2.0 do
(checked against a fixture written by h5py 3.16 / libhdf5 2.0.0);
dump --json keeps the {r, i} compound (hdf5-json has no complex class).
- clawhdf5-wasm reads native complex datasets as [re, im] pairs.
- Python: clawhdf5.File(path, 'w', libver=...) with h5py's values,
mapped to FileBuilder::libver_bounds; 'v108' output opens in HDF5 1.8.23.
- Docs: known-issues entry moved to Fixed (history), CHANGELOG, READMEs.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- DataLayout::Chunked::chunk_dimensions is Vec<u64> (was Vec<u32>), and
the chunk index readers, writers and serializers take &[u64]: layout
messages of version 4/5 store each dimension in up to 8 bytes, and
libhdf5 2.x writes dimensions of 2^32 or more (layout version 5). Such
dimensions were refused on read (InvalidChunkDimensions) and write. A
version-3 layout (4-byte dimensions) is never written for them; a chunk
whose size overflows 64 bits is refused when opened.
- The file writer lays chunked datasets out as pieces referring to the
chunks instead of copying them into one buffer per pass, and an
unfiltered chunk that is a contiguous run of the dataset's data (a
dataset stored as one chunk of its shape, row blocks) borrows it; a
filtered one is compressed straight from it. Contiguous datasets are not
copied either. FileWriter::finish_with streams the file to a callback;
FileBuilder::write uses it, so the file is never assembled in memory.
DatasetBuilder::with_u8_data_owned takes the data without a copy.
Peak RSS writing one unfiltered 1 GiB chunk (with_u8_data_owned +
write): 5.0 GiB before, 1.0 GiB after; with_u8_data: 6.0 -> 2.0 GiB.
- extract_chunk no longer panics on data shorter than the shape.
- Tests: huge_chunk_dims.h5 fixture (libhdf5 2.0.0 via h5py 3.16, u8
chunks of 2^32 + 7), and opt-in end-to-end tests of chunk dims >= 2^32
(filtered and unfiltered, read and written, h5py and h5dump 2.2.0), of
an unfiltered 4 GiB+ chunk written by clawhdf5, and of LZ4/Zstd chunks
of that size; example write_one_chunk for memory measurements.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
known-issues: two fixed entries (chunk dimensions of 2^32 or more were
refused; a 4 GiB unfiltered chunk was held several times when written),
the open "Chunks of 4 GiB or more: limits" updated (LZ4/Zstd tested end
to end, memory figures of 2026-09-29) and the open-issues table in sync.
CHANGELOG (Unreleased) entry with the public API changes and the peak
RSS measurements; README capability table. The wasm package test lists
the new fixture and expects an Overflow error reading its chunks.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A 1 MiB BufWriter sits above glibc's 128 KiB mmap threshold; depending on
allocator history it was mapped afresh by every write, and its page faults
(1.6 M vs 13 K over the criterion write_2d_chunked group) made 1 MiB
chunked writes 1.35x-1.83x slower than main. With 64 KiB they match main.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Merge order: 3 of 3. This PR is stacked on #31.
This closes the remaining limits of 4 GiB+ chunks from #29.
DataLayout::Chunked::chunk_dimensionsisVec<u64>, and layout v4/v5 encode each dimension in 1 to 8 bytes. Layout v3 still refuses them, as libhdf5 does.u8chunks of 2^32 + 7 in the single-chunk and extensible-array indexes.FileBuilder::writestreams to a temporary file throughFileWriter::finish_with./usr/bin/time -v, tank).CLAWHDF5_HUGE_CHUNKS=1): unfiltered 4 GiB+ chunks, and LZ4/Zstd at that size (h5py reads them through hdf5plugin). All 12 pass under a 12 GB memory cap, peaking at 10.04 GiB.extract_chunkpanicked on data shorter than the shape. It now zero-fills.Speed, found and fixed while stacking: the streaming writer's first version used a 1 MiB write buffer.
h5bench_writesuite. Run alone, the case showed nothing.API break:
chunk_dimensionsis nowVec<u64>, and so are several&[u64]parameters (listed in the CHANGELOG). New:FileWriter::finish_withandDatasetBuilder::with_u8_data_owned.Verification:
scripts/ci-test.shon the stack passes 27/27 (tank, 2026-09-29), and ClawBrainHub passes 204/204.Still open:
FileEditordoesn't rewrite such chunks.🤖 Generated with Claude Code
- Datatype::parse returns Datatype::Complex for class 11 (also inside compounds, arrays and VL types) instead of the {r, i} compound view. - Facade: DType::Complex(Box<DType>); read_complex_f32/f64 accept it. - h5rs dump/ls/diff print native complex as h5dump/h5ls/h5diff 2.2.0 do (checked against a fixture written by h5py 3.16 / libhdf5 2.0.0); dump --json keeps the {r, i} compound (hdf5-json has no complex class). - clawhdf5-wasm reads native complex datasets as [re, im] pairs. - Python: clawhdf5.File(path, 'w', libver=...) with h5py's values, mapped to FileBuilder::libver_bounds; 'v108' output opens in HDF5 1.8.23. - Docs: known-issues entry moved to Fixed (history), CHANGELOG, READMEs. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>View command line instructions
Checkout
From your project repository, check out a new branch and test the changes.