Chunk dimensions of 2^32 or more; unfiltered 4 GiB chunks written without copies

- DataLayout::Chunked::chunk_dimensions is Vec<u64> (was Vec<u32>), and
  the chunk index readers, writers and serializers take &[u64]: layout
  messages of version 4/5 store each dimension in up to 8 bytes, and
  libhdf5 2.x writes dimensions of 2^32 or more (layout version 5). Such
  dimensions were refused on read (InvalidChunkDimensions) and write. A
  version-3 layout (4-byte dimensions) is never written for them; a chunk
  whose size overflows 64 bits is refused when opened.
- The file writer lays chunked datasets out as pieces referring to the
  chunks instead of copying them into one buffer per pass, and an
  unfiltered chunk that is a contiguous run of the dataset's data (a
  dataset stored as one chunk of its shape, row blocks) borrows it; a
  filtered one is compressed straight from it. Contiguous datasets are not
  copied either. FileWriter::finish_with streams the file to a callback;
  FileBuilder::write uses it, so the file is never assembled in memory.
  DatasetBuilder::with_u8_data_owned takes the data without a copy.
  Peak RSS writing one unfiltered 1 GiB chunk (with_u8_data_owned +
  write): 5.0 GiB before, 1.0 GiB after; with_u8_data: 6.0 -> 2.0 GiB.
- extract_chunk no longer panics on data shorter than the shape.
- Tests: huge_chunk_dims.h5 fixture (libhdf5 2.0.0 via h5py 3.16, u8
  chunks of 2^32 + 7), and opt-in end-to-end tests of chunk dims >= 2^32
  (filtered and unfiltered, read and written, h5py and h5dump 2.2.0), of
  an unfiltered 4 GiB+ chunk written by clawhdf5, and of LZ4/Zstd chunks
  of that size; example write_one_chunk for memory measurements.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-29 20:47:33 -05:00
co-authored by Claude Opus 5.5
parent e4ba09946f
commit e4b3a76dc7
17 changed files with 1156 additions and 256 deletions
+4 -4
View File
@@ -873,9 +873,9 @@ impl Checker<'_> {
self.counts.chunk_index_checksummed += 1;
}
let filtered = matches!(&info.filters, Ok(Some(p)) if !p.filters.is_empty());
let chunk_bytes = cdims.iter().try_fold(u64::from(dt.type_size()), |a, &d| {
a.checked_mul(u64::from(d))
});
let chunk_bytes = cdims
.iter()
.try_fold(u64::from(dt.type_size()), |a, &d| a.checked_mul(d));
let max = ds.max_dimensions.clone();
let mut seen: HashSet<Vec<u64>> = HashSet::with_capacity(chunks.len().min(1 << 20));
let mut reported = 0usize;
@@ -893,7 +893,7 @@ impl Checker<'_> {
));
} else {
for (i, (&o, &cd)) in c.offsets.iter().zip(cdims).enumerate() {
if o % u64::from(cd) != 0 {
if o % cd != 0 {
bad.push(format!(
"offset {o} in dimension {i} is not a multiple of the chunk size {cd}"
));
+1 -1
View File
@@ -292,7 +292,7 @@ impl Ls<'_> {
.unwrap_or(0);
let bytes = dims
.iter()
.try_fold(esize, |a, &d| a.checked_mul(u64::from(d)))
.try_fold(esize, |a, &d| a.checked_mul(d))
.map(|b| b.to_string())
.unwrap_or_else(|| "?".into());
writeln!(