Chunk dimensions of 2^32 or more; copy-free unfiltered 4 GiB writes #32

Open
osobh wants to merge 7 commits from feat/huge-chunk-dims into main
7 Commits
Author SHA1 Message Date
osobhandClaude Opus 5.5 4cc3111151 docs: known-issues — cite PRs #30, #31, #32
CI / test-arm64 (pull_request) Successful in 1m39s
CI / test (pull_request) Successful in 21m9s
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 21:33:48 -05:00
osobhandClaude Opus 5.5 7fba2f3996 BENCHMARKS: streaming FileBuilder::write A/B (the 1 MiB buffer regression and the fix)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 21:33:37 -05:00
osobhandClaude Opus 5.5 549e442aff FileBuilder::write: 64 KiB write buffer, not 1 MiB
A 1 MiB BufWriter sits above glibc's 128 KiB mmap threshold; depending on
allocator history it was mapped afresh by every write, and its page faults
(1.6 M vs 13 K over the criterion write_2d_chunked group) made 1 MiB
chunked writes 1.35x-1.83x slower than main. With 64 KiB they match main.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 21:16:01 -05:00
osobhandClaude Opus 5.5 728ffeef16 docs: chunk dimensions of 2^32 or more and copy-free 4 GiB writes
known-issues: two fixed entries (chunk dimensions of 2^32 or more were
refused; a 4 GiB unfiltered chunk was held several times when written),
the open "Chunks of 4 GiB or more: limits" updated (LZ4/Zstd tested end
to end, memory figures of 2026-09-29) and the open-issues table in sync.
CHANGELOG (Unreleased) entry with the public API changes and the peak
RSS measurements; README capability table. The wasm package test lists
the new fixture and expects an Overflow error reading its chunks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:40 -05:00
osobhandClaude Opus 5.5 e4b3a76dc7 Chunk dimensions of 2^32 or more; unfiltered 4 GiB chunks written without copies
- DataLayout::Chunked::chunk_dimensions is Vec<u64> (was Vec<u32>), and
  the chunk index readers, writers and serializers take &[u64]: layout
  messages of version 4/5 store each dimension in up to 8 bytes, and
  libhdf5 2.x writes dimensions of 2^32 or more (layout version 5). Such
  dimensions were refused on read (InvalidChunkDimensions) and write. A
  version-3 layout (4-byte dimensions) is never written for them; a chunk
  whose size overflows 64 bits is refused when opened.
- The file writer lays chunked datasets out as pieces referring to the
  chunks instead of copying them into one buffer per pass, and an
  unfiltered chunk that is a contiguous run of the dataset's data (a
  dataset stored as one chunk of its shape, row blocks) borrows it; a
  filtered one is compressed straight from it. Contiguous datasets are not
  copied either. FileWriter::finish_with streams the file to a callback;
  FileBuilder::write uses it, so the file is never assembled in memory.
  DatasetBuilder::with_u8_data_owned takes the data without a copy.
  Peak RSS writing one unfiltered 1 GiB chunk (with_u8_data_owned +
  write): 5.0 GiB before, 1.0 GiB after; with_u8_data: 6.0 -> 2.0 GiB.
- extract_chunk no longer panics on data shorter than the shape.
- Tests: huge_chunk_dims.h5 fixture (libhdf5 2.0.0 via h5py 3.16, u8
  chunks of 2^32 + 7), and opt-in end-to-end tests of chunk dims >= 2^32
  (filtered and unfiltered, read and written, h5py and h5dump 2.2.0), of
  an unfiltered 4 GiB+ chunk written by clawhdf5, and of LZ4/Zstd chunks
  of that size; example write_one_chunk for memory measurements.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:33 -05:00
osobhandClaude Opus 5.5 e4ba09946f HDF5 2.0 native complex as a first-class type on read; Python libver=
CI / test-arm64 (pull_request) Successful in 1m38s
CI / test (pull_request) Successful in 21m11s
- Datatype::parse returns Datatype::Complex for class 11 (also inside
  compounds, arrays and VL types) instead of the {r, i} compound view.
- Facade: DType::Complex(Box<DType>); read_complex_f32/f64 accept it.
- h5rs dump/ls/diff print native complex as h5dump/h5ls/h5diff 2.2.0 do
  (checked against a fixture written by h5py 3.16 / libhdf5 2.0.0);
  dump --json keeps the {r, i} compound (hdf5-json has no complex class).
- clawhdf5-wasm reads native complex datasets as [re, im] pairs.
- Python: clawhdf5.File(path, 'w', libver=...) with h5py's values,
  mapped to FileBuilder::libver_bounds; 'v108' output opens in HDF5 1.8.23.
- Docs: known-issues entry moved to Fixed (history), CHANGELOG, READMEs.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:33 -05:00
osobhandClaude Opus 5.5 e5d6f59e12 clawhdf5-netcdf4: phony dimensions, skipped types and order as netCDF-C
CI / test-arm64 (pull_request) Successful in 1m43s
CI / test (pull_request) Successful in 20m28s
Read a file's metadata the way netCDF-C 4.9.3 does (libhdf5/hdf5open.c),
for the whole file on first use (src/model.rs, replacing src/scope.rs):

- links in creation order when the group tracks it, else name order;
  a group's datasets before its subgroups; dimension ids file-wide;
- variables' dimensions from _Netcdf4Coordinates (file-wide ids), else
  the scales DIMENSION_LIST attaches when the first axis has one, else
  netCDF-C's phony dimensions phony_dim_<id> (create_phony_dims: shared
  by length and unlimitedness within a group, not between two axes of
  one variable, numbered subgroups first, a zero length unlimited);
- datasets of types netCDF-C cannot represent are not variables
  (references, bit fields, time, arrays, compounds/enums/VLENs over
  them), replaying netCDF-C's file-wide type list, failed types
  included;
- unlimited lengths as nc4_find_dim_len (its group and below).

NcType gains Enum, Compound, VLen, Opaque and is #[non_exhaustive];
Variable::nc_type is netCDF-C's type (1-byte strings NC_CHAR). New
clawhdf5_format::group_v2::links_in_creation_order_in.

Tests compare with netCDF-C itself (tests/netcdf_c_view.py calls the
libnetcdf netCDF4-python bundles through ctypes): new interop cases for
h5py files without dimension scales, every type class, link order; and
the gated corpus_vs_netcdf_c (CLAWHDF5_NETCDF_CORPUS): 420 of the 429
conformance-corpus files netCDF-C opens match (main: 68); the other 9
are explained in tests/corpus_known_differences.txt and known-issues.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:38:51 -05:00