Files
clawhdf5/crates
osobhandClaude Fable 5.1 05c665a898 feat(format): choose chunk dimensions automatically for large datasets
Requesting a filter without chunk dimensions made the whole dataset a single
chunk. Any read, even one row, then decompresses everything, and a large
dataset cannot be decoded in parallel — which also made the new partial reads
pointless for such files.

auto_chunk_dims keeps datasets up to 1 MiB as one chunk (unchanged behaviour)
and splits larger ones by halving the dimensions in turn, so chunks keep
roughly the dataset's proportions, until a chunk is at most 1 MiB — h5py's
approach. An empty (unlimited, unwritten) dimension is treated as 1024. The
writer passes the element size through resolve_chunk_dims_for; the old
resolve_chunk_dims assumes 8-byte elements. Explicit with_chunks always wins.

Interop test: h5py reads an auto-chunked 13 MB deflate dataset, sees chunks
between 128 KiB and 1 MiB, and a small dataset still has one chunk.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-19 14:21:04 -07:00
..
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00
2026-09-19 13:32:47 -07:00