edit: record the maximum before resizing a chunked dataset that has none

FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.

libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).

Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 07:40:01 -05:00
co-authored by Claude Opus 5.5
parent 92fb0830e0
commit 1f7651644b
8 changed files with 482 additions and 7 deletions
+27
View File
@@ -7,6 +7,33 @@ deleting it.
---
## Shrinking a chunked dataset with no recorded maximum scrambled it
**Status:** fixed 2026-09-27, before any release (`FileEditor::resize`
shipped on main in PR #18, a4c2ace). Files the writer produced before the
fix still lack the maximum; see *Existing files*.
clawhdf5's writer stored no maximum dimensions for a chunked dataset
created without a `maxshape` (libhdf5 always stores one, equal to the
dimensions when none is given). With none recorded, the maximum is the
current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index
places chunks by the maximum. `FileEditor::resize` changed only the
current dimensions, so a shrink moved every existing chunk and every
reader returned wrong values; after a shrink the dataset could not grow
back. Found by the review of the Python editing work.
**Fix:** before the first resize of such a dataset the editor records the
maximum libhdf5 would have written (the dimensions the index was built
with); the writer now records it for every chunked dataset. **Test:**
`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:**
datasets written before the fix have no recorded maximum. The fixed
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
does not** — it scrambles them the same way, and lets them grow past
their Fixed Array. Resize them once with a fixed `FileEditor` (a resize
to the same shape changes nothing; shrink and grow back) before letting
libhdf5 resize them. A dataset already shrunk by the unfixed
editor holds misplaced chunks; rewrite it from a good copy.
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to