py: in-place editing (clawhdf5.File(path, 'r+')) through FileEditor

clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():

- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
  slices and integers with extra leading 1-axes allowed; the exact shape
  for an index list, a scalar only where h5py expands it). Arrays are
  converted as libhdf5 converts them in native byte order (integers
  saturate, floats truncate toward zero and clip, integers go into h5py's
  bool enum by value); other values through
  numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
  dataset is a ValueError instead of libhdf5's arbitrary value. The value
  preparation is a small Python module compiled into the extension
  (src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
  attrs.modify: numeric, bool, complex, bytes and str data of any shape,
  with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
  cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.

Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.

Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.

Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:57:36 -05:00
co-authored by Claude Opus 5.5
parent 910d81904c
commit 92fb0830e0
15 changed files with 1978 additions and 102 deletions
+39 -2
View File
@@ -98,6 +98,41 @@ chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
written on `close()`.
## Editing a file in place
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
where it is, through clawhdf5's `FileEditor`; the file is locked until
`close()`, and every edit is written and synced before the statement
returns.
```python
with clawhdf5.File("data.h5", "r+") as f:
ds = f["grid"]
ds[10:20, ::2] = 0 # h5py keys and broadcasting
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
f["series"].attrs["units"] = "K"
f.attrs.create("version", 2, dtype="u1")
```
- Values: a numpy array is converted to the dataset's dtype as libhdf5
converts it (integers saturate at the target's limits; floats are
truncated toward zero and clipped); anything else goes through
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
value). A few libhdf5 edge cases differ on purpose; see
`docs/known-issues.md`.
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
see its result.
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
`str` is stored as a fixed-length UTF-8 string (h5py stores a
variable-length one), so h5py reads it back as `bytes`.
- Not supported (`NotImplementedError`, nothing written): creating or
deleting datasets, groups and attributes, writing compound fields by
name, variable-length data, HDF5 array types, and whatever
`FileEditor` refuses (listed in `docs/known-issues.md`).
## Tests
```bash
@@ -108,8 +143,10 @@ pytest crates/clawhdf5-py/tests
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
writes, opened locally and over HTTP (an in-process range server,
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
failures, the GIL). `scripts/ci-test.sh` builds the wheel and runs these
in CI.
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
the wheel and runs these in CI.
## License