py: in-place editing (clawhdf5.File(path, 'r+')) through FileEditor
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a FileEditor, and with it the file's exclusive lock, until close(): - ds[key] = value: h5py's keys and broadcasting (numpy's rules for slices and integers with extra leading 1-axes allowed; the exact shape for an index list, a scalar only where h5py expands it). Arrays are converted as libhdf5 converts them in native byte order (integers saturate, floats truncate toward zero and clip, integers go into h5py's bool enum by value); other values through numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer dataset is a ValueError instead of libhdf5's arbitrary value. The value preparation is a small Python module compiled into the extension (src/edit_helpers.py). - ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules. - attrs[name] = value, attrs.create(name, data, shape, dtype), attrs.modify: numeric, bool, complex, bytes and str data of any shape, with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor cannot write variable-length strings). - File.mode, File.flush(), Dataset.chunks. Each edit runs with the GIL released under the file handle's write lock (no read sees a half-written edit), then the file is reopened; datasets and attrs objects re-read their shape and attributes when the handle's edit generation moved. What the editor cannot do is NotImplementedError before anything is written: deleting attributes or objects, creating datasets or groups, compound fields by name, variable-length data, and FileEditor's own limits. Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft conversions in non-native byte order (a float in (-1, 0) becomes the integer minimum, same-size unsigned->signed wraps) and native casts that are undefined in C (half floats into unsigned, float(max) rounded up) -- clawhdf5 saturates as libhdf5's native path does; listed in docs/known-issues.md. Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to copies of the same file and both read back through h5py after each edit, on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed sequence over every chunk index kind, compact/contiguous/gzip layouts and numeric, bool, enum, complex, string and compound types, 16 random sequences of 40 edits, and a numeric conversion matrix; a refused edit must be refused by both and leave the file unchanged. Also dense attributes, locking, objects seeing edits, readers racing a writer, and h5dump (plus h5rs check in ci-test.sh) on every edited file. The read-vs-h5py suite also runs on a file opened 'r+'. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -116,6 +116,50 @@ space it leaves is too small for its next, larger version.
|
||||
**No journal.** A crash while an edit patches existing structures can leave
|
||||
the file inconsistent; see the `FileEditor` documentation.
|
||||
|
||||
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
|
||||
|
||||
**Status:** open (added 2026-09-27). The Python bindings edit through
|
||||
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
|
||||
before anything is written. On top of them:
|
||||
|
||||
- **No new or deleted objects:** `create_dataset`/`create_group` in an
|
||||
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
|
||||
`NotImplementedError` (the editor changes values, shapes and
|
||||
attributes only). Mode `'a'` works on an existing file only.
|
||||
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
|
||||
whole elements of the same structured dtype are), HDF5 array-type
|
||||
elements, variable-length data, strings padded with spaces or
|
||||
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
|
||||
ones, h5py's, are writable), compounds containing such strings, null
|
||||
dataspaces, and index-list writes of more than 2²² elements (write them
|
||||
in slices).
|
||||
- **`str` attributes are fixed-length UTF-8**, where h5py writes
|
||||
variable-length strings: h5py reads them back as `bytes`
|
||||
(`numpy.bytes_`), not `str`.
|
||||
- **Numeric conversion follows libhdf5's native-order results, not its
|
||||
bugs.** Arrays are converted as libhdf5 converts them (integers
|
||||
saturate, floats are truncated toward zero and clipped), checked value by
|
||||
value against h5py 3.16 / HDF5 2.0 in
|
||||
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
|
||||
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
|
||||
differs from h5py on purpose:
|
||||
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
|
||||
minimum or 2⁶³ depending on the type);
|
||||
- when the dataset or the array is not in native byte order, libhdf5's
|
||||
"soft" conversions store a float in (-1, 0) as the integer minimum and
|
||||
wrap an unsigned value too large for the signed type of the same size
|
||||
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
|
||||
native order;
|
||||
- libhdf5's native casts that are undefined in C: half floats into
|
||||
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
|
||||
integers (→ minimum), a float equal to the integer maximum rounded up
|
||||
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
|
||||
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
|
||||
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
|
||||
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
|
||||
- **Each edit reopens the file** (a new memory map) so that reads see it;
|
||||
reads from other threads wait while an edit is written.
|
||||
|
||||
## Selection reads that decode more than the selection
|
||||
|
||||
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
|
||||
|
||||
Reference in New Issue
Block a user