py: in-place editing (clawhdf5.File(path, 'r+')) through FileEditor
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a FileEditor, and with it the file's exclusive lock, until close(): - ds[key] = value: h5py's keys and broadcasting (numpy's rules for slices and integers with extra leading 1-axes allowed; the exact shape for an index list, a scalar only where h5py expands it). Arrays are converted as libhdf5 converts them in native byte order (integers saturate, floats truncate toward zero and clip, integers go into h5py's bool enum by value); other values through numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer dataset is a ValueError instead of libhdf5's arbitrary value. The value preparation is a small Python module compiled into the extension (src/edit_helpers.py). - ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules. - attrs[name] = value, attrs.create(name, data, shape, dtype), attrs.modify: numeric, bool, complex, bytes and str data of any shape, with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor cannot write variable-length strings). - File.mode, File.flush(), Dataset.chunks. Each edit runs with the GIL released under the file handle's write lock (no read sees a half-written edit), then the file is reopened; datasets and attrs objects re-read their shape and attributes when the handle's edit generation moved. What the editor cannot do is NotImplementedError before anything is written: deleting attributes or objects, creating datasets or groups, compound fields by name, variable-length data, and FileEditor's own limits. Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft conversions in non-native byte order (a float in (-1, 0) becomes the integer minimum, same-size unsigned->signed wraps) and native casts that are undefined in C (half floats into unsigned, float(max) rounded up) -- clawhdf5 saturates as libhdf5's native path does; listed in docs/known-issues.md. Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to copies of the same file and both read back through h5py after each edit, on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed sequence over every chunk index kind, compact/contiguous/gzip layouts and numeric, bool, enum, complex, string and compound types, 16 random sequences of 40 edits, and a numeric conversion matrix; a refused edit must be refused by both and leave the file unchanged. Also dense attributes, locking, objects seeing edits, readers racing a writer, and h5dump (plus h5rs check in ci-test.sh) on every edited file. The read-vs-h5py suite also runs on a file opened 'r+'. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -2,6 +2,48 @@
|
||||
|
||||
## Unreleased
|
||||
|
||||
### Python bindings: in-place editing (2026-09-27)
|
||||
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
|
||||
file for editing through `clawhdf5::FileEditor`, holding its exclusive
|
||||
lock until `close()`:
|
||||
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
|
||||
one increasing index list) and broadcasting (numpy's rules for slices
|
||||
and integers, allowing extra leading length-1 axes; the exact shape for
|
||||
an index list, a scalar only where h5py expands it). A numpy array is
|
||||
converted to the dataset's dtype as libhdf5 converts it in native byte
|
||||
order (integers saturate; floats are truncated toward zero and clipped;
|
||||
integers go into h5py's bool enum by value, as libhdf5 stores them);
|
||||
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
|
||||
does. NaN into an integer dataset is a `ValueError`.
|
||||
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
|
||||
and errors (`TypeError` for a dataset that is not chunked).
|
||||
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
|
||||
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
|
||||
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
|
||||
enum, complex as the `r`/`i` compound), except that `str` becomes
|
||||
fixed-length UTF-8.
|
||||
- Every edit is written and synced before it returns, then the file is
|
||||
reopened: datasets and attrs objects taken earlier see new shapes and
|
||||
attributes, and reads on other threads wait while an edit is written.
|
||||
- What the editor cannot do raises `NotImplementedError` and writes
|
||||
nothing (deleting attributes or objects, creating datasets or groups,
|
||||
compound fields by name, variable-length data, ...;
|
||||
`docs/known-issues.md`).
|
||||
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
|
||||
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
|
||||
by h5py and by clawhdf5, and both read back through h5py after every
|
||||
edit, on files h5py writes with `libver` earliest, v114 and latest and on
|
||||
a clawhdf5-written one: a fixed sequence over every chunk index kind,
|
||||
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
|
||||
and compound types, and 16 random sequences of 40 edits (writes,
|
||||
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
|
||||
and leave its file unchanged. A matrix of every numeric source dtype into
|
||||
every numeric dataset dtype at the edge values, dense attribute storage,
|
||||
locking, objects seeing each other's edits, readers racing a writer
|
||||
(never a partly written dataset). Every edited file must pass `h5dump`
|
||||
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
|
||||
a file opened `'r+'`.
|
||||
|
||||
### Python bindings: remote files (2026-09-27)
|
||||
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
|
||||
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
|
||||
|
||||
Reference in New Issue
Block a user