Lazy remote files in the browser (M4), SWMR reader (M5), Python remote reads and editing #19
@@ -2,6 +2,48 @@
|
|||||||
|
|
||||||
## Unreleased
|
## Unreleased
|
||||||
|
|
||||||
|
### Python bindings: in-place editing (2026-09-27)
|
||||||
|
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
|
||||||
|
file for editing through `clawhdf5::FileEditor`, holding its exclusive
|
||||||
|
lock until `close()`:
|
||||||
|
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
|
||||||
|
one increasing index list) and broadcasting (numpy's rules for slices
|
||||||
|
and integers, allowing extra leading length-1 axes; the exact shape for
|
||||||
|
an index list, a scalar only where h5py expands it). A numpy array is
|
||||||
|
converted to the dataset's dtype as libhdf5 converts it in native byte
|
||||||
|
order (integers saturate; floats are truncated toward zero and clipped;
|
||||||
|
integers go into h5py's bool enum by value, as libhdf5 stores them);
|
||||||
|
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
|
||||||
|
does. NaN into an integer dataset is a `ValueError`.
|
||||||
|
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
|
||||||
|
and errors (`TypeError` for a dataset that is not chunked).
|
||||||
|
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
|
||||||
|
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
|
||||||
|
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
|
||||||
|
enum, complex as the `r`/`i` compound), except that `str` becomes
|
||||||
|
fixed-length UTF-8.
|
||||||
|
- Every edit is written and synced before it returns, then the file is
|
||||||
|
reopened: datasets and attrs objects taken earlier see new shapes and
|
||||||
|
attributes, and reads on other threads wait while an edit is written.
|
||||||
|
- What the editor cannot do raises `NotImplementedError` and writes
|
||||||
|
nothing (deleting attributes or objects, creating datasets or groups,
|
||||||
|
compound fields by name, variable-length data, ...;
|
||||||
|
`docs/known-issues.md`).
|
||||||
|
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
|
||||||
|
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
|
||||||
|
by h5py and by clawhdf5, and both read back through h5py after every
|
||||||
|
edit, on files h5py writes with `libver` earliest, v114 and latest and on
|
||||||
|
a clawhdf5-written one: a fixed sequence over every chunk index kind,
|
||||||
|
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
|
||||||
|
and compound types, and 16 random sequences of 40 edits (writes,
|
||||||
|
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
|
||||||
|
and leave its file unchanged. A matrix of every numeric source dtype into
|
||||||
|
every numeric dataset dtype at the edge values, dense attribute storage,
|
||||||
|
locking, objects seeing each other's edits, readers racing a writer
|
||||||
|
(never a partly written dataset). Every edited file must pass `h5dump`
|
||||||
|
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
|
||||||
|
a file opened `'r+'`.
|
||||||
|
|
||||||
### Python bindings: remote files (2026-09-27)
|
### Python bindings: remote files (2026-09-27)
|
||||||
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
|
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
|
||||||
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
|
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
|
||||||
|
|||||||
@@ -542,6 +542,22 @@ f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 *
|
|||||||
headers={"Authorization": "Bearer ..."})
|
headers={"Authorization": "Bearer ..."})
|
||||||
```
|
```
|
||||||
|
|
||||||
|
An existing file opened with `"r+"` is edited in place (through
|
||||||
|
`clawhdf5::FileEditor`), with h5py's indexing, broadcasting and numeric
|
||||||
|
conversion; each edit is on disk when the statement returns:
|
||||||
|
|
||||||
|
```python
|
||||||
|
with clawhdf5.File("data.h5", "r+") as f:
|
||||||
|
f["group/temperatures"][100:200, ::4] = 0.0
|
||||||
|
f["series"].resize(5000, axis=0) # chunked datasets, within maxshape
|
||||||
|
f["series"][4000:] = new_values
|
||||||
|
f["group"].attrs["calibrated"] = True
|
||||||
|
```
|
||||||
|
|
||||||
|
Creating or deleting datasets, groups and attributes in an existing file is
|
||||||
|
not supported (`NotImplementedError`); limits are in
|
||||||
|
[known issues](docs/known-issues.md).
|
||||||
|
|
||||||
The default build reads `http://` URLs only; build with
|
The default build reads `http://` URLs only; build with
|
||||||
`maturin develop --release --features https` (rustls with ring, which
|
`maturin develop --release --features https` (rustls with ring, which
|
||||||
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
|
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
|
||||||
@@ -561,7 +577,8 @@ non-default fill value (`docs/known-issues.md`). An index list is read one
|
|||||||
group of neighbouring chunks at a time.
|
group of neighbouring chunks at a time.
|
||||||
Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`)
|
Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`)
|
||||||
covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests
|
covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests
|
||||||
in `crates/clawhdf5-py/tests` compare every read with h5py; run them with
|
in `crates/clawhdf5-py/tests` compare every read and every in-place edit
|
||||||
|
with h5py; run them with
|
||||||
`pip install pytest h5py && pytest crates/clawhdf5-py/tests`.
|
`pip install pytest h5py && pytest crates/clawhdf5-py/tests`.
|
||||||
|
|
||||||
### Agent Memory
|
### Agent Memory
|
||||||
|
|||||||
@@ -98,6 +98,41 @@ chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
|
|||||||
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
|
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
|
||||||
written on `close()`.
|
written on `close()`.
|
||||||
|
|
||||||
|
## Editing a file in place
|
||||||
|
|
||||||
|
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
|
||||||
|
where it is, through clawhdf5's `FileEditor`; the file is locked until
|
||||||
|
`close()`, and every edit is written and synced before the statement
|
||||||
|
returns.
|
||||||
|
|
||||||
|
```python
|
||||||
|
with clawhdf5.File("data.h5", "r+") as f:
|
||||||
|
ds = f["grid"]
|
||||||
|
ds[10:20, ::2] = 0 # h5py keys and broadcasting
|
||||||
|
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
|
||||||
|
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
|
||||||
|
f["series"].attrs["units"] = "K"
|
||||||
|
f.attrs.create("version", 2, dtype="u1")
|
||||||
|
```
|
||||||
|
|
||||||
|
- Values: a numpy array is converted to the dataset's dtype as libhdf5
|
||||||
|
converts it (integers saturate at the target's limits; floats are
|
||||||
|
truncated toward zero and clipped); anything else goes through
|
||||||
|
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
|
||||||
|
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
|
||||||
|
value). A few libhdf5 edge cases differ on purpose; see
|
||||||
|
`docs/known-issues.md`.
|
||||||
|
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
|
||||||
|
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
|
||||||
|
see its result.
|
||||||
|
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
|
||||||
|
`str` is stored as a fixed-length UTF-8 string (h5py stores a
|
||||||
|
variable-length one), so h5py reads it back as `bytes`.
|
||||||
|
- Not supported (`NotImplementedError`, nothing written): creating or
|
||||||
|
deleting datasets, groups and attributes, writing compound fields by
|
||||||
|
name, variable-length data, HDF5 array types, and whatever
|
||||||
|
`FileEditor` refuses (listed in `docs/known-issues.md`).
|
||||||
|
|
||||||
## Tests
|
## Tests
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
@@ -108,8 +143,10 @@ pytest crates/clawhdf5-py/tests
|
|||||||
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
|
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
|
||||||
writes, opened locally and over HTTP (an in-process range server,
|
writes, opened locally and over HTTP (an in-process range server,
|
||||||
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
|
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
|
||||||
failures, the GIL). `scripts/ci-test.sh` builds the wheel and runs these
|
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
|
||||||
in CI.
|
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
|
||||||
|
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
|
||||||
|
the wheel and runs these in CI.
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
+152
-55
@@ -1,23 +1,62 @@
|
|||||||
//! PyAttrs — dict-like access to HDF5 attributes.
|
//! PyAttrs — dict-like access to HDF5 attributes.
|
||||||
|
|
||||||
use std::sync::{Arc, Mutex};
|
use std::sync::{Arc, Mutex, PoisonError};
|
||||||
|
|
||||||
use clawhdf5_format::attribute::AttributeMessage;
|
use clawhdf5_format::attribute::AttributeMessage;
|
||||||
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
|
use pyo3::exceptions::{PyKeyError, PyNotImplementedError, PyTypeError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::{PyList, PyTuple};
|
use pyo3::types::{PyList, PyTuple};
|
||||||
|
|
||||||
use crate::convert::{Converter, Elements, resolve_vl};
|
use crate::convert::{Converter, Elements, resolve_vl};
|
||||||
use crate::handle::Handle;
|
use crate::handle::Handle;
|
||||||
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value};
|
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, edit, node, py_to_attr_value};
|
||||||
|
|
||||||
|
/// The attributes of an object in a file opened for reading (or editing).
|
||||||
|
struct ReadAttrs {
|
||||||
|
handle: Arc<Handle>,
|
||||||
|
addr: u64,
|
||||||
|
path: String,
|
||||||
|
/// Sorted by name, with the file generation they were read at: an edit
|
||||||
|
/// (`attrs[name] = value`, here or through another handle on the same
|
||||||
|
/// object) makes them re-read.
|
||||||
|
cache: Mutex<(u64, Arc<Vec<AttributeMessage>>)>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl ReadAttrs {
|
||||||
|
fn current(&self, py: Python<'_>) -> PyResult<Arc<Vec<AttributeMessage>>> {
|
||||||
|
let generation = self.handle.generation();
|
||||||
|
{
|
||||||
|
let cached = self.cache.lock().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
if cached.0 == generation {
|
||||||
|
return Ok(Arc::clone(&cached.1));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let (addr, path) = (self.addr, &self.path);
|
||||||
|
let attrs = Arc::new(self.handle.with(py, |f| node::attributes(f, addr, path))?);
|
||||||
|
*self.cache.lock().unwrap_or_else(PoisonError::into_inner) =
|
||||||
|
(generation, Arc::clone(&attrs));
|
||||||
|
Ok(attrs)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn check_writable(&self) -> PyResult<()> {
|
||||||
|
if self.handle.is_writable() {
|
||||||
|
return Ok(());
|
||||||
|
}
|
||||||
|
Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||||
|
"cannot set attributes on a read-only file (open it with mode 'r+')",
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn set(&self, py: Python<'_>, name: &str, value: clawhdf5_rs::AttrValue) -> PyResult<()> {
|
||||||
|
self.check_writable()?;
|
||||||
|
let path = node::name(&self.path);
|
||||||
|
self.handle.edit(py, |ed| ed.set_attr(&path, name, &value))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Backing storage for attributes.
|
/// Backing storage for attributes.
|
||||||
enum AttrsInner {
|
enum AttrsInner {
|
||||||
/// Attributes of an object in a file opened for reading, sorted by name.
|
Read(ReadAttrs),
|
||||||
Read {
|
|
||||||
handle: Arc<Handle>,
|
|
||||||
attrs: Vec<AttributeMessage>,
|
|
||||||
},
|
|
||||||
/// Writable attribute list shared with a parent (PyFile or PyGroup).
|
/// Writable attribute list shared with a parent (PyFile or PyGroup).
|
||||||
Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>),
|
Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>),
|
||||||
}
|
}
|
||||||
@@ -27,8 +66,11 @@ enum AttrsInner {
|
|||||||
/// In read mode, values are what h5py returns: numpy scalars for scalar
|
/// In read mode, values are what h5py returns: numpy scalars for scalar
|
||||||
/// attributes, numpy arrays otherwise, `str` for variable-length strings,
|
/// attributes, numpy arrays otherwise, `str` for variable-length strings,
|
||||||
/// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace.
|
/// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace.
|
||||||
/// In write mode, attributes set here are accumulated and written when
|
/// In a file opened with `'r+'`, `attrs[name] = value` adds or replaces an
|
||||||
/// the parent file is closed.
|
/// attribute in the file at once (as h5py stores it, except that `str`
|
||||||
|
/// values become fixed-length UTF-8 strings). In write mode (`'w'`),
|
||||||
|
/// attributes set here are accumulated and written when the parent file is
|
||||||
|
/// closed.
|
||||||
#[pyclass(name = "Attrs")]
|
#[pyclass(name = "Attrs")]
|
||||||
pub struct PyAttrs {
|
pub struct PyAttrs {
|
||||||
inner: AttrsInner,
|
inner: AttrsInner,
|
||||||
@@ -43,9 +85,15 @@ impl PyAttrs {
|
|||||||
addr: u64,
|
addr: u64,
|
||||||
path: &str,
|
path: &str,
|
||||||
) -> PyResult<Self> {
|
) -> PyResult<Self> {
|
||||||
let attrs = handle.with(py, |f| node::attributes(f, addr, path))?;
|
let generation = handle.generation();
|
||||||
|
let attrs = Arc::new(handle.with(py, |f| node::attributes(f, addr, path))?);
|
||||||
Ok(Self {
|
Ok(Self {
|
||||||
inner: AttrsInner::Read { handle, attrs },
|
inner: AttrsInner::Read(ReadAttrs {
|
||||||
|
handle,
|
||||||
|
addr,
|
||||||
|
path: path.to_string(),
|
||||||
|
cache: Mutex::new((generation, attrs)),
|
||||||
|
}),
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -55,14 +103,48 @@ impl PyAttrs {
|
|||||||
inner: AttrsInner::Write(store),
|
inner: AttrsInner::Write(store),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn set_value(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
key: &str,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
dtype: Option<&Bound<'_, PyAny>>,
|
||||||
|
shape: Option<&Bound<'_, PyAny>>,
|
||||||
|
) -> PyResult<()> {
|
||||||
|
match &self.inner {
|
||||||
|
AttrsInner::Read(r) => {
|
||||||
|
r.check_writable()?;
|
||||||
|
let value = edit::attr_value(py, value, dtype, shape)?;
|
||||||
|
r.set(py, key, value)
|
||||||
|
}
|
||||||
|
AttrsInner::Write(store) => {
|
||||||
|
if dtype.is_some() || shape.is_some() {
|
||||||
|
return Err(PyNotImplementedError::new_err(
|
||||||
|
"attrs.create with a dtype or shape is only supported in a file opened \
|
||||||
|
with 'r+'",
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let owned = py_to_attr_value(value)?;
|
||||||
|
let mut guard = store.lock().unwrap();
|
||||||
|
// Replace existing key if present.
|
||||||
|
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
|
||||||
|
entry.1 = owned;
|
||||||
|
} else {
|
||||||
|
guard.push((key.to_string(), owned));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
#[pymethods]
|
#[pymethods]
|
||||||
impl PyAttrs {
|
impl PyAttrs {
|
||||||
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { handle, attrs } => match attrs.iter().find(|a| a.name == key) {
|
AttrsInner::Read(r) => match r.current(py)?.iter().find(|a| a.name == key) {
|
||||||
Some(attr) => Ok(attr_to_py(py, handle, attr)?.unbind()),
|
Some(attr) => Ok(attr_to_py(py, &r.handle, attr)?.unbind()),
|
||||||
None => Err(PyKeyError::new_err(format!(
|
None => Err(PyKeyError::new_err(format!(
|
||||||
"Can't open attribute (can't locate attribute: '{key}')"
|
"Can't open attribute (can't locate attribute: '{key}')"
|
||||||
))),
|
))),
|
||||||
@@ -80,36 +162,63 @@ impl PyAttrs {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __setitem__(&self, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
/// `attrs[name] = value`. In a file opened with `'r+'` this writes the
|
||||||
|
/// attribute (numeric, bool, complex, bytes and str data, any shape)
|
||||||
|
/// into the file before returning; see the class docs.
|
||||||
|
fn __setitem__(&self, py: Python<'_>, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
||||||
|
self.set_value(py, key, value, None, None)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Deleting attributes is not supported in a file (the in-place editor
|
||||||
|
/// cannot remove them); in write mode it removes a pending attribute.
|
||||||
|
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { .. } => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
AttrsInner::Read(_) => Err(PyNotImplementedError::new_err(format!(
|
||||||
"cannot set attributes on a read-only file",
|
"cannot delete attribute '{key}': deleting attributes is not supported by \
|
||||||
)),
|
clawhdf5's in-place editor"
|
||||||
|
))),
|
||||||
AttrsInner::Write(store) => {
|
AttrsInner::Write(store) => {
|
||||||
let owned = py_to_attr_value(value)?;
|
|
||||||
let mut guard = store.lock().unwrap();
|
let mut guard = store.lock().unwrap();
|
||||||
// Replace existing key if present.
|
let before = guard.len();
|
||||||
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
|
guard.retain(|(k, _)| k != key);
|
||||||
entry.1 = owned;
|
if guard.len() == before {
|
||||||
} else {
|
return Err(PyKeyError::new_err(key.to_string()));
|
||||||
guard.push((key.to_string(), owned));
|
|
||||||
}
|
}
|
||||||
Ok(())
|
Ok(())
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __len__(&self) -> usize {
|
/// h5py's `attrs.create(name, data, shape=None, dtype=None)`: `data`
|
||||||
|
/// converted to `dtype` and reshaped to `shape` first.
|
||||||
|
#[pyo3(signature = (name, data, shape=None, dtype=None))]
|
||||||
|
fn create(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
name: &str,
|
||||||
|
data: &Bound<'_, PyAny>,
|
||||||
|
shape: Option<&Bound<'_, PyAny>>,
|
||||||
|
dtype: Option<&Bound<'_, PyAny>>,
|
||||||
|
) -> PyResult<()> {
|
||||||
|
self.set_value(py, name, data, dtype, shape)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// h5py's `attrs.modify(name, value)`: same as `attrs[name] = value`.
|
||||||
|
fn modify(&self, py: Python<'_>, name: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
||||||
|
self.set_value(py, name, value, None, None)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { attrs, .. } => attrs.len(),
|
AttrsInner::Read(r) => Ok(r.current(py)?.len()),
|
||||||
AttrsInner::Write(store) => store.lock().unwrap().len(),
|
AttrsInner::Write(store) => Ok(store.lock().unwrap().len()),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __contains__(&self, key: &str) -> bool {
|
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { attrs, .. } => attrs.iter().any(|a| a.name == key),
|
AttrsInner::Read(r) => Ok(r.current(py)?.iter().any(|a| a.name == key)),
|
||||||
AttrsInner::Write(store) => store.lock().unwrap().iter().any(|(k, _)| k == key),
|
AttrsInner::Write(store) => Ok(store.lock().unwrap().iter().any(|(k, _)| k == key)),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -119,15 +228,17 @@ impl PyAttrs {
|
|||||||
Ok(iter)
|
Ok(iter)
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __repr__(&self) -> String {
|
fn __repr__(&self, py: Python<'_>) -> String {
|
||||||
let n = self.__len__();
|
match self.__len__(py) {
|
||||||
format!("<HDF5 Attrs ({n} members)>")
|
Ok(n) => format!("<HDF5 Attrs ({n} members)>"),
|
||||||
|
Err(_) => "<HDF5 Attrs>".to_string(),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The value of `key`, or `default` if there is no such attribute.
|
/// The value of `key`, or `default` if there is no such attribute.
|
||||||
#[pyo3(signature = (key, default=None))]
|
#[pyo3(signature = (key, default=None))]
|
||||||
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
|
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
|
||||||
if self.__contains__(key) {
|
if self.__contains__(py, key)? {
|
||||||
self.__getitem__(py, key)
|
self.__getitem__(py, key)
|
||||||
} else {
|
} else {
|
||||||
Ok(default.unwrap_or_else(|| py.None()))
|
Ok(default.unwrap_or_else(|| py.None()))
|
||||||
@@ -137,7 +248,7 @@ impl PyAttrs {
|
|||||||
/// Return attribute names as a list.
|
/// Return attribute names as a list.
|
||||||
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
let names: Vec<String> = match &self.inner {
|
let names: Vec<String> = match &self.inner {
|
||||||
AttrsInner::Read { attrs, .. } => attrs.iter().map(|a| a.name.clone()).collect(),
|
AttrsInner::Read(r) => r.current(py)?.iter().map(|a| a.name.clone()).collect(),
|
||||||
AttrsInner::Write(store) => store
|
AttrsInner::Write(store) => store
|
||||||
.lock()
|
.lock()
|
||||||
.unwrap()
|
.unwrap()
|
||||||
@@ -152,9 +263,10 @@ impl PyAttrs {
|
|||||||
/// Return attribute values as a list.
|
/// Return attribute values as a list.
|
||||||
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
let vals: Vec<Py<PyAny>> = match &self.inner {
|
let vals: Vec<Py<PyAny>> = match &self.inner {
|
||||||
AttrsInner::Read { handle, attrs } => attrs
|
AttrsInner::Read(r) => r
|
||||||
|
.current(py)?
|
||||||
.iter()
|
.iter()
|
||||||
.map(|a| attr_to_py(py, handle, a).map(Bound::unbind))
|
.map(|a| attr_to_py(py, &r.handle, a).map(Bound::unbind))
|
||||||
.collect::<PyResult<_>>()?,
|
.collect::<PyResult<_>>()?,
|
||||||
AttrsInner::Write(store) => store
|
AttrsInner::Write(store) => store
|
||||||
.lock()
|
.lock()
|
||||||
@@ -173,9 +285,10 @@ impl PyAttrs {
|
|||||||
/// Return attribute (key, value) pairs as a list of tuples.
|
/// Return attribute (key, value) pairs as a list of tuples.
|
||||||
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
|
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
|
||||||
AttrsInner::Read { handle, attrs } => attrs
|
AttrsInner::Read(r) => r
|
||||||
|
.current(py)?
|
||||||
.iter()
|
.iter()
|
||||||
.map(|a| Ok((a.name.clone(), attr_to_py(py, handle, a)?.unbind())))
|
.map(|a| Ok((a.name.clone(), attr_to_py(py, &r.handle, a)?.unbind())))
|
||||||
.collect::<PyResult<_>>()?,
|
.collect::<PyResult<_>>()?,
|
||||||
AttrsInner::Write(store) => store
|
AttrsInner::Write(store) => store
|
||||||
.lock()
|
.lock()
|
||||||
@@ -248,19 +361,3 @@ fn prefix_err(py: Python<'_>, name: &str, e: PyErr) -> PyErr {
|
|||||||
PyValueError::new_err(msg)
|
PyValueError::new_err(msg)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
#[cfg(test)]
|
|
||||||
mod tests {
|
|
||||||
use super::*;
|
|
||||||
|
|
||||||
#[test]
|
|
||||||
fn write_attrs_len() {
|
|
||||||
let store = Arc::new(Mutex::new(Vec::new()));
|
|
||||||
store
|
|
||||||
.lock()
|
|
||||||
.unwrap()
|
|
||||||
.push(("key".into(), OwnedAttrValue::I64(99)));
|
|
||||||
let attrs = PyAttrs::from_write(store);
|
|
||||||
assert_eq!(attrs.__len__(), 1);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|||||||
@@ -9,12 +9,12 @@
|
|||||||
//! Python threads reading the same or different datasets run in parallel,
|
//! Python threads reading the same or different datasets run in parallel,
|
||||||
//! and a remote file's network reads never hold the GIL.
|
//! and a remote file's network reads never hold the GIL.
|
||||||
|
|
||||||
use std::sync::Arc;
|
use std::sync::{Arc, Mutex, PoisonError};
|
||||||
|
|
||||||
use clawhdf5_format::datatype::Datatype;
|
use clawhdf5_format::datatype::Datatype;
|
||||||
use clawhdf5_format::object_header::ObjectHeader;
|
use clawhdf5_format::object_header::ObjectHeader;
|
||||||
use clawhdf5_rs::File;
|
use clawhdf5_rs::File;
|
||||||
use pyo3::exceptions::{PyTypeError, PyValueError};
|
use pyo3::exceptions::{PyNotImplementedError, PyOSError, PyTypeError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::{PyList, PyTuple};
|
use pyo3::types::{PyList, PyTuple};
|
||||||
|
|
||||||
@@ -22,7 +22,7 @@ use crate::attrs::PyAttrs;
|
|||||||
use crate::convert::{Converter, Elements, VlError, resolve_vl};
|
use crate::convert::{Converter, Elements, VlError, resolve_vl};
|
||||||
use crate::handle::Handle;
|
use crate::handle::Handle;
|
||||||
use crate::select::{self, Plan};
|
use crate::select::{self, Plan};
|
||||||
use crate::{PyEmpty, node, to_py_err};
|
use crate::{PyEmpty, edit, node, to_py_err};
|
||||||
|
|
||||||
/// What opening a dataset reads from the file (without the GIL).
|
/// What opening a dataset reads from the file (without the GIL).
|
||||||
pub(crate) struct DatasetMeta {
|
pub(crate) struct DatasetMeta {
|
||||||
@@ -67,8 +67,9 @@ pub struct PyDataset {
|
|||||||
/// Where the dataset's object header is: reads open it from here rather
|
/// Where the dataset's object header is: reads open it from here rather
|
||||||
/// than resolve `path` again.
|
/// than resolve `path` again.
|
||||||
addr: u64,
|
addr: u64,
|
||||||
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
|
/// The shape (`None` for a null dataspace, h5py's `Empty`), with the
|
||||||
shape: Option<Vec<u64>>,
|
/// file generation it was read at: an edit (a resize) may change it.
|
||||||
|
shape: Mutex<(u64, Option<Vec<u64>>)>,
|
||||||
/// The chunk shape, for a chunked dataset.
|
/// The chunk shape, for a chunked dataset.
|
||||||
chunks: Option<Vec<u64>>,
|
chunks: Option<Vec<u64>>,
|
||||||
datatype: Datatype,
|
datatype: Datatype,
|
||||||
@@ -84,13 +85,14 @@ impl PyDataset {
|
|||||||
addr: u64,
|
addr: u64,
|
||||||
meta: DatasetMeta,
|
meta: DatasetMeta,
|
||||||
) -> Self {
|
) -> Self {
|
||||||
|
let generation = handle.generation();
|
||||||
let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
|
let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
|
||||||
.map_err(|e| e.value(py).to_string());
|
.map_err(|e| e.value(py).to_string());
|
||||||
Self {
|
Self {
|
||||||
handle,
|
handle,
|
||||||
path,
|
path,
|
||||||
addr,
|
addr,
|
||||||
shape: meta.shape,
|
shape: Mutex::new((generation, meta.shape)),
|
||||||
chunks: meta.chunks,
|
chunks: meta.chunks,
|
||||||
datatype: meta.datatype,
|
datatype: meta.datatype,
|
||||||
conv,
|
conv,
|
||||||
@@ -103,10 +105,54 @@ impl PyDataset {
|
|||||||
.map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path))))
|
.map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path))))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The current shape: the one read at open, or re-read after an edit.
|
||||||
|
fn dims(&self, py: Python<'_>) -> PyResult<Option<Vec<u64>>> {
|
||||||
|
let generation = self.handle.generation();
|
||||||
|
{
|
||||||
|
let cached = self.shape.lock().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
if cached.0 == generation {
|
||||||
|
return Ok(cached.1.clone());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let addr = self.addr;
|
||||||
|
let null = self
|
||||||
|
.shape
|
||||||
|
.lock()
|
||||||
|
.unwrap_or_else(PoisonError::into_inner)
|
||||||
|
.1
|
||||||
|
.is_none();
|
||||||
|
let shape = if null {
|
||||||
|
None
|
||||||
|
} else {
|
||||||
|
Some(self.handle.with(py, |f| {
|
||||||
|
f.dataset_at(addr)
|
||||||
|
.and_then(|ds| ds.shape())
|
||||||
|
.map_err(to_py_err)
|
||||||
|
})?)
|
||||||
|
};
|
||||||
|
*self.shape.lock().unwrap_or_else(PoisonError::into_inner) = (generation, shape.clone());
|
||||||
|
Ok(shape)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn check_writable(&self) -> PyResult<()> {
|
||||||
|
if self.handle.is_writable() {
|
||||||
|
Ok(())
|
||||||
|
} else {
|
||||||
|
Err(PyOSError::new_err(format!(
|
||||||
|
"{}: the file is open read-only; open it with mode 'r+' to change it",
|
||||||
|
node::name(&self.path)
|
||||||
|
)))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Read the selection described by `plan` into a numpy array.
|
/// Read the selection described by `plan` into a numpy array.
|
||||||
fn read_plan<'py>(&self, py: Python<'py>, plan: &Plan) -> PyResult<Bound<'py, PyAny>> {
|
fn read_plan<'py>(
|
||||||
|
&self,
|
||||||
|
py: Python<'py>,
|
||||||
|
plan: &Plan,
|
||||||
|
dims: &[u64],
|
||||||
|
) -> PyResult<Bound<'py, PyAny>> {
|
||||||
let conv = self.converter()?;
|
let conv = self.converter()?;
|
||||||
let dims = self.shape.as_deref().unwrap_or(&[]);
|
|
||||||
let out_shape = plan.out_shape();
|
let out_shape = plan.out_shape();
|
||||||
|
|
||||||
let arr = if plan.is_empty() {
|
let arr = if plan.is_empty() {
|
||||||
@@ -272,7 +318,7 @@ impl PyDataset {
|
|||||||
/// The shape of the dataset (`None` for an empty/null dataspace).
|
/// The shape of the dataset (`None` for an empty/null dataspace).
|
||||||
#[getter]
|
#[getter]
|
||||||
fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||||
match &self.shape {
|
match self.dims(py)? {
|
||||||
Some(s) => Ok(PyTuple::new(py, s)?.into_any()),
|
Some(s) => Ok(PyTuple::new(py, s)?.into_any()),
|
||||||
None => Ok(py.None().into_bound(py)),
|
None => Ok(py.None().into_bound(py)),
|
||||||
}
|
}
|
||||||
@@ -281,7 +327,7 @@ impl PyDataset {
|
|||||||
/// The maximum shape (`None` per unlimited dimension), like h5py.
|
/// The maximum shape (`None` per unlimited dimension), like h5py.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||||
let Some(shape) = &self.shape else {
|
let Some(shape) = self.dims(py)? else {
|
||||||
return Ok(py.None().into_bound(py));
|
return Ok(py.None().into_bound(py));
|
||||||
};
|
};
|
||||||
let addr = self.addr;
|
let addr = self.addr;
|
||||||
@@ -292,7 +338,7 @@ impl PyDataset {
|
|||||||
.and_then(|ds| ds.max_dimensions())
|
.and_then(|ds| ds.max_dimensions())
|
||||||
.map_err(to_py_err)
|
.map_err(to_py_err)
|
||||||
})?
|
})?
|
||||||
.unwrap_or_else(|| shape.clone());
|
.unwrap_or(shape);
|
||||||
let items: Vec<Option<u64>> = max
|
let items: Vec<Option<u64>> = max
|
||||||
.into_iter()
|
.into_iter()
|
||||||
.map(|d| (d != u64::MAX).then_some(d))
|
.map(|d| (d != u64::MAX).then_some(d))
|
||||||
@@ -300,6 +346,15 @@ impl PyDataset {
|
|||||||
Ok(PyTuple::new(py, items)?.into_any())
|
Ok(PyTuple::new(py, items)?.into_any())
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The chunk shape, or `None` for a dataset that is not chunked.
|
||||||
|
#[getter]
|
||||||
|
fn chunks<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||||
|
match &self.chunks {
|
||||||
|
Some(c) => Ok(PyTuple::new(py, c)?.into_any()),
|
||||||
|
None => Ok(py.None().into_bound(py)),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// The dataset's numpy dtype, as h5py reports it.
|
/// The dataset's numpy dtype, as h5py reports it.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn dtype<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
fn dtype<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||||
@@ -307,14 +362,14 @@ impl PyDataset {
|
|||||||
}
|
}
|
||||||
|
|
||||||
#[getter]
|
#[getter]
|
||||||
fn ndim(&self) -> usize {
|
fn ndim(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
self.shape.as_ref().map_or(0, Vec::len)
|
Ok(self.dims(py)?.map_or(0, |s| s.len()))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Number of elements (`None` for an empty/null dataspace, as h5py).
|
/// Number of elements (`None` for an empty/null dataspace, as h5py).
|
||||||
#[getter]
|
#[getter]
|
||||||
fn size(&self) -> Option<u64> {
|
fn size(&self, py: Python<'_>) -> PyResult<Option<u64>> {
|
||||||
self.shape.as_ref().map(|s| s.iter().product())
|
Ok(self.dims(py)?.map(|s| s.iter().product()))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The dataset's full name, e.g. `/group/data`.
|
/// The dataset's full name, e.g. `/group/data`.
|
||||||
@@ -323,7 +378,8 @@ impl PyDataset {
|
|||||||
node::name(&self.path)
|
node::name(&self.path)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The dataset's attributes (read-only, dict-like).
|
/// The dataset's attributes (dict-like; writable in a file opened with
|
||||||
|
/// `'r+'`).
|
||||||
#[getter]
|
#[getter]
|
||||||
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||||
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
|
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
|
||||||
@@ -338,7 +394,7 @@ impl PyDataset {
|
|||||||
py: Python<'py>,
|
py: Python<'py>,
|
||||||
key: &Bound<'py, PyAny>,
|
key: &Bound<'py, PyAny>,
|
||||||
) -> PyResult<Bound<'py, PyAny>> {
|
) -> PyResult<Bound<'py, PyAny>> {
|
||||||
let Some(dims) = &self.shape else {
|
let Some(dims) = self.dims(py)? else {
|
||||||
let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty());
|
let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty());
|
||||||
let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>();
|
let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>();
|
||||||
if is_empty_tuple || is_ellipsis {
|
if is_empty_tuple || is_ellipsis {
|
||||||
@@ -347,8 +403,109 @@ impl PyDataset {
|
|||||||
}
|
}
|
||||||
return Err(PyValueError::new_err("Empty datasets cannot be sliced"));
|
return Err(PyValueError::new_err("Empty datasets cannot be sliced"));
|
||||||
};
|
};
|
||||||
let plan = select::parse(key, dims)?;
|
let plan = select::parse(key, &dims)?;
|
||||||
self.read_plan(py, &plan)
|
self.read_plan(py, &plan, &dims)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Write with h5py indexing (file opened with `'r+'`): `ds[key] = value`.
|
||||||
|
///
|
||||||
|
/// The key is what `ds[key]` reads (without compound field names). The
|
||||||
|
/// value is converted to the dataset's dtype as h5py converts it (a
|
||||||
|
/// numpy array as libhdf5 does, clipping out-of-range numbers; anything
|
||||||
|
/// else through `numpy.asarray(value, dtype=ds.dtype)`), and broadcast
|
||||||
|
/// to the selection as h5py broadcasts. The edit is written and synced
|
||||||
|
/// before this returns; what the in-place editor cannot write raises
|
||||||
|
/// `NotImplementedError` and leaves the file as it was.
|
||||||
|
fn __setitem__(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
key: &Bound<'_, PyAny>,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
) -> PyResult<()> {
|
||||||
|
self.check_writable()?;
|
||||||
|
let Some(dims) = self.dims(py)? else {
|
||||||
|
return Err(PyNotImplementedError::new_err(
|
||||||
|
"writing to an empty (null dataspace) dataset is not supported",
|
||||||
|
));
|
||||||
|
};
|
||||||
|
let plan = select::parse(key, &dims)?;
|
||||||
|
if !plan.fields.is_empty() {
|
||||||
|
return Err(PyNotImplementedError::new_err(
|
||||||
|
"writing compound fields by name is not supported by clawhdf5's in-place editor; \
|
||||||
|
write whole elements",
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let category = edit::category(&self.datatype)?;
|
||||||
|
let conv = self.converter()?;
|
||||||
|
let bytes = edit::dataset_bytes(
|
||||||
|
py,
|
||||||
|
value,
|
||||||
|
conv.dtype.bind(py),
|
||||||
|
category,
|
||||||
|
&plan,
|
||||||
|
self.chunks.as_deref(),
|
||||||
|
)?;
|
||||||
|
if plan.is_empty() {
|
||||||
|
return Ok(());
|
||||||
|
}
|
||||||
|
let sel = edit::selection(&plan, &dims)?;
|
||||||
|
let path = node::name(&self.path);
|
||||||
|
self.handle
|
||||||
|
.edit(py, |ed| ed.write_selection(&path, &sel, &bytes))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Change the dataset's shape (file opened with `'r+'`), as h5py's
|
||||||
|
/// `Dataset.resize`: `ds.resize((100, 20))`, or `ds.resize(100, axis=0)`.
|
||||||
|
/// Only chunked datasets, within their maximum shape; new elements read
|
||||||
|
/// as the fill value.
|
||||||
|
#[pyo3(signature = (size, axis=None))]
|
||||||
|
fn resize(&self, py: Python<'_>, size: &Bound<'_, PyAny>, axis: Option<isize>) -> PyResult<()> {
|
||||||
|
self.check_writable()?;
|
||||||
|
let Some(dims) = self.dims(py)? else {
|
||||||
|
return Err(PyTypeError::new_err("Empty datasets cannot be resized"));
|
||||||
|
};
|
||||||
|
if self.chunks.is_none() {
|
||||||
|
return Err(PyTypeError::new_err("Only chunked datasets can be resized"));
|
||||||
|
}
|
||||||
|
let shape: Vec<u64> = match axis {
|
||||||
|
Some(axis) => {
|
||||||
|
let rank = dims.len();
|
||||||
|
let a = usize::try_from(axis)
|
||||||
|
.ok()
|
||||||
|
.filter(|&a| a < rank)
|
||||||
|
.ok_or_else(|| {
|
||||||
|
PyValueError::new_err(format!(
|
||||||
|
"Invalid axis (0 to {} allowed)",
|
||||||
|
rank.saturating_sub(1)
|
||||||
|
))
|
||||||
|
})?;
|
||||||
|
let n: u64 = size.extract().map_err(|_| {
|
||||||
|
PyTypeError::new_err("Argument must be a single int if axis is specified")
|
||||||
|
})?;
|
||||||
|
let mut s = dims.clone();
|
||||||
|
s[a] = n;
|
||||||
|
s
|
||||||
|
}
|
||||||
|
// As h5py: without `axis` the size is a sequence (`tuple(size)`).
|
||||||
|
None => size.extract().map_err(|_| {
|
||||||
|
PyTypeError::new_err(format!(
|
||||||
|
"'{}' object is not iterable",
|
||||||
|
size.get_type()
|
||||||
|
.name()
|
||||||
|
.map(|n| n.to_string())
|
||||||
|
.unwrap_or_default()
|
||||||
|
))
|
||||||
|
})?,
|
||||||
|
};
|
||||||
|
if shape.len() != dims.len() {
|
||||||
|
return Err(PyValueError::new_err(format!(
|
||||||
|
"new shape {shape:?} has {} dimensions, the dataset {}",
|
||||||
|
shape.len(),
|
||||||
|
dims.len()
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
let path = node::name(&self.path);
|
||||||
|
self.handle.edit(py, |ed| ed.resize(&path, &shape))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// `numpy.asarray(ds)` reads the whole dataset.
|
/// `numpy.asarray(ds)` reads the whole dataset.
|
||||||
@@ -360,20 +517,20 @@ impl PyDataset {
|
|||||||
copy: Option<bool>,
|
copy: Option<bool>,
|
||||||
) -> PyResult<Bound<'py, PyAny>> {
|
) -> PyResult<Bound<'py, PyAny>> {
|
||||||
let _ = copy; // every read is a fresh array
|
let _ = copy; // every read is a fresh array
|
||||||
let Some(dims) = &self.shape else {
|
let Some(dims) = self.dims(py)? else {
|
||||||
return Err(PyValueError::new_err("an empty dataset has no array value"));
|
return Err(PyValueError::new_err("an empty dataset has no array value"));
|
||||||
};
|
};
|
||||||
let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any();
|
let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any();
|
||||||
let plan = select::parse(&ellipsis, dims)?;
|
let plan = select::parse(&ellipsis, &dims)?;
|
||||||
let arr = self.read_plan(py, &plan)?;
|
let arr = self.read_plan(py, &plan, &dims)?;
|
||||||
match dtype {
|
match dtype {
|
||||||
Some(dt) => arr.call_method1("astype", (dt,)),
|
Some(dt) => arr.call_method1("astype", (dt,)),
|
||||||
None => Ok(arr),
|
None => Ok(arr),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __len__(&self) -> PyResult<usize> {
|
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
match self.shape.as_deref() {
|
match self.dims(py)?.as_deref() {
|
||||||
Some([first, ..]) => Ok(*first as usize),
|
Some([first, ..]) => Ok(*first as usize),
|
||||||
_ => Err(PyTypeError::new_err(
|
_ => Err(PyTypeError::new_err(
|
||||||
"Attempt to take len() of scalar dataset",
|
"Attempt to take len() of scalar dataset",
|
||||||
@@ -391,9 +548,10 @@ impl PyDataset {
|
|||||||
.unwrap_or_default(),
|
.unwrap_or_default(),
|
||||||
Err(_) => format!("{:?}", self.datatype),
|
Err(_) => format!("{:?}", self.datatype),
|
||||||
};
|
};
|
||||||
let shape = match &self.shape {
|
let shape = match self.dims(py) {
|
||||||
Some(s) => format!("{s:?}"),
|
Ok(Some(s)) => format!("{s:?}"),
|
||||||
None => "None".to_string(),
|
Ok(None) => "None".to_string(),
|
||||||
|
Err(_) => "?".to_string(),
|
||||||
};
|
};
|
||||||
format!(
|
format!(
|
||||||
"<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">",
|
"<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">",
|
||||||
|
|||||||
@@ -0,0 +1,387 @@
|
|||||||
|
//! In-place editing (`clawhdf5.File(path, 'r+')`) through `FileEditor`:
|
||||||
|
//! turning what Python assigns into the bytes, selections and attribute
|
||||||
|
//! values the editor takes.
|
||||||
|
//!
|
||||||
|
//! Value conversion follows h5py (see `edit_helpers.py`, run inside the
|
||||||
|
//! extension module); what `FileEditor` cannot do is `NotImplementedError`
|
||||||
|
//! before anything is written.
|
||||||
|
|
||||||
|
use std::ffi::CString;
|
||||||
|
|
||||||
|
use clawhdf5_format::datatype::{
|
||||||
|
CharacterSet, CompoundMember, Datatype, DatatypeByteOrder, EnumMember, StringPadding,
|
||||||
|
};
|
||||||
|
use clawhdf5_format::selection::Selection;
|
||||||
|
use clawhdf5_rs::AttrValue;
|
||||||
|
use pyo3::exceptions::{PyNotImplementedError, PyTypeError};
|
||||||
|
use pyo3::prelude::*;
|
||||||
|
use pyo3::sync::PyOnceLock;
|
||||||
|
use pyo3::types::{PyBytes, PyModule, PyTuple};
|
||||||
|
|
||||||
|
use crate::select::{Axis, Plan};
|
||||||
|
|
||||||
|
/// The helper module, compiled once.
|
||||||
|
pub(crate) fn helpers(py: Python<'_>) -> PyResult<&Bound<'_, PyModule>> {
|
||||||
|
static HELPERS: PyOnceLock<Py<PyModule>> = PyOnceLock::new();
|
||||||
|
let module = HELPERS.get_or_try_init(py, || -> PyResult<Py<PyModule>> {
|
||||||
|
let code = CString::new(include_str!("edit_helpers.py"))
|
||||||
|
.map_err(|e| PyTypeError::new_err(e.to_string()))?;
|
||||||
|
Ok(PyModule::from_code(
|
||||||
|
py,
|
||||||
|
&code,
|
||||||
|
c"clawhdf5/edit_helpers.py",
|
||||||
|
c"clawhdf5._edit_helpers",
|
||||||
|
)?
|
||||||
|
.unbind())
|
||||||
|
})?;
|
||||||
|
Ok(module.bind(py))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn not_implemented(what: impl std::fmt::Display) -> PyErr {
|
||||||
|
PyNotImplementedError::new_err(format!(
|
||||||
|
"{what} is not supported by clawhdf5's in-place editor"
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// How values for a dataset of type `dt` are converted (a category of
|
||||||
|
/// `edit_helpers._convert_array`), or why they cannot be written.
|
||||||
|
pub(crate) fn category(dt: &Datatype) -> PyResult<&'static str> {
|
||||||
|
match dt {
|
||||||
|
Datatype::FixedPoint { .. } => Ok("int"),
|
||||||
|
Datatype::FloatingPoint { .. } => Ok("float"),
|
||||||
|
Datatype::Enumeration {
|
||||||
|
base_type, members, ..
|
||||||
|
} => {
|
||||||
|
let is_bool = base_type.type_size() == 1
|
||||||
|
&& members.len() == 2
|
||||||
|
&& members
|
||||||
|
.iter()
|
||||||
|
.any(|m| m.name == "FALSE" && m.value.first() == Some(&0))
|
||||||
|
&& members
|
||||||
|
.iter()
|
||||||
|
.any(|m| m.name == "TRUE" && m.value.first() == Some(&1));
|
||||||
|
Ok(if is_bool { "bool" } else { "enum" })
|
||||||
|
}
|
||||||
|
Datatype::String {
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
..
|
||||||
|
} => Ok("string"),
|
||||||
|
Datatype::String { padding, .. } => Err(not_implemented(format!(
|
||||||
|
"writing fixed-length strings padded {padding:?} (libhdf5 converts them \
|
||||||
|
differently from numpy)"
|
||||||
|
))),
|
||||||
|
Datatype::Compound { size, members } => {
|
||||||
|
if is_complex(*size, members) {
|
||||||
|
return Ok("complex");
|
||||||
|
}
|
||||||
|
check_exact(dt)?;
|
||||||
|
Ok("exact")
|
||||||
|
}
|
||||||
|
Datatype::Opaque { .. } => Ok("exact"),
|
||||||
|
Datatype::Array { .. } => Err(not_implemented("writing HDF5 array-type elements")),
|
||||||
|
Datatype::VariableLength { .. } => Err(not_implemented("writing variable-length data")),
|
||||||
|
Datatype::Reference { .. } => Err(not_implemented("writing references")),
|
||||||
|
Datatype::BitField { .. } => Err(not_implemented("writing bitfields")),
|
||||||
|
Datatype::Time { .. } => Err(not_implemented("writing time values")),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// h5py's complex numbers: a compound of two identical floats `r`, `i`.
|
||||||
|
fn is_complex(size: u32, members: &[CompoundMember]) -> bool {
|
||||||
|
matches!(members, [r, i] if r.name == "r" && i.name == "i"
|
||||||
|
&& r.datatype == i.datatype
|
||||||
|
&& matches!(r.datatype, Datatype::FloatingPoint { size: fs, .. }
|
||||||
|
if r.byte_offset == 0 && i.byte_offset == u64::from(fs) && size == 2 * fs))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Compound members written byte for byte from the same numpy dtype: fine
|
||||||
|
/// unless libhdf5 would convert them on the way (strings padded other than
|
||||||
|
/// with NULs), or the editor cannot write them at all.
|
||||||
|
fn check_exact(dt: &Datatype) -> PyResult<()> {
|
||||||
|
match dt {
|
||||||
|
Datatype::Compound { members, .. } => {
|
||||||
|
members.iter().try_for_each(|m| check_exact(&m.datatype))
|
||||||
|
}
|
||||||
|
Datatype::Array { base_type, .. } => check_exact(base_type),
|
||||||
|
Datatype::String {
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
..
|
||||||
|
}
|
||||||
|
| Datatype::FixedPoint { .. }
|
||||||
|
| Datatype::FloatingPoint { .. }
|
||||||
|
| Datatype::Enumeration { .. }
|
||||||
|
| Datatype::Opaque { .. }
|
||||||
|
| Datatype::BitField { .. } => Ok(()),
|
||||||
|
Datatype::String { .. } => Err(not_implemented(
|
||||||
|
"writing compounds with strings not padded with NULs",
|
||||||
|
)),
|
||||||
|
Datatype::VariableLength { .. } => Err(not_implemented(
|
||||||
|
"writing compounds with variable-length members",
|
||||||
|
)),
|
||||||
|
Datatype::Reference { .. } => Err(not_implemented("writing references")),
|
||||||
|
Datatype::Time { .. } => Err(not_implemented("writing time values")),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Largest point selection an index-list write builds (one coordinate
|
||||||
|
/// vector per element).
|
||||||
|
const MAX_POINTS: usize = 1 << 22;
|
||||||
|
|
||||||
|
/// The selection `plan` writes, whose elements are numbered as the value's
|
||||||
|
/// (row-major over the selection's shape).
|
||||||
|
pub(crate) fn selection(plan: &Plan, dims: &[u64]) -> PyResult<Selection> {
|
||||||
|
if plan.axes.is_empty() {
|
||||||
|
return Ok(Selection::All);
|
||||||
|
}
|
||||||
|
if plan.list_axis().is_none() {
|
||||||
|
let (reads, _) = plan.reads(dims, None, 1);
|
||||||
|
return match <[_; 1]>::try_from(reads) {
|
||||||
|
Ok([read]) => Ok(read.sel),
|
||||||
|
Err(_) => Err(PyTypeError::new_err("internal error: several hyperslabs")),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
// An index list: the points, in the value's order.
|
||||||
|
let per_axis: Vec<Vec<u64>> = plan
|
||||||
|
.axes
|
||||||
|
.iter()
|
||||||
|
.map(|a| match a {
|
||||||
|
Axis::Index(i) => vec![*i],
|
||||||
|
Axis::Slice { start, step, count } => (0..*count).map(|k| start + k * step).collect(),
|
||||||
|
Axis::List(v) => v.clone(),
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
let n = per_axis
|
||||||
|
.iter()
|
||||||
|
.try_fold(1usize, |acc, v| acc.checked_mul(v.len()))
|
||||||
|
.filter(|&n| n <= MAX_POINTS)
|
||||||
|
.ok_or_else(|| {
|
||||||
|
not_implemented(format!(
|
||||||
|
"an index-list write of more than {MAX_POINTS} elements (write it in slices)"
|
||||||
|
))
|
||||||
|
})?;
|
||||||
|
let mut points = Vec::with_capacity(n);
|
||||||
|
let mut at = vec![0usize; per_axis.len()];
|
||||||
|
for _ in 0..n {
|
||||||
|
points.push(at.iter().zip(&per_axis).map(|(&i, v)| v[i]).collect());
|
||||||
|
for d in (0..at.len()).rev() {
|
||||||
|
at[d] += 1;
|
||||||
|
if at[d] < per_axis[d].len() {
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
at[d] = 0;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Ok(Selection::Points(points))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The bytes to write for `value` under `plan`, in the dataset's dtype.
|
||||||
|
pub(crate) fn dataset_bytes(
|
||||||
|
py: Python<'_>,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
dtype: &Bound<'_, PyAny>,
|
||||||
|
category: &str,
|
||||||
|
plan: &Plan,
|
||||||
|
chunks: Option<&[u64]>,
|
||||||
|
) -> PyResult<Vec<u8>> {
|
||||||
|
let shape = PyTuple::new(py, plan.out_shape())?;
|
||||||
|
let fancy = plan.list_axis().is_some();
|
||||||
|
let chunk_elems = chunks.map_or(0, |c| c.iter().fold(1u64, |a, &d| a.saturating_mul(d)));
|
||||||
|
let bytes = helpers(py)?.call_method1(
|
||||||
|
"dataset_values",
|
||||||
|
(value, dtype, category, shape, fancy, chunk_elems),
|
||||||
|
)?;
|
||||||
|
Ok(bytes.cast::<PyBytes>()?.as_bytes().to_vec())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn ieee_float(size: u32, byte_order: DatatypeByteOrder) -> Option<Datatype> {
|
||||||
|
let (exponent_location, exponent_size, mantissa_size, exponent_bias) = match size {
|
||||||
|
2 => (10, 5, 10, 15),
|
||||||
|
4 => (23, 8, 23, 127),
|
||||||
|
8 => (52, 11, 52, 1023),
|
||||||
|
_ => return None,
|
||||||
|
};
|
||||||
|
Some(Datatype::FloatingPoint {
|
||||||
|
size,
|
||||||
|
byte_order,
|
||||||
|
bit_offset: 0,
|
||||||
|
bit_precision: (size * 8) as u16,
|
||||||
|
exponent_location,
|
||||||
|
exponent_size,
|
||||||
|
mantissa_location: 0,
|
||||||
|
mantissa_size,
|
||||||
|
exponent_bias,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The HDF5 datatype h5py writes for a numpy dtype string (`'<i4'`,
|
||||||
|
/// `'|b1'`, `'>f8'`, `'<c16'`, `'|S5'`).
|
||||||
|
fn datatype_of(dtype: &str) -> Option<Datatype> {
|
||||||
|
let order = match dtype.as_bytes().first()? {
|
||||||
|
b'<' | b'|' | b'=' => DatatypeByteOrder::LittleEndian,
|
||||||
|
b'>' => DatatypeByteOrder::BigEndian,
|
||||||
|
_ => return None,
|
||||||
|
};
|
||||||
|
let kind = dtype.as_bytes().get(1)?;
|
||||||
|
let size: u32 = dtype.get(2..)?.parse().ok()?;
|
||||||
|
match kind {
|
||||||
|
b'b' if size == 1 => Some(Datatype::Enumeration {
|
||||||
|
size: 1,
|
||||||
|
base_type: Box::new(Datatype::FixedPoint {
|
||||||
|
size: 1,
|
||||||
|
byte_order: DatatypeByteOrder::LittleEndian,
|
||||||
|
signed: true,
|
||||||
|
bit_offset: 0,
|
||||||
|
bit_precision: 8,
|
||||||
|
}),
|
||||||
|
members: vec![
|
||||||
|
EnumMember {
|
||||||
|
name: "FALSE".into(),
|
||||||
|
value: vec![0],
|
||||||
|
},
|
||||||
|
EnumMember {
|
||||||
|
name: "TRUE".into(),
|
||||||
|
value: vec![1],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
}),
|
||||||
|
b'i' | b'u' if matches!(size, 1 | 2 | 4 | 8) => Some(Datatype::FixedPoint {
|
||||||
|
size,
|
||||||
|
byte_order: order,
|
||||||
|
signed: *kind == b'i',
|
||||||
|
bit_offset: 0,
|
||||||
|
bit_precision: (size * 8) as u16,
|
||||||
|
}),
|
||||||
|
b'f' => ieee_float(size, order),
|
||||||
|
b'c' => {
|
||||||
|
let part = ieee_float(size / 2, order)?;
|
||||||
|
Some(Datatype::Compound {
|
||||||
|
size,
|
||||||
|
members: vec![
|
||||||
|
CompoundMember {
|
||||||
|
name: "r".into(),
|
||||||
|
byte_offset: 0,
|
||||||
|
datatype: part.clone(),
|
||||||
|
},
|
||||||
|
CompoundMember {
|
||||||
|
name: "i".into(),
|
||||||
|
byte_offset: u64::from(size / 2),
|
||||||
|
datatype: part,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
})
|
||||||
|
}
|
||||||
|
b'S' if size > 0 => Some(Datatype::String {
|
||||||
|
size,
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
charset: CharacterSet::Ascii,
|
||||||
|
}),
|
||||||
|
_ => None,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// An attribute value as h5py would store it (`attrs[name] = value`, or
|
||||||
|
/// `attrs.create(name, data, shape, dtype)`), except that `str` data is
|
||||||
|
/// stored as fixed-length UTF-8 strings (h5py stores variable-length ones,
|
||||||
|
/// which the editor cannot write).
|
||||||
|
pub(crate) fn attr_value(
|
||||||
|
py: Python<'_>,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
dtype: Option<&Bound<'_, PyAny>>,
|
||||||
|
shape: Option<&Bound<'_, PyAny>>,
|
||||||
|
) -> PyResult<AttrValue> {
|
||||||
|
if value.is_instance_of::<crate::PyEmpty>() {
|
||||||
|
return Err(not_implemented(
|
||||||
|
"writing an empty (null dataspace) attribute",
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let (kind, dt, dims, data): (String, String, Vec<u64>, Vec<u8>) = helpers(py)?
|
||||||
|
.call_method1("attr_value", (value, dtype, shape))?
|
||||||
|
.extract()?;
|
||||||
|
let datatype = if kind == "str" {
|
||||||
|
let size: u32 = dt
|
||||||
|
.parse()
|
||||||
|
.map_err(|_| PyTypeError::new_err("bad string size"))?;
|
||||||
|
Datatype::String {
|
||||||
|
size,
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
charset: CharacterSet::Utf8,
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
datatype_of(&dt).ok_or_else(|| not_implemented(format!("an attribute of dtype {dt}")))?
|
||||||
|
};
|
||||||
|
Ok(AttrValue::Raw {
|
||||||
|
datatype,
|
||||||
|
shape: dims,
|
||||||
|
data,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod tests {
|
||||||
|
use super::*;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn numpy_dtypes_map_to_h5py_types() {
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of("<i4"),
|
||||||
|
Some(Datatype::FixedPoint {
|
||||||
|
size: 4,
|
||||||
|
signed: true,
|
||||||
|
byte_order: DatatypeByteOrder::LittleEndian,
|
||||||
|
..
|
||||||
|
})
|
||||||
|
));
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of(">u2"),
|
||||||
|
Some(Datatype::FixedPoint {
|
||||||
|
size: 2,
|
||||||
|
signed: false,
|
||||||
|
byte_order: DatatypeByteOrder::BigEndian,
|
||||||
|
..
|
||||||
|
})
|
||||||
|
));
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of("<f2"),
|
||||||
|
Some(Datatype::FloatingPoint { size: 2, .. })
|
||||||
|
));
|
||||||
|
let c = datatype_of("<c16").unwrap();
|
||||||
|
assert!(matches!(&c, Datatype::Compound { size: 16, members } if is_complex(16, members)));
|
||||||
|
assert_eq!(category(&c).unwrap(), "complex");
|
||||||
|
let b = datatype_of("|b1").unwrap();
|
||||||
|
assert_eq!(category(&b).unwrap(), "bool");
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of("|S5"),
|
||||||
|
Some(Datatype::String { size: 5, .. })
|
||||||
|
));
|
||||||
|
assert!(datatype_of("<f16").is_none());
|
||||||
|
assert!(datatype_of("<M8").is_none());
|
||||||
|
assert!(datatype_of("|S0").is_none());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn list_writes_become_points_in_value_order() {
|
||||||
|
let plan = Plan {
|
||||||
|
axes: vec![
|
||||||
|
Axis::List(vec![1, 4]),
|
||||||
|
Axis::Slice {
|
||||||
|
start: 0,
|
||||||
|
step: 2,
|
||||||
|
count: 2,
|
||||||
|
},
|
||||||
|
Axis::Index(3),
|
||||||
|
],
|
||||||
|
fields: vec![],
|
||||||
|
scalar: false,
|
||||||
|
};
|
||||||
|
let sel = selection(&plan, &[5, 4, 4]).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
sel,
|
||||||
|
Selection::Points(vec![
|
||||||
|
vec![1, 0, 3],
|
||||||
|
vec![1, 2, 3],
|
||||||
|
vec![4, 0, 3],
|
||||||
|
vec![4, 2, 3]
|
||||||
|
])
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,175 @@
|
|||||||
|
"""Values for in-place writes (clawhdf5.File(path, 'r+')), prepared the way
|
||||||
|
h5py prepares them, so `ds[key] = value` stores what h5py would store.
|
||||||
|
|
||||||
|
Loaded by the extension module (src/edit.rs); not a public API.
|
||||||
|
|
||||||
|
h5py converts in two ways, and so does this module:
|
||||||
|
|
||||||
|
- a value that is not a numpy array (a list, a Python or numpy scalar) is
|
||||||
|
converted by numpy straight to the dataset's dtype
|
||||||
|
(`numpy.asarray(value, dtype=ds.dtype)`), with numpy's rules and errors;
|
||||||
|
- a numpy array is converted by libhdf5, whose numeric conversions clip to
|
||||||
|
the target's range instead of wrapping: integers saturate, floats are
|
||||||
|
truncated toward zero and clipped, a double too large for a float becomes
|
||||||
|
infinity. That is what `_convert_array` reproduces. Where libhdf5 has no
|
||||||
|
meaningful answer — NaN into an integer, for which it writes a different
|
||||||
|
arbitrary value per type — this raises ValueError instead of guessing.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
|
||||||
|
def _no_path(src, dst):
|
||||||
|
return TypeError(f"No conversion path for dtype: {src!r} -> {dst!r}")
|
||||||
|
|
||||||
|
|
||||||
|
def _to_int(arr, dtype):
|
||||||
|
"""Integer target: libhdf5's saturating conversion."""
|
||||||
|
info = np.iinfo(dtype)
|
||||||
|
kind = arr.dtype.kind
|
||||||
|
if kind == "b":
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if kind in "iu":
|
||||||
|
src = np.iinfo(arr.dtype)
|
||||||
|
lo = max(info.min, src.min)
|
||||||
|
hi = min(info.max, src.max)
|
||||||
|
clipped = np.clip(arr, np.array(lo, arr.dtype), np.array(hi, arr.dtype))
|
||||||
|
return clipped.astype(dtype)
|
||||||
|
if kind == "f":
|
||||||
|
if np.isnan(arr).any():
|
||||||
|
raise ValueError(
|
||||||
|
"cannot write NaN to an integer dataset (libhdf5 would store an arbitrary value)"
|
||||||
|
)
|
||||||
|
t = np.trunc(arr.astype(np.float64))
|
||||||
|
# info.max + 1 and info.min are powers of two: exact as floats.
|
||||||
|
over = t >= float(info.max + 1)
|
||||||
|
under = t < float(info.min)
|
||||||
|
out = np.where(over | under, 0.0, t).astype(dtype)
|
||||||
|
out[over] = info.max
|
||||||
|
out[under] = info.min
|
||||||
|
return out
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
|
||||||
|
|
||||||
|
def _convert_array(arr, dtype, category):
|
||||||
|
kind = arr.dtype.kind
|
||||||
|
if category in ("int", "enum"):
|
||||||
|
if arr.dtype == dtype and kind in "iu":
|
||||||
|
return arr
|
||||||
|
if category == "enum" and kind not in "iu":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
return _to_int(arr, dtype)
|
||||||
|
if category == "bool":
|
||||||
|
if kind == "b":
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if kind in "iu":
|
||||||
|
# h5py's bool is an enum over int8; libhdf5 converts integers
|
||||||
|
# into it by value (saturating), not to FALSE/TRUE, so 3 is
|
||||||
|
# stored as 3. Keep those bytes: a view, not a cast.
|
||||||
|
return _to_int(arr, np.dtype("i1")).view(dtype)
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
if category == "float":
|
||||||
|
if kind not in "biuf":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
with np.errstate(over="ignore", invalid="ignore"):
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if category == "complex":
|
||||||
|
if kind != "c":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
with np.errstate(over="ignore", invalid="ignore"):
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if category == "string":
|
||||||
|
if kind != "S":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
return arr.astype(dtype)
|
||||||
|
# "exact": compound and opaque types, written only from the same dtype.
|
||||||
|
if arr.dtype == dtype:
|
||||||
|
return arr
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
|
||||||
|
|
||||||
|
def _convert_other(value, dtype, category):
|
||||||
|
if category == "string":
|
||||||
|
items = np.asarray(value, dtype=object)
|
||||||
|
if any(isinstance(x, str) for x in items.flat):
|
||||||
|
meta = dtype.metadata or {}
|
||||||
|
if meta.get("h5py_encoding") == "utf-8":
|
||||||
|
enc = [x.encode("utf-8") if isinstance(x, str) else x for x in items.flat]
|
||||||
|
return np.array(enc, dtype=dtype).reshape(items.shape)
|
||||||
|
return np.asarray(value, dtype=dtype)
|
||||||
|
|
||||||
|
|
||||||
|
def _broadcast(arr, shape, fancy, chunk_elems):
|
||||||
|
"""h5py's broadcasting: numpy's rules against the selection's shape
|
||||||
|
(extra leading length-1 axes allowed) for slices and integers. For an
|
||||||
|
index list, the exact shape; a scalar only where h5py expands it to the
|
||||||
|
whole selection (a chunked dataset whose chunk holds at least as many
|
||||||
|
elements as the selection)."""
|
||||||
|
if arr.shape == shape:
|
||||||
|
return arr
|
||||||
|
if fancy:
|
||||||
|
size = int(np.prod(shape))
|
||||||
|
if arr.ndim == 0 and ((chunk_elems > 0 and size <= chunk_elems) or len(shape) == 1):
|
||||||
|
return np.broadcast_to(arr, shape)
|
||||||
|
raise TypeError("Broadcasting is not supported for complex selections")
|
||||||
|
if arr.ndim == 0:
|
||||||
|
return np.broadcast_to(arr, shape)
|
||||||
|
err = TypeError(f"Can't broadcast {arr.shape} -> {shape}")
|
||||||
|
src = arr.shape
|
||||||
|
while len(src) > len(shape) and src[0] == 1:
|
||||||
|
src = src[1:]
|
||||||
|
if len(src) > len(shape):
|
||||||
|
raise err
|
||||||
|
try:
|
||||||
|
return np.broadcast_to(arr.reshape(src), shape)
|
||||||
|
except ValueError:
|
||||||
|
raise err from None
|
||||||
|
|
||||||
|
|
||||||
|
def dataset_values(value, dtype, category, shape, fancy, chunk_elems):
|
||||||
|
"""The bytes to write for `value` under a selection of `shape`, as a
|
||||||
|
C-ordered array of the dataset's dtype."""
|
||||||
|
if isinstance(value, np.ndarray):
|
||||||
|
arr = _convert_array(value, dtype, category)
|
||||||
|
else:
|
||||||
|
arr = _convert_other(value, dtype, category)
|
||||||
|
arr = _broadcast(arr, tuple(shape), fancy, chunk_elems)
|
||||||
|
return np.ascontiguousarray(arr, dtype=dtype).tobytes()
|
||||||
|
|
||||||
|
|
||||||
|
def attr_value(value, dtype=None, shape=None):
|
||||||
|
"""(kind, dtype string, shape, bytes) for an attribute value, h5py's
|
||||||
|
`attrs[name] = value` / `attrs.create(name, data, shape, dtype)`:
|
||||||
|
|
||||||
|
- "str": `str` data (h5py would store a variable-length string; this
|
||||||
|
stores a fixed-length UTF-8 string, which clawhdf5 can write);
|
||||||
|
the dtype string is the byte length of the longest element;
|
||||||
|
- "raw": a numeric, bool or bytes array, as numpy lays it out.
|
||||||
|
"""
|
||||||
|
if dtype is not None:
|
||||||
|
arr = np.asarray(value, dtype=dtype, order="C")
|
||||||
|
else:
|
||||||
|
arr = np.asarray(value, order="C")
|
||||||
|
if shape is not None:
|
||||||
|
arr = arr.reshape(shape)
|
||||||
|
kind = arr.dtype.kind
|
||||||
|
if kind == "O":
|
||||||
|
if arr.size and all(isinstance(x, str) for x in arr.flat):
|
||||||
|
kind = "U"
|
||||||
|
elif arr.size and all(isinstance(x, bytes) for x in arr.flat):
|
||||||
|
arr = arr.astype(bytes)
|
||||||
|
kind = "S"
|
||||||
|
else:
|
||||||
|
raise TypeError(
|
||||||
|
f"clawhdf5 cannot write an attribute of Python objects ({value!r:.60})"
|
||||||
|
)
|
||||||
|
if kind == "U":
|
||||||
|
enc = [str(x).encode("utf-8") for x in arr.flat]
|
||||||
|
size = max([len(b) for b in enc] + [1])
|
||||||
|
data = np.array(enc, dtype=f"S{size}").reshape(arr.shape)
|
||||||
|
return ("str", str(size), arr.shape, data.tobytes())
|
||||||
|
if kind in "biufcS":
|
||||||
|
return ("raw", arr.dtype.str, arr.shape, np.ascontiguousarray(arr).tobytes())
|
||||||
|
raise NotImplementedError(
|
||||||
|
f"clawhdf5 cannot write an attribute of dtype {arr.dtype} in place"
|
||||||
|
)
|
||||||
@@ -5,7 +5,7 @@ use std::path::PathBuf;
|
|||||||
use std::sync::{Arc, Mutex};
|
use std::sync::{Arc, Mutex};
|
||||||
use std::time::Duration;
|
use std::time::Duration;
|
||||||
|
|
||||||
use pyo3::exceptions::PyValueError;
|
use pyo3::exceptions::{PyNotImplementedError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::{PyDict, PyList};
|
use pyo3::types::{PyDict, PyList};
|
||||||
|
|
||||||
@@ -95,6 +95,18 @@ impl PyFile {
|
|||||||
}
|
}
|
||||||
match mode {
|
match mode {
|
||||||
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
|
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
|
||||||
|
"r+" => Ok(Self::from_handle(
|
||||||
|
Handle::open_editable(py, path)?,
|
||||||
|
filename,
|
||||||
|
)),
|
||||||
|
"a" if std::path::Path::new(path).exists() => Ok(Self::from_handle(
|
||||||
|
Handle::open_editable(py, path)?,
|
||||||
|
filename,
|
||||||
|
)),
|
||||||
|
"a" => Err(PyNotImplementedError::new_err(format!(
|
||||||
|
"mode 'a' on {path}, which does not exist: clawhdf5 can only edit an existing \
|
||||||
|
file in place; create a new one with mode 'w'"
|
||||||
|
))),
|
||||||
"w" => Ok(Self {
|
"w" => Ok(Self {
|
||||||
filename,
|
filename,
|
||||||
inner: Some(FileInner::Write(WriteState {
|
inner: Some(FileInner::Write(WriteState {
|
||||||
@@ -105,7 +117,7 @@ impl PyFile {
|
|||||||
})),
|
})),
|
||||||
}),
|
}),
|
||||||
other => Err(PyValueError::new_err(format!(
|
other => Err(PyValueError::new_err(format!(
|
||||||
"unsupported mode '{other}'; expected 'r' or 'w'"
|
"unsupported mode '{other}'; expected 'r', 'r+', 'a' or 'w'"
|
||||||
))),
|
))),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -226,7 +238,10 @@ impl PyFile {
|
|||||||
PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed")
|
PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed")
|
||||||
})?;
|
})?;
|
||||||
match inner {
|
match inner {
|
||||||
FileInner::Read(_) => Ok(()),
|
FileInner::Read(root) => {
|
||||||
|
root.handle.close();
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
FileInner::Write(state) => finalize_write(state),
|
FileInner::Write(state) => finalize_write(state),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -289,6 +304,30 @@ impl PyFile {
|
|||||||
"/"
|
"/"
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// `'r'` for a file opened read-only (a local file or a URL), `'r+'`
|
||||||
|
/// for one open for editing or writing, as h5py reports it.
|
||||||
|
#[getter]
|
||||||
|
fn mode(&self) -> PyResult<&'static str> {
|
||||||
|
match &self.inner {
|
||||||
|
Some(FileInner::Read(root)) if !root.handle.is_writable() => Ok("r"),
|
||||||
|
Some(_) => Ok("r+"),
|
||||||
|
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||||
|
"file is closed",
|
||||||
|
)),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Nothing to do: every edit is written and synced when it is made, and
|
||||||
|
/// a file opened with 'w' is written on `close()`.
|
||||||
|
fn flush(&self) {}
|
||||||
|
|
||||||
|
/// Deleting objects is not supported (h5py's `del f[name]`).
|
||||||
|
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||||
|
Err(PyNotImplementedError::new_err(format!(
|
||||||
|
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
|
||||||
|
)))
|
||||||
|
}
|
||||||
|
|
||||||
/// The path (or URL) the file was opened with.
|
/// The path (or URL) the file was opened with.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn filename(&self) -> &str {
|
fn filename(&self) -> &str {
|
||||||
@@ -389,6 +428,13 @@ impl PyFile {
|
|||||||
fn write_state_mut(&mut self) -> PyResult<&mut WriteState> {
|
fn write_state_mut(&mut self) -> PyResult<&mut WriteState> {
|
||||||
match &mut self.inner {
|
match &mut self.inner {
|
||||||
Some(FileInner::Write(s)) => Ok(s),
|
Some(FileInner::Write(s)) => Ok(s),
|
||||||
|
Some(FileInner::Read(root)) if root.handle.is_writable() => {
|
||||||
|
Err(PyNotImplementedError::new_err(
|
||||||
|
"creating datasets or groups in an existing file is not supported by \
|
||||||
|
clawhdf5's in-place editor (mode 'r+' changes values, shapes and \
|
||||||
|
attributes)",
|
||||||
|
))
|
||||||
|
}
|
||||||
Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||||
"cannot write to a file opened for reading",
|
"cannot write to a file opened for reading",
|
||||||
)),
|
)),
|
||||||
|
|||||||
@@ -3,7 +3,7 @@
|
|||||||
use std::collections::HashMap;
|
use std::collections::HashMap;
|
||||||
use std::sync::{Arc, Mutex, OnceLock};
|
use std::sync::{Arc, Mutex, OnceLock};
|
||||||
|
|
||||||
use pyo3::exceptions::{PyIOError, PyKeyError, PyOSError, PyValueError};
|
use pyo3::exceptions::{PyIOError, PyKeyError, PyNotImplementedError, PyOSError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::PyList;
|
use pyo3::types::PyList;
|
||||||
|
|
||||||
@@ -301,12 +301,23 @@ impl PyGroup {
|
|||||||
state.lock().unwrap().datasets.push(spec);
|
state.lock().unwrap().datasets.push(spec);
|
||||||
Ok(())
|
Ok(())
|
||||||
}
|
}
|
||||||
GroupInner::Read { .. } => Err(PyIOError::new_err(
|
GroupInner::Read(g) if g.handle.is_writable() => Err(PyNotImplementedError::new_err(
|
||||||
|
"creating datasets or groups in an existing file is not supported by \
|
||||||
|
clawhdf5's in-place editor (mode 'r+' changes values, shapes and attributes)",
|
||||||
|
)),
|
||||||
|
GroupInner::Read(_) => Err(PyIOError::new_err(
|
||||||
"cannot create datasets on a read-only group",
|
"cannot create datasets on a read-only group",
|
||||||
)),
|
)),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Deleting objects is not supported (h5py's `del group[name]`).
|
||||||
|
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||||
|
Err(PyNotImplementedError::new_err(format!(
|
||||||
|
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
|
||||||
|
)))
|
||||||
|
}
|
||||||
|
|
||||||
/// Attribute access.
|
/// Attribute access.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||||
|
|||||||
@@ -2,25 +2,34 @@
|
|||||||
//!
|
//!
|
||||||
//! Every read goes through [`Handle::with`], which releases the GIL and
|
//! Every read goes through [`Handle::with`], which releases the GIL and
|
||||||
//! parses through `File::storage()`, so the same code serves a local file
|
//! parses through `File::storage()`, so the same code serves a local file
|
||||||
//! (memory-mapped) and a remote one (`clawhdf5-remote`: range requests
|
//! (memory-mapped), a remote one (`clawhdf5-remote`: range requests through
|
||||||
//! through a block cache, so a network read never holds the GIL).
|
//! a block cache, so a network read never holds the GIL) and a file open
|
||||||
|
//! for editing.
|
||||||
|
//!
|
||||||
|
//! A file opened with `'r+'` also holds a [`FileEditor`]. An edit takes the
|
||||||
|
//! file's write lock, so no read runs while the file changes underneath it,
|
||||||
|
//! and reopens the file afterwards: reads after an edit see the new bytes
|
||||||
|
//! (a grown file, a new dataspace), never a stale mapping or chunk cache.
|
||||||
|
//! Objects that cache something an edit can change compare
|
||||||
|
//! [`Handle::generation`] with the value they cached it at.
|
||||||
//!
|
//!
|
||||||
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
|
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
|
||||||
//! with the GIL released, and code that holds it never touches Python.
|
//! with the GIL released, and code that holds it never touches Python.
|
||||||
|
|
||||||
use std::path::PathBuf;
|
use std::path::PathBuf;
|
||||||
use std::sync::{Arc, PoisonError, RwLock};
|
use std::sync::atomic::{AtomicU64, Ordering};
|
||||||
|
use std::sync::{Arc, Mutex, PoisonError, RwLock};
|
||||||
|
|
||||||
use clawhdf5_rs::File;
|
use clawhdf5_rs::{File, FileEditor};
|
||||||
use pyo3::exceptions::PyOSError;
|
use pyo3::exceptions::PyOSError;
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
|
|
||||||
use crate::to_py_err;
|
use crate::{panic_text, to_py_err};
|
||||||
|
|
||||||
/// Where the file's bytes come from.
|
/// Where the file's bytes come from.
|
||||||
pub(crate) enum Source {
|
pub(crate) enum Source {
|
||||||
/// A local path (memory-mapped).
|
/// A local path (memory-mapped).
|
||||||
Local(#[allow(dead_code)] PathBuf),
|
Local(PathBuf),
|
||||||
/// A URL, read through `clawhdf5-remote`'s block cache.
|
/// A URL, read through `clawhdf5-remote`'s block cache.
|
||||||
Remote {
|
Remote {
|
||||||
url: String,
|
url: String,
|
||||||
@@ -29,8 +38,14 @@ pub(crate) enum Source {
|
|||||||
}
|
}
|
||||||
|
|
||||||
pub(crate) struct Handle {
|
pub(crate) struct Handle {
|
||||||
|
/// The file as last opened; `None` if reopening it after an edit failed
|
||||||
|
/// (every read is then an error rather than a read of stale bytes).
|
||||||
file: RwLock<Option<File>>,
|
file: RwLock<Option<File>>,
|
||||||
|
/// For `'r+'`: the editor, until the file is closed.
|
||||||
|
editor: Option<Mutex<Option<FileEditor>>>,
|
||||||
source: Source,
|
source: Source,
|
||||||
|
/// Bumped by every edit.
|
||||||
|
generation: AtomicU64,
|
||||||
pub offset_size: u8,
|
pub offset_size: u8,
|
||||||
pub length_size: u8,
|
pub length_size: u8,
|
||||||
pub root: u64,
|
pub root: u64,
|
||||||
@@ -41,13 +56,15 @@ fn closed_after_failed_reopen() -> PyErr {
|
|||||||
}
|
}
|
||||||
|
|
||||||
impl Handle {
|
impl Handle {
|
||||||
fn new(file: File, source: Source) -> Arc<Self> {
|
fn new(file: File, source: Source, editor: Option<FileEditor>) -> Arc<Self> {
|
||||||
let sb = file.superblock();
|
let sb = file.superblock();
|
||||||
let (offset_size, length_size, root) =
|
let (offset_size, length_size, root) =
|
||||||
(sb.offset_size, sb.length_size, sb.root_group_address);
|
(sb.offset_size, sb.length_size, sb.root_group_address);
|
||||||
Arc::new(Self {
|
Arc::new(Self {
|
||||||
file: RwLock::new(Some(file)),
|
file: RwLock::new(Some(file)),
|
||||||
|
editor: editor.map(|e| Mutex::new(Some(e))),
|
||||||
source,
|
source,
|
||||||
|
generation: AtomicU64::new(0),
|
||||||
offset_size,
|
offset_size,
|
||||||
length_size,
|
length_size,
|
||||||
root,
|
root,
|
||||||
@@ -57,7 +74,25 @@ impl Handle {
|
|||||||
/// A local file, read-only.
|
/// A local file, read-only.
|
||||||
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
|
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
|
||||||
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
|
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
|
||||||
Ok(Self::new(file, Source::Local(PathBuf::from(path))))
|
Ok(Self::new(file, Source::Local(PathBuf::from(path)), None))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A local file, open for in-place editing (`'r+'`): the editor takes
|
||||||
|
/// the file's exclusive lock and checks that it can edit the file, then
|
||||||
|
/// the file is opened for reading.
|
||||||
|
pub(crate) fn open_editable(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
|
||||||
|
let (file, editor) = py.detach(|| {
|
||||||
|
crate::no_panic(|| {
|
||||||
|
let editor = FileEditor::open(path).map_err(to_py_err)?;
|
||||||
|
let file = File::open(path).map_err(to_py_err)?;
|
||||||
|
Ok((file, editor))
|
||||||
|
})
|
||||||
|
})?;
|
||||||
|
Ok(Self::new(
|
||||||
|
file,
|
||||||
|
Source::Local(PathBuf::from(path)),
|
||||||
|
Some(editor),
|
||||||
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// A remote file (`http(s)://`, `s3://`, ...).
|
/// A remote file (`http(s)://`, `s3://`, ...).
|
||||||
@@ -79,6 +114,7 @@ impl Handle {
|
|||||||
url: url.to_string(),
|
url: url.to_string(),
|
||||||
storage,
|
storage,
|
||||||
},
|
},
|
||||||
|
None,
|
||||||
))
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -102,6 +138,17 @@ impl Handle {
|
|||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Edits so far: objects that cache something an edit can change (a
|
||||||
|
/// dataset's shape, an object's attributes) re-read it when this moved.
|
||||||
|
pub(crate) fn generation(&self) -> u64 {
|
||||||
|
self.generation.load(Ordering::Acquire)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Whether the file was opened for editing (`'r+'`), even if closed since.
|
||||||
|
pub(crate) fn is_writable(&self) -> bool {
|
||||||
|
self.editor.is_some()
|
||||||
|
}
|
||||||
|
|
||||||
/// The remote file's block cache.
|
/// The remote file's block cache.
|
||||||
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
|
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
|
||||||
match &self.source {
|
match &self.source {
|
||||||
@@ -117,6 +164,61 @@ impl Handle {
|
|||||||
Source::Local(_) => None,
|
Source::Local(_) => None,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Release the editor, and with it the file's lock. Objects still
|
||||||
|
/// open keep reading the file as it was last written; an edit through
|
||||||
|
/// them is an error.
|
||||||
|
pub(crate) fn close(&self) {
|
||||||
|
if let Some(ed) = &self.editor {
|
||||||
|
ed.lock().unwrap_or_else(PoisonError::into_inner).take();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Apply one edit with the GIL released. No read runs while it writes,
|
||||||
|
/// and the file is reopened afterwards — also after a failed edit, since
|
||||||
|
/// a commit that failed part-way may have changed the file.
|
||||||
|
pub(crate) fn edit<R: Send>(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
f: impl FnOnce(&mut FileEditor) -> Result<R, clawhdf5_rs::Error> + Send,
|
||||||
|
) -> PyResult<R> {
|
||||||
|
let Some(editor) = &self.editor else {
|
||||||
|
return Err(PyOSError::new_err(match self.source {
|
||||||
|
Source::Remote { .. } => "remote files are read-only",
|
||||||
|
Source::Local(_) => {
|
||||||
|
"the file is open read-only; open it with mode 'r+' to change it"
|
||||||
|
}
|
||||||
|
}));
|
||||||
|
};
|
||||||
|
let Source::Local(path) = &self.source else {
|
||||||
|
return Err(PyOSError::new_err("remote files are read-only"));
|
||||||
|
};
|
||||||
|
py.detach(|| {
|
||||||
|
let mut ed = editor.lock().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
let ed = ed
|
||||||
|
.as_mut()
|
||||||
|
.ok_or_else(|| PyOSError::new_err("the file is closed"))?;
|
||||||
|
let mut file = self.file.write().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| f(ed)));
|
||||||
|
// Drop the old mapping and its chunk cache before reopening.
|
||||||
|
*file = None;
|
||||||
|
let reopened = std::panic::catch_unwind(|| File::open(path));
|
||||||
|
self.generation.fetch_add(1, Ordering::AcqRel);
|
||||||
|
match reopened {
|
||||||
|
Ok(Ok(f)) => *file = Some(f),
|
||||||
|
Ok(Err(e)) => return Err(to_py_err(e)),
|
||||||
|
Err(_) => return Err(closed_after_failed_reopen()),
|
||||||
|
}
|
||||||
|
drop(file);
|
||||||
|
match result {
|
||||||
|
Ok(r) => r.map_err(to_py_err),
|
||||||
|
Err(p) => Err(crate::InternalError::new_err(format!(
|
||||||
|
"clawhdf5 internal error (please report it): {}",
|
||||||
|
panic_text(&*p)
|
||||||
|
))),
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// A `clawhdf5_remote::Error` as a Python exception: the network side
|
/// A `clawhdf5_remote::Error` as a Python exception: the network side
|
||||||
|
|||||||
@@ -7,11 +7,19 @@
|
|||||||
//!
|
//!
|
||||||
//! with clawhdf5.File('data.h5', 'r') as f:
|
//! with clawhdf5.File('data.h5', 'r') as f:
|
||||||
//! data = f['dataset_name'][:]
|
//! data = f['dataset_name'][:]
|
||||||
|
//!
|
||||||
|
//! with clawhdf5.File('http://host/data.h5') as f: # range requests
|
||||||
|
//! block = f['dataset_name'][10:20]
|
||||||
|
//!
|
||||||
|
//! with clawhdf5.File('data.h5', 'r+') as f: # in-place edits
|
||||||
|
//! f['dataset_name'][0] = 1.5
|
||||||
|
//! f.attrs['note'] = 'edited'
|
||||||
//! ```
|
//! ```
|
||||||
|
|
||||||
mod attrs;
|
mod attrs;
|
||||||
mod convert;
|
mod convert;
|
||||||
mod dataset;
|
mod dataset;
|
||||||
|
mod edit;
|
||||||
mod file;
|
mod file;
|
||||||
mod group;
|
mod group;
|
||||||
mod handle;
|
mod handle;
|
||||||
|
|||||||
@@ -0,0 +1,742 @@
|
|||||||
|
"""In-place editing: clawhdf5.File(path, 'r+') against h5py.
|
||||||
|
|
||||||
|
Every edit is applied twice, to two copies of the same file: once through
|
||||||
|
h5py (libhdf5) and once through clawhdf5 (FileEditor). After every edit both
|
||||||
|
files are read back with h5py and must hold the same shapes, values and
|
||||||
|
attributes; clawhdf5's own view must agree; when h5py refuses an edit,
|
||||||
|
clawhdf5 must refuse it too and leave its file as it was. Files are written
|
||||||
|
by h5py (libver earliest and latest, so every chunk index kind) and by
|
||||||
|
clawhdf5; `h5dump` must read every result."""
|
||||||
|
|
||||||
|
import io
|
||||||
|
import os
|
||||||
|
import shutil
|
||||||
|
import subprocess
|
||||||
|
import threading
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
import clawhdf5
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Files
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
ENUM = {"RED": 0, "GREEN": 1, "BLUE": 7}
|
||||||
|
|
||||||
|
|
||||||
|
def _h5py_file(h5py, path, libver):
|
||||||
|
rng = np.random.default_rng(1)
|
||||||
|
with h5py.File(path, "w", libver=libver) as f:
|
||||||
|
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
|
||||||
|
f.create_dataset("be_i2", data=np.arange(24, dtype=">i2").reshape(4, 6))
|
||||||
|
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
|
||||||
|
f.create_dataset("f2", data=rng.standard_normal(12).astype("<f2"))
|
||||||
|
f.create_dataset("f4_2d", data=rng.standard_normal((8, 9)).astype("<f4"))
|
||||||
|
f.create_dataset("f8_3d", data=rng.standard_normal((4, 5, 6)))
|
||||||
|
f.create_dataset("u8", data=np.arange(10, dtype="<u8"))
|
||||||
|
f.create_dataset("c8", data=(np.arange(6) + 1j * np.arange(6)).astype("<c8"))
|
||||||
|
f.create_dataset("bool", data=np.array([True, False, True, True]))
|
||||||
|
f.create_dataset("enum", data=np.array([0, 1, 7, 0], dtype="i1"),
|
||||||
|
dtype=h5py.enum_dtype(ENUM, basetype="i1"))
|
||||||
|
f.create_dataset("s5", data=np.array([b"ab", b"cdefg", b""], dtype="S5"))
|
||||||
|
cmp_dt = np.dtype([("id", "<i4"), ("x", "<f8"), ("tag", "S3")])
|
||||||
|
f.create_dataset("cmp", data=np.array([(i, i / 2, b"t%d" % i) for i in range(5)], dtype=cmp_dt))
|
||||||
|
f.create_dataset("scalar", data=np.float64(3.5))
|
||||||
|
# Chunked: fixed maxshape (v4 fixed array under latest), one
|
||||||
|
# unlimited dimension (extensible array), two (v2 B-tree), one chunk.
|
||||||
|
f.create_dataset("chunk_fixed", data=np.arange(100, dtype="<i8").reshape(10, 10), chunks=(3, 4))
|
||||||
|
f.create_dataset("chunk_ext", data=rng.standard_normal((12, 7)), chunks=(5, 7), maxshape=(None, 7))
|
||||||
|
f.create_dataset("chunk_bt2", data=np.arange(30, dtype="<i4").reshape(5, 6), chunks=(2, 2),
|
||||||
|
maxshape=(None, None))
|
||||||
|
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20), chunks=(6, 6),
|
||||||
|
compression="gzip", maxshape=(40, 40))
|
||||||
|
f.create_dataset("chunk_single", data=np.arange(12, dtype="<u2").reshape(3, 4), chunks=(3, 4),
|
||||||
|
maxshape=(3, 4))
|
||||||
|
f.create_dataset("chunk_fill", shape=(8,), dtype="<i4", chunks=(3,), maxshape=(20,), fillvalue=-1)
|
||||||
|
f.create_dataset("vlen", data=["a", "bb"], dtype=h5py.string_dtype())
|
||||||
|
# Compact layout (low-level API).
|
||||||
|
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
|
||||||
|
dcpl.set_layout(h5py.h5d.COMPACT)
|
||||||
|
space = h5py.h5s.create_simple((7,))
|
||||||
|
dsid = h5py.h5d.create(f.id, b"compact", h5py.h5t.STD_I32LE, space, dcpl=dcpl)
|
||||||
|
dsid.write(h5py.h5s.ALL, h5py.h5s.ALL, np.arange(7, dtype="<i4"))
|
||||||
|
g = f.create_group("grp")
|
||||||
|
g.create_dataset("leaf", data=np.arange(5.0))
|
||||||
|
g.attrs["units"] = "m"
|
||||||
|
f.attrs["version"] = np.int32(1)
|
||||||
|
|
||||||
|
|
||||||
|
def _clawhdf5_file(path):
|
||||||
|
with clawhdf5.File(str(path), "w") as f:
|
||||||
|
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
|
||||||
|
f.create_dataset("f8", data=np.linspace(0, 1, 30).reshape(5, 6))
|
||||||
|
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20),
|
||||||
|
chunks=[6, 6], compression="gzip")
|
||||||
|
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
|
||||||
|
g = f.create_group("grp")
|
||||||
|
g.create_dataset("leaf", data=np.arange(5.0))
|
||||||
|
g.attrs["units"] = "m"
|
||||||
|
f.attrs["version"] = 1
|
||||||
|
|
||||||
|
|
||||||
|
# h5py's libver: "earliest" (v1 B-tree chunk indexes), "v114" (the 1.10+
|
||||||
|
# indexes: fixed and extensible arrays, v2 B-trees, single chunk) and
|
||||||
|
# "latest" (HDF5 2.0's newest format, which h5dump 1.14 cannot read).
|
||||||
|
SOURCES = ["h5py-earliest", "h5py-v114", "h5py-latest", "clawhdf5"]
|
||||||
|
|
||||||
|
|
||||||
|
def _make(h5py, tmp_path, source):
|
||||||
|
base = tmp_path / f"base-{source}.h5"
|
||||||
|
if source == "clawhdf5":
|
||||||
|
_clawhdf5_file(base)
|
||||||
|
else:
|
||||||
|
_h5py_file(h5py, str(base), source.split("-")[1])
|
||||||
|
theirs = tmp_path / f"theirs-{source}.h5"
|
||||||
|
ours = tmp_path / f"ours-{source}.h5"
|
||||||
|
shutil.copy(base, theirs)
|
||||||
|
shutil.copy(base, ours)
|
||||||
|
return str(theirs), str(ours), str(base)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Comparing files through h5py
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _norm_attr(v):
|
||||||
|
"""Attribute values comparable across the two writers: clawhdf5 stores
|
||||||
|
`str` as fixed-length UTF-8 (h5py reads bytes), h5py as variable-length
|
||||||
|
(h5py reads str)."""
|
||||||
|
if isinstance(v, bytes):
|
||||||
|
return ("str", v.decode("utf-8"))
|
||||||
|
if isinstance(v, str):
|
||||||
|
return ("str", v)
|
||||||
|
arr = np.asarray(v)
|
||||||
|
if arr.dtype.kind in "SO":
|
||||||
|
return ("strs", [x.decode() if isinstance(x, bytes) else x for x in arr.ravel().tolist()], arr.shape)
|
||||||
|
return (arr.dtype.str, arr.shape, arr.tobytes())
|
||||||
|
|
||||||
|
|
||||||
|
def snapshot(h5py, path):
|
||||||
|
"""What h5py sees in the file: every dataset's shape, dtype, bytes and
|
||||||
|
attributes (read without locking: clawhdf5 may hold the file open)."""
|
||||||
|
out = {}
|
||||||
|
with h5py.File(path, "r", locking=False) as f:
|
||||||
|
def visit(name, obj):
|
||||||
|
attrs = {k: _norm_attr(obj.attrs[k]) for k in obj.attrs}
|
||||||
|
if isinstance(obj, h5py.Dataset):
|
||||||
|
if obj.dtype.kind == "O":
|
||||||
|
data = [x for x in obj[...].ravel().tolist()]
|
||||||
|
else:
|
||||||
|
data = obj[()].tobytes() if obj.shape is not None else None
|
||||||
|
out[name] = (obj.shape, obj.dtype.str, obj.maxshape, data, attrs)
|
||||||
|
else:
|
||||||
|
out[name] = ("group", attrs)
|
||||||
|
visit("/", f)
|
||||||
|
f.visititems(visit)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def assert_same_files(h5py, theirs, ours, what):
|
||||||
|
a, b = snapshot(h5py, theirs), snapshot(h5py, ours)
|
||||||
|
assert a.keys() == b.keys(), what
|
||||||
|
for k in a:
|
||||||
|
assert a[k] == b[k], f"{what}: {k} differs\n h5py: {a[k]}\n clawhdf5: {b[k]}"
|
||||||
|
|
||||||
|
|
||||||
|
def assert_ours_reads_like_h5py(h5py, f, path, what):
|
||||||
|
"""clawhdf5's own view of the file it is editing matches h5py's."""
|
||||||
|
with h5py.File(path, "r", locking=False) as t:
|
||||||
|
for name in ["i4", "chunk_ext", "chunk_bt2", "chunk_gzip", "f8_3d", "cmp", "bool", "enum", "scalar"]:
|
||||||
|
if name not in t:
|
||||||
|
continue
|
||||||
|
o = f[name]
|
||||||
|
assert o.shape == t[name].shape, f"{what}: {name} shape"
|
||||||
|
assert o.maxshape == t[name].maxshape, f"{what}: {name} maxshape"
|
||||||
|
np.testing.assert_array_equal(o[()], t[name][()], err_msg=f"{what}: {name}")
|
||||||
|
for obj in ["/", "grp"]:
|
||||||
|
assert sorted(f[obj].attrs.keys()) == sorted(t[obj].attrs.keys()), what
|
||||||
|
for k in t[obj].attrs:
|
||||||
|
assert _norm_attr(f[obj].attrs[k]) == _norm_attr(t[obj].attrs[k]), f"{what}: {obj}.attrs[{k}]"
|
||||||
|
|
||||||
|
|
||||||
|
def h5dump_reads(path, base=None):
|
||||||
|
"""h5dump (libhdf5 1.14) reads every object and value of `path` — when it
|
||||||
|
reads the unedited `base` (it cannot read HDF5 2.0's newest format)."""
|
||||||
|
exe = shutil.which("h5dump")
|
||||||
|
if exe is None:
|
||||||
|
if os.environ.get("CLAWHDF5_REQUIRE_INTEROP") == "1":
|
||||||
|
pytest.fail("h5dump is required (CLAWHDF5_REQUIRE_INTEROP=1)")
|
||||||
|
return
|
||||||
|
h5rs = os.environ.get("CLAWHDF5_H5RS")
|
||||||
|
if h5rs:
|
||||||
|
# clawhdf5's structural and checksum validator (scripts/ci-test.sh
|
||||||
|
# points this at the h5rs it built).
|
||||||
|
r = subprocess.run([h5rs, "check", path], capture_output=True, text=True)
|
||||||
|
assert r.returncode == 0, (r.stdout + r.stderr)[-2000:]
|
||||||
|
if base is not None and subprocess.run([exe, "-H", base], capture_output=True).returncode != 0:
|
||||||
|
return
|
||||||
|
r = subprocess.run([exe, path], capture_output=True, text=True)
|
||||||
|
assert r.returncode == 0, r.stderr[-2000:]
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Applying one edit both ways
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _native_conversion(h5py, value, ds_dtype):
|
||||||
|
"""`value` as libhdf5 converts it to `ds_dtype` in native byte order.
|
||||||
|
|
||||||
|
libhdf5 2.0 (h5py 3.16) converts numbers differently when either side is
|
||||||
|
not in native byte order (its "soft" conversions): a float in (-1, 0)
|
||||||
|
becomes the integer type's minimum instead of 0, and an unsigned integer
|
||||||
|
too large for the signed type of the same size wraps instead of
|
||||||
|
saturating. clawhdf5 applies the native-order results to every byte
|
||||||
|
order, so the reference is h5py converting into a native dataset of the
|
||||||
|
same kind; the result then reaches the real dataset by a plain byte
|
||||||
|
swap."""
|
||||||
|
if not isinstance(value, np.ndarray):
|
||||||
|
return value
|
||||||
|
if value.dtype.kind not in "biuf" or ds_dtype.kind not in "biuf":
|
||||||
|
return value
|
||||||
|
if value.dtype.isnative and ds_dtype.isnative:
|
||||||
|
return value
|
||||||
|
with h5py.File(io.BytesIO(), "w") as tmp:
|
||||||
|
d = tmp.create_dataset("t", shape=value.shape, dtype=ds_dtype.newbyteorder("="))
|
||||||
|
d[...] = value.astype(value.dtype.newbyteorder("="))
|
||||||
|
return np.asarray(d[()])
|
||||||
|
|
||||||
|
|
||||||
|
def _apply(f, op, h5py=None):
|
||||||
|
"""Apply `op` to `f`; with `h5py`, `f` is an h5py file and a numpy array
|
||||||
|
value is first converted as libhdf5 converts in native byte order (see
|
||||||
|
`_native_conversion`)."""
|
||||||
|
kind = op[0]
|
||||||
|
if kind == "set":
|
||||||
|
_, name, key, value = op
|
||||||
|
if h5py is not None:
|
||||||
|
value = _native_conversion(h5py, value, f[name].dtype)
|
||||||
|
f[name][key] = value
|
||||||
|
elif kind == "resize":
|
||||||
|
_, name, size, axis = op
|
||||||
|
if axis is None:
|
||||||
|
f[name].resize(size)
|
||||||
|
else:
|
||||||
|
f[name].resize(size, axis=axis)
|
||||||
|
elif kind == "attr":
|
||||||
|
_, obj, name, value = op
|
||||||
|
f[obj].attrs[name] = value
|
||||||
|
else:
|
||||||
|
raise AssertionError(op)
|
||||||
|
|
||||||
|
|
||||||
|
def edit_both(h5py, theirs, ours_path, ours, op):
|
||||||
|
"""Apply `op` with h5py and with clawhdf5 (`ours`, open 'r+'); the two
|
||||||
|
files must then read the same through h5py. If h5py refuses, clawhdf5
|
||||||
|
must refuse and its file must be unchanged. Returns h5py's error."""
|
||||||
|
before = snapshot(h5py, ours_path)
|
||||||
|
try:
|
||||||
|
with h5py.File(theirs, "r+") as t:
|
||||||
|
_apply(t, op, h5py)
|
||||||
|
except Exception as e: # noqa: BLE001 - h5py refuses: so must we
|
||||||
|
try:
|
||||||
|
_apply(ours, op)
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
pass
|
||||||
|
else:
|
||||||
|
pytest.fail(f"{op!r:.300}: h5py refused ({type(e).__name__}: {e}), clawhdf5 did not")
|
||||||
|
assert snapshot(h5py, ours_path) == before, f"{op!r}: clawhdf5 changed the file while failing"
|
||||||
|
return e
|
||||||
|
_apply(ours, op)
|
||||||
|
assert_same_files(h5py, theirs, ours_path, repr(op)[:200])
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Tests
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("source", SOURCES)
|
||||||
|
def test_edit_sequence_matches_h5py(h5py, tmp_path, source):
|
||||||
|
theirs, ours_path, base = _make(h5py, tmp_path, source)
|
||||||
|
ops = [
|
||||||
|
("set", "i4", 0, 99),
|
||||||
|
("set", "i4", (slice(1, 5, 2), slice(None, None, 3)), np.array([[1.5, -2.5, 1e12, -1e12]])),
|
||||||
|
("set", "i4", (slice(None), 2), np.arange(6, dtype="<i8") * 1000),
|
||||||
|
("set", "i4", ([0, 2, 5], slice(4, 6)), np.array([[1, 2], [3, 4], [5, 6]])),
|
||||||
|
("set", "i4", (Ellipsis, -1), np.int8(-7)),
|
||||||
|
("set", "i4", (3, 3), 12345.9),
|
||||||
|
("set", "u1", slice(2, 8), np.array([-5, 0, 300, 255, 256, 1], dtype="<i4")),
|
||||||
|
("set", "u1", slice(0, 3), [1, 2, 3]),
|
||||||
|
("set", "chunk_gzip", (slice(0, 20, 7), slice(3, 17)), 42.25),
|
||||||
|
("set", "chunk_gzip", (slice(5, 11), slice(5, 11)), np.ones((6, 6), dtype="<f8") * np.pi),
|
||||||
|
("set", "grp/leaf", slice(None), np.array([1, 2, 3, 4, 5], dtype="<u2")),
|
||||||
|
("attr", "/", "version", np.int32(2)),
|
||||||
|
("attr", "/", "count", 7),
|
||||||
|
("attr", "grp", "scale", np.array([0.5, 0.25], dtype="<f4")),
|
||||||
|
("attr", "grp", "matrix", np.arange(6, dtype=">i8").reshape(2, 3)),
|
||||||
|
("attr", "i4", "flag", np.bool_(True)),
|
||||||
|
("attr", "i4", "z", np.complex64(1 - 2j)),
|
||||||
|
("attr", "i4", "raw", np.bytes_(b"abc")),
|
||||||
|
("set", "i4", slice(0, 2), np.zeros((3, 10))), # shape mismatch: refused by both
|
||||||
|
]
|
||||||
|
if source != "clawhdf5":
|
||||||
|
ops += [
|
||||||
|
("set", "be_i2", (slice(None), slice(1, 3)), np.array([70000, -70000], dtype="<i8")),
|
||||||
|
("set", "f2", slice(None, None, 4), np.array([1e6, -3.25, 0.1])),
|
||||||
|
("set", "f4_2d", (2, slice(None)), np.linspace(-1, 1, 9)),
|
||||||
|
("set", "f8_3d", (slice(1, 3), 2, slice(None, None, 2)), np.arange(3, dtype="<i2")),
|
||||||
|
("set", "u8", slice(None), np.array([-1, 0, 2**63, 1e30, -1e30, 5.5, 2, 3, 4, 5])),
|
||||||
|
("set", "c8", slice(1, 3), np.array([1 + 1j, 2 - 2j], dtype="<c16")),
|
||||||
|
("set", "c8", 0, np.float64(1.0)), # h5py: no conversion path
|
||||||
|
("set", "bool", slice(None), np.array([0, 3, 0, -1], dtype="<i4")),
|
||||||
|
("set", "bool", 1, np.array(True)),
|
||||||
|
("set", "enum", slice(0, 2), np.array([7, 1], dtype="<i4")),
|
||||||
|
("set", "s5", 0, np.bytes_(b"xyzuvw")),
|
||||||
|
("set", "s5", slice(1, 3), [b"q", b"rs"]),
|
||||||
|
("set", "s5", 2, np.array("uni")), # h5py: no conversion from 'U'
|
||||||
|
("set", "cmp", 2, np.array((9, 9.5, b"zz"), dtype=[("id", "<i4"), ("x", "<f8"), ("tag", "S3")])),
|
||||||
|
("set", "cmp", slice(3, 5), [(1, 0.5, b"a"), (2, 1.5, b"b")]),
|
||||||
|
("set", "scalar", (), 7.25),
|
||||||
|
("set", "scalar", Ellipsis, np.float32(-1.5)),
|
||||||
|
("set", "compact", slice(1, 6, 2), np.array([10, 20, 30])),
|
||||||
|
("set", "chunk_fixed", (slice(2, 9), slice(1, 10, 4)), np.arange(21).reshape(7, 3)),
|
||||||
|
("set", "chunk_single", (1, slice(None)), np.array([9, 8, 7, 6])),
|
||||||
|
("resize", "chunk_ext", (20, 7), None),
|
||||||
|
("set", "chunk_ext", slice(12, 20), np.full((8, 7), 2.5)),
|
||||||
|
("resize", "chunk_ext", 9, 0),
|
||||||
|
("resize", "chunk_ext", 16, 0),
|
||||||
|
("resize", "chunk_bt2", (9, 11), None),
|
||||||
|
("set", "chunk_bt2", (slice(4, 9), slice(5, 11)), np.arange(30).reshape(5, 6)),
|
||||||
|
("resize", "chunk_bt2", (3, 3), None),
|
||||||
|
("resize", "chunk_bt2", (7, 8), None),
|
||||||
|
("resize", "chunk_gzip", (40, 25), None),
|
||||||
|
("resize", "chunk_gzip", (41, 25), None), # beyond maxshape: refused
|
||||||
|
("resize", "chunk_fill", 15, None), # h5py: a size without axis must be a tuple
|
||||||
|
("resize", "chunk_fill", (15,), None),
|
||||||
|
("set", "chunk_fill", slice(10, 12), [1, 2]),
|
||||||
|
("resize", "chunk_fill", (4,), None),
|
||||||
|
("resize", "chunk_fill", (20,), None),
|
||||||
|
("resize", "i4", (7, 10), None), # not chunked: refused
|
||||||
|
]
|
||||||
|
with clawhdf5.File(ours_path, "r+") as ours:
|
||||||
|
assert ours.mode == "r+"
|
||||||
|
for i, op in enumerate(ops):
|
||||||
|
edit_both(h5py, theirs, ours_path, ours, op)
|
||||||
|
if i % 5 == 0:
|
||||||
|
assert_ours_reads_like_h5py(h5py, ours, ours_path, repr(op))
|
||||||
|
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
|
||||||
|
h5dump_reads(ours_path, base)
|
||||||
|
# Reopened, clawhdf5 reads what h5py reads.
|
||||||
|
with clawhdf5.File(ours_path, "r") as f:
|
||||||
|
assert f.mode == "r"
|
||||||
|
assert_ours_reads_like_h5py(h5py, f, ours_path, "reopened")
|
||||||
|
|
||||||
|
|
||||||
|
def _random_key(rng, shape):
|
||||||
|
key = []
|
||||||
|
for n in shape:
|
||||||
|
r = rng.random()
|
||||||
|
if n == 0 or r < 0.3:
|
||||||
|
start = int(rng.integers(0, n + 1)) if n else 0
|
||||||
|
stop = int(rng.integers(start, n + 1)) if n else 0
|
||||||
|
step = int(rng.integers(1, 4))
|
||||||
|
key.append(slice(start, stop, step))
|
||||||
|
elif r < 0.55:
|
||||||
|
key.append(int(rng.integers(-n, n)))
|
||||||
|
elif r < 0.7:
|
||||||
|
k = int(rng.integers(1, min(n, 4) + 1))
|
||||||
|
key.append(sorted(rng.choice(n, size=k, replace=False).tolist()))
|
||||||
|
else:
|
||||||
|
key.append(slice(None))
|
||||||
|
# Only one index list per key.
|
||||||
|
lists = [i for i, k in enumerate(key) if isinstance(k, list)]
|
||||||
|
for i in lists[1:]:
|
||||||
|
key[i] = slice(None)
|
||||||
|
return tuple(key)
|
||||||
|
|
||||||
|
|
||||||
|
def _selection_shape(key, shape):
|
||||||
|
out = []
|
||||||
|
fancy = False
|
||||||
|
for k, n in zip(key, shape):
|
||||||
|
if isinstance(k, slice):
|
||||||
|
out.append(len(range(*k.indices(n))))
|
||||||
|
elif isinstance(k, list):
|
||||||
|
out.append(len(k))
|
||||||
|
fancy = True
|
||||||
|
return tuple(out), fancy
|
||||||
|
|
||||||
|
|
||||||
|
def _random_value(rng, sel_shape, fancy, dtype):
|
||||||
|
r = rng.random()
|
||||||
|
if r < 0.2 or not sel_shape:
|
||||||
|
v = rng.standard_normal() * 1000
|
||||||
|
return np.float64(v) if rng.random() < 0.5 else int(v)
|
||||||
|
shape = list(sel_shape)
|
||||||
|
if not fancy and r < 0.35 and shape:
|
||||||
|
shape[0] = 1 # broadcast along the first axis
|
||||||
|
kind = rng.choice(["same", "f8", "i8", "u1", "f4"])
|
||||||
|
if kind == "same" and np.dtype(dtype).kind in "iuf":
|
||||||
|
dt = np.dtype(dtype)
|
||||||
|
else:
|
||||||
|
dt = np.dtype(str(kind) if kind != "same" else "f8")
|
||||||
|
base = rng.standard_normal(size=shape) * (10 ** rng.integers(0, 6))
|
||||||
|
with np.errstate(all="ignore"):
|
||||||
|
return base.astype(dt)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("source", SOURCES)
|
||||||
|
@pytest.mark.parametrize("seed", [0, 1, 2, 3])
|
||||||
|
def test_random_edits_match_h5py(h5py, tmp_path, source, seed):
|
||||||
|
"""Random writes (slices, steps, integers, index lists, broadcasts,
|
||||||
|
other dtypes and out-of-range values), resizes and attributes, each
|
||||||
|
compared with h5py applying the same edit."""
|
||||||
|
rng = np.random.default_rng(1000 * seed + SOURCES.index(source))
|
||||||
|
theirs, ours_path, base = _make(h5py, tmp_path, source)
|
||||||
|
with h5py.File(theirs, "r") as t:
|
||||||
|
names = [n for n in ["i4", "u1", "f8", "f4_2d", "f8_3d", "be_i2", "chunk_fixed", "chunk_ext",
|
||||||
|
"chunk_bt2", "chunk_gzip", "chunk_fill", "compact", "grp/leaf"] if n in t]
|
||||||
|
resizable = [n for n in ["chunk_ext", "chunk_bt2", "chunk_gzip", "chunk_fill"] if n in names]
|
||||||
|
refused = 0
|
||||||
|
with clawhdf5.File(ours_path, "r+") as ours:
|
||||||
|
for step in range(40):
|
||||||
|
r = rng.random()
|
||||||
|
if r < 0.15 and resizable:
|
||||||
|
name = str(rng.choice(resizable))
|
||||||
|
with h5py.File(theirs, "r") as t:
|
||||||
|
maxshape = t[name].maxshape
|
||||||
|
shape = t[name].shape
|
||||||
|
new = tuple(int(rng.integers(0, (m if m is not None else s + 10) + 1)) for m, s in zip(maxshape, shape))
|
||||||
|
op = ("resize", name, new, None)
|
||||||
|
elif r < 0.25:
|
||||||
|
obj = str(rng.choice(["/", "grp", names[0]]))
|
||||||
|
choices = [np.int16(rng.integers(-100, 100)), rng.standard_normal(3),
|
||||||
|
np.arange(int(rng.integers(1, 5)), dtype=">u4"), np.float32(0.5)]
|
||||||
|
value = choices[int(rng.integers(0, len(choices)))]
|
||||||
|
op = ("attr", obj, f"a{int(rng.integers(0, 4))}", value)
|
||||||
|
else:
|
||||||
|
name = str(rng.choice(names))
|
||||||
|
with h5py.File(theirs, "r") as t:
|
||||||
|
shape, dtype = t[name].shape, t[name].dtype
|
||||||
|
key = _random_key(rng, shape)
|
||||||
|
sel_shape, fancy = _selection_shape(key, shape)
|
||||||
|
op = ("set", name, key, _random_value(rng, sel_shape, fancy, dtype))
|
||||||
|
if edit_both(h5py, theirs, ours_path, ours, op) is not None:
|
||||||
|
refused += 1
|
||||||
|
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
|
||||||
|
assert refused < 30
|
||||||
|
h5dump_reads(ours_path, base)
|
||||||
|
|
||||||
|
|
||||||
|
NUMERIC = ["<i1", "<u1", "<i2", ">u2", "<i4", "<u4", ">i8", "<u8", "<f2", "<f4", ">f8"]
|
||||||
|
|
||||||
|
|
||||||
|
def _quiet(f):
|
||||||
|
with np.errstate(all="ignore"):
|
||||||
|
return f()
|
||||||
|
|
||||||
|
|
||||||
|
def _libhdf5_undefined(vals, target):
|
||||||
|
"""The values whose conversion to `target` libhdf5 2.0 (h5py 3.16) gets
|
||||||
|
wrong even in native byte order, where its C casts are undefined
|
||||||
|
behaviour; clawhdf5 saturates them as libhdf5's range handling
|
||||||
|
intends (docs/known-issues.md):
|
||||||
|
|
||||||
|
- half floats into unsigned integers: negatives wrap (-1 -> 65535) and
|
||||||
|
+inf becomes 0; into signed integers, +-inf becomes the minimum;
|
||||||
|
- a float equal to the integer maximum rounded up in the float's
|
||||||
|
precision (float32(2**31 - 1) == 2**31 -> int32, float64(2**64 - 1)
|
||||||
|
-> uint64) becomes the minimum (or 0);
|
||||||
|
- a double between 65504 and 65520 into a half float becomes infinity
|
||||||
|
(IEEE rounds it down to 65504, as numpy does).
|
||||||
|
"""
|
||||||
|
bad = np.zeros(vals.shape, dtype=bool)
|
||||||
|
if vals.dtype.kind != "f":
|
||||||
|
return bad
|
||||||
|
if target.kind in "iu":
|
||||||
|
if vals.dtype.itemsize == 2:
|
||||||
|
bad |= np.isinf(vals)
|
||||||
|
if target.kind == "u":
|
||||||
|
bad |= vals <= -1
|
||||||
|
top = _quiet(lambda: np.array(np.iinfo(target).max).astype(vals.dtype))
|
||||||
|
if float(top) > np.iinfo(target).max:
|
||||||
|
bad |= vals == top
|
||||||
|
if target.kind == "f" and target.itemsize == 2 and vals.dtype.itemsize > 2:
|
||||||
|
bad |= (np.abs(vals) > 65504) & (np.abs(vals) < 65520)
|
||||||
|
return bad
|
||||||
|
|
||||||
|
|
||||||
|
def test_numeric_conversions_match_h5py(h5py, tmp_path):
|
||||||
|
"""Every numeric source dtype into every numeric dataset dtype, with
|
||||||
|
values at and beyond the targets' limits, as libhdf5 converts them."""
|
||||||
|
edge = np.array([0, 1, -1, -0.3, 2.5, -2.5, 3.7, -3.7, 127.9, -128.9, 200.5, 255.5, 256, -129,
|
||||||
|
32767.5, 40000, 65504, 70000, -70000, 2**31 - 1, 2**31, -2**31 - 1,
|
||||||
|
4e9, 1e15, -1e15, 1e19, 1e300, -1e300, np.inf, -np.inf])
|
||||||
|
sources = {
|
||||||
|
"f8": edge,
|
||||||
|
"f4": _quiet(lambda: edge.astype("<f4")),
|
||||||
|
"f2": np.array([0, 1, -1, 2.5, -3.5, 65504, -65504, np.inf, -np.inf, 100.5], dtype="<f2"),
|
||||||
|
"i8": np.array([0, 1, -1, 127, 128, -129, 255, 256, 32768, -32769, 65536, 2**31, -2**31 - 1,
|
||||||
|
2**32, 2**62, -2**63, 2**63 - 1], dtype="<i8"),
|
||||||
|
"u8": np.array([0, 1, 127, 128, 255, 256, 65535, 65536, 2**31, 2**32, 2**63, 2**64 - 1], dtype="<u8"),
|
||||||
|
"i1": np.array([-128, -1, 0, 1, 127], dtype="i1"),
|
||||||
|
"u2": np.array([0, 255, 256, 65535], dtype=">u2"),
|
||||||
|
"b": np.array([True, False, True]),
|
||||||
|
}
|
||||||
|
for target in NUMERIC:
|
||||||
|
for sname, src in sources.items():
|
||||||
|
vals = src[~_libhdf5_undefined(src, np.dtype(target))]
|
||||||
|
path_t = str(tmp_path / f"t_{target[1:]}_{sname}.h5")
|
||||||
|
path_o = str(tmp_path / f"o_{target[1:]}_{sname}.h5")
|
||||||
|
with h5py.File(path_t, "w") as f:
|
||||||
|
f.create_dataset("d", shape=vals.shape, dtype=target)
|
||||||
|
shutil.copy(path_t, path_o)
|
||||||
|
what = f"{vals.dtype} -> {target}"
|
||||||
|
try:
|
||||||
|
with h5py.File(path_t, "r+") as f:
|
||||||
|
f["d"][...] = _native_conversion(h5py, vals, f["d"].dtype)
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
with clawhdf5.File(path_o, "r+") as f, pytest.raises(Exception):
|
||||||
|
f["d"][...] = vals
|
||||||
|
continue
|
||||||
|
with clawhdf5.File(path_o, "r+") as f:
|
||||||
|
f["d"][...] = vals
|
||||||
|
with h5py.File(path_t, "r") as a, h5py.File(path_o, "r") as b:
|
||||||
|
assert a["d"][...].tobytes() == b["d"][...].tobytes(), (
|
||||||
|
f"{what}: h5py {a['d'][...].tolist()} clawhdf5 {b['d'][...].tolist()}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_nan_into_an_integer_dataset_is_refused(h5py, tmp_path):
|
||||||
|
"""libhdf5 stores NaN as an arbitrary integer (0, the minimum or 2**63,
|
||||||
|
depending on the type); clawhdf5 refuses and writes nothing."""
|
||||||
|
path = str(tmp_path / "nan.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
with pytest.raises(ValueError, match="NaN"):
|
||||||
|
f["d"][...] = np.array([1.0, np.nan, 2.0, 3.0])
|
||||||
|
# A Python list goes through numpy, which refuses NaN too.
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
f["d"][0:2] = [np.nan, 1.0]
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
np.testing.assert_array_equal(f["d"][...], np.arange(4))
|
||||||
|
|
||||||
|
|
||||||
|
def test_unsupported_edits_are_clear_errors(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "u.h5")
|
||||||
|
_h5py_file(h5py, path, "earliest")
|
||||||
|
before = snapshot(h5py, path)
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
with pytest.raises(NotImplementedError, match="delet"):
|
||||||
|
del f.attrs["version"]
|
||||||
|
with pytest.raises(NotImplementedError, match="delet"):
|
||||||
|
del f["i4"]
|
||||||
|
with pytest.raises(NotImplementedError, match="delet"):
|
||||||
|
del f["grp"]["leaf"]
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f.create_dataset("new", data=np.arange(3.0))
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f.create_group("newgrp")
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f["grp"].create_dataset("new", data=np.arange(3.0))
|
||||||
|
with pytest.raises(NotImplementedError, match="variable-length"):
|
||||||
|
f["vlen"][0] = "x"
|
||||||
|
with pytest.raises(NotImplementedError, match="field"):
|
||||||
|
f["cmp"]["id"] = np.arange(5)
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f.attrs["empty"] = clawhdf5.Empty("f8")
|
||||||
|
with pytest.raises(TypeError, match="chunked"):
|
||||||
|
f["i4"].resize((7, 10))
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
f["chunk_gzip"].resize((41, 20))
|
||||||
|
with pytest.raises(ValueError, match="axis"):
|
||||||
|
f["chunk_ext"].resize(3, axis=2)
|
||||||
|
with pytest.raises(TypeError):
|
||||||
|
f["i4"][0] = np.array(["a"] * 10)
|
||||||
|
assert snapshot(h5py, path) == before
|
||||||
|
h5dump_reads(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_read_only_files_and_modes(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "m.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"), chunks=(2,), maxshape=(None,))
|
||||||
|
with clawhdf5.File(path, "r") as f:
|
||||||
|
with pytest.raises(OSError, match="r\\+"):
|
||||||
|
f["d"][0] = 1
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
f["d"].resize((8,))
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
f.attrs["x"] = 1
|
||||||
|
with pytest.raises(NotImplementedError, match="does not exist"):
|
||||||
|
clawhdf5.File(str(tmp_path / "missing.h5"), "a")
|
||||||
|
with pytest.raises(ValueError, match="mode"):
|
||||||
|
clawhdf5.File(path, "rw")
|
||||||
|
with clawhdf5.File(path, "a") as f:
|
||||||
|
assert f.mode == "r+"
|
||||||
|
f["d"][1] = 10
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
assert f["d"][1] == 10
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_file_is_locked_while_open_for_editing(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "lock.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||||
|
f = clawhdf5.File(path, "r+")
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
clawhdf5.File(path, "r+")
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
h5py.File(path, "r+")
|
||||||
|
f.close()
|
||||||
|
with h5py.File(path, "r+") as g:
|
||||||
|
g["d"][0] = 5
|
||||||
|
with clawhdf5.File(path, "r+") as g:
|
||||||
|
g["d"][1] = 6
|
||||||
|
with h5py.File(path, "r") as g:
|
||||||
|
np.testing.assert_array_equal(g["d"][...], [5, 6, 2, 3])
|
||||||
|
|
||||||
|
|
||||||
|
def test_objects_see_edits_made_through_others(h5py, tmp_path):
|
||||||
|
"""A dataset or attrs object taken before an edit reports the file as
|
||||||
|
it is after it: the new shape, the new attribute."""
|
||||||
|
path = str(tmp_path / "live.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(6.0), chunks=(4,), maxshape=(None,))
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
d1 = f["d"]
|
||||||
|
d2 = f["d"]
|
||||||
|
attrs = d1.attrs
|
||||||
|
assert len(attrs) == 0 and "u" not in attrs
|
||||||
|
d2.resize((10,))
|
||||||
|
assert d1.shape == (10,) and d1.size == 10 and len(d1) == 10
|
||||||
|
np.testing.assert_array_equal(d1[6:], np.zeros(4))
|
||||||
|
d2.attrs["u"] = "m/s"
|
||||||
|
assert "u" in attrs and attrs["u"] == b"m/s" and len(attrs) == 1
|
||||||
|
f.attrs.create("shaped", np.arange(6), shape=(2, 3), dtype="<i2")
|
||||||
|
f.attrs.modify("shaped2", [1.5, 2.5])
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
assert f["d"].shape == (10,)
|
||||||
|
assert f["d"].attrs["u"] == b"m/s"
|
||||||
|
assert f.attrs["shaped"].dtype == np.dtype("<i2") and f.attrs["shaped"].shape == (2, 3)
|
||||||
|
np.testing.assert_array_equal(f.attrs["shaped2"], [1.5, 2.5])
|
||||||
|
|
||||||
|
|
||||||
|
def test_attribute_types_as_h5py_reads_them(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "attrs.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_group("g")
|
||||||
|
values = {
|
||||||
|
"i8": 5,
|
||||||
|
"f8": 2.5,
|
||||||
|
"i1": np.int8(-3),
|
||||||
|
"u8": np.uint64(2**64 - 1),
|
||||||
|
">f4": np.array([1.5, 2.5], dtype=">f4"),
|
||||||
|
"f2": np.float16(0.5),
|
||||||
|
"b": True,
|
||||||
|
"barr": np.array([True, False]),
|
||||||
|
"c16": np.complex128(1 + 2j),
|
||||||
|
"bytes": b"raw",
|
||||||
|
"sarr": np.array([b"a", b"bcd"]),
|
||||||
|
"str": "héllo",
|
||||||
|
"strs": ["x", "yz"],
|
||||||
|
"2d": np.arange(12, dtype="<u2").reshape(3, 4),
|
||||||
|
"empty": np.zeros((0,), dtype="<i4"),
|
||||||
|
}
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
for k, v in values.items():
|
||||||
|
f["g"].attrs[k] = v
|
||||||
|
# Replace one, with another type and size.
|
||||||
|
f["g"].attrs["i8"] = np.arange(100.0)
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
a = f["g"].attrs
|
||||||
|
np.testing.assert_array_equal(a["i8"], np.arange(100.0))
|
||||||
|
assert a["f8"] == 2.5 and a["f8"].dtype == np.float64
|
||||||
|
assert a["i1"] == -3 and a["i1"].dtype == np.int8
|
||||||
|
assert a["u8"] == 2**64 - 1 and a["u8"].dtype == np.uint64
|
||||||
|
assert a[">f4"].dtype == np.dtype(">f4")
|
||||||
|
assert a["f2"].dtype == np.float16
|
||||||
|
assert a["b"] is np.True_ or a["b"] == True # noqa: E712
|
||||||
|
assert a["barr"].dtype == np.bool_
|
||||||
|
assert a["c16"] == 1 + 2j
|
||||||
|
assert a["bytes"] == b"raw"
|
||||||
|
assert list(a["sarr"]) == [b"a", b"bcd"]
|
||||||
|
# str is stored as fixed-length UTF-8: h5py reads bytes.
|
||||||
|
assert a["str"].decode("utf-8") == "héllo"
|
||||||
|
assert [x.decode() for x in a["strs"]] == ["x", "yz"]
|
||||||
|
assert a["2d"].shape == (3, 4) and a["2d"].dtype == np.dtype("<u2")
|
||||||
|
assert a["empty"].shape == (0,)
|
||||||
|
with clawhdf5.File(path, "r") as f:
|
||||||
|
assert f["g"].attrs["c16"] == 1 + 2j
|
||||||
|
assert f["g"].attrs["str"].decode("utf-8") == "héllo"
|
||||||
|
h5dump_reads(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_many_attributes_move_to_dense_storage(h5py, tmp_path):
|
||||||
|
"""Past the compact limit (8 attributes under libver v114) the object's
|
||||||
|
attributes move to dense storage; h5py reads all of them."""
|
||||||
|
path = str(tmp_path / "dense.h5")
|
||||||
|
with h5py.File(path, "w", libver="v114") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(3))
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
for i in range(20):
|
||||||
|
f["d"].attrs[f"a{i:02d}"] = np.full(i + 1, i, dtype="<i2")
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
assert sorted(f["d"].attrs.keys()) == [f"a{i:02d}" for i in range(20)]
|
||||||
|
for i in range(20):
|
||||||
|
np.testing.assert_array_equal(f["d"].attrs[f"a{i:02d}"], np.full(i + 1, i))
|
||||||
|
h5dump_reads(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_reads_never_see_a_half_written_edit(h5py, tmp_path):
|
||||||
|
"""Readers on other threads while one thread rewrites a dataset: every
|
||||||
|
read returns one whole version (all elements equal), never a mix."""
|
||||||
|
path = str(tmp_path / "race.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.zeros((64, 64)), chunks=(16, 16), compression="gzip")
|
||||||
|
errors = []
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
stop = threading.Event()
|
||||||
|
|
||||||
|
def read():
|
||||||
|
ds = f["d"]
|
||||||
|
while not stop.is_set():
|
||||||
|
a = ds[...]
|
||||||
|
if not (a == a.flat[0]).all():
|
||||||
|
errors.append(a)
|
||||||
|
return
|
||||||
|
|
||||||
|
readers = [threading.Thread(target=read) for _ in range(3)]
|
||||||
|
for t in readers:
|
||||||
|
t.start()
|
||||||
|
try:
|
||||||
|
for k in range(1, 25):
|
||||||
|
f["d"][...] = float(k)
|
||||||
|
finally:
|
||||||
|
stop.set()
|
||||||
|
for t in readers:
|
||||||
|
t.join()
|
||||||
|
np.testing.assert_array_equal(f["d"][...], np.full((64, 64), 24.0))
|
||||||
|
assert not errors, "a read saw a partly written dataset"
|
||||||
|
|
||||||
|
|
||||||
|
def test_close_releases_the_file(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "close.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||||
|
f = clawhdf5.File(path, "r+")
|
||||||
|
ds = f["d"]
|
||||||
|
f.close()
|
||||||
|
# The handle still reads the file as last written, but cannot edit it.
|
||||||
|
np.testing.assert_array_equal(ds[...], np.arange(4))
|
||||||
|
with pytest.raises(OSError, match="closed"):
|
||||||
|
ds[0] = 1
|
||||||
|
with h5py.File(path, "r+") as g:
|
||||||
|
g["d"][0] = 9
|
||||||
@@ -178,12 +178,15 @@ def _write_fixture(h5py, path):
|
|||||||
g.attrs["depth"] = np.int8(3)
|
g.attrs["depth"] = np.int8(3)
|
||||||
|
|
||||||
|
|
||||||
@pytest.fixture(scope="module", params=["local", "http", "http-1k-blocks"])
|
@pytest.fixture(scope="module", params=["local", "r+", "http", "http-1k-blocks"])
|
||||||
def pair(request, h5py, tmp_path_factory):
|
def pair(request, h5py, tmp_path_factory):
|
||||||
"""The fixture file through h5py and through clawhdf5: opened locally,
|
"""The fixture file through h5py and through clawhdf5: opened locally
|
||||||
and over HTTP range requests (a local server in this process) with the
|
(read-only, and for editing: a copy, since editing locks the file), and
|
||||||
|
over HTTP range requests (a local server in this process) with the
|
||||||
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
|
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
|
||||||
through many small ranges."""
|
through many small ranges."""
|
||||||
|
import shutil
|
||||||
|
|
||||||
from conftest import RangeServer
|
from conftest import RangeServer
|
||||||
|
|
||||||
root = tmp_path_factory.mktemp("h5")
|
root = tmp_path_factory.mktemp("h5")
|
||||||
@@ -193,6 +196,10 @@ def pair(request, h5py, tmp_path_factory):
|
|||||||
server = None
|
server = None
|
||||||
if request.param == "local":
|
if request.param == "local":
|
||||||
ours = clawhdf5.File(path, "r")
|
ours = clawhdf5.File(path, "r")
|
||||||
|
elif request.param == "r+":
|
||||||
|
copy = str(root / "editable.h5")
|
||||||
|
shutil.copy(path, copy)
|
||||||
|
ours = clawhdf5.File(copy, "r+")
|
||||||
else:
|
else:
|
||||||
server = RangeServer(root)
|
server = RangeServer(root)
|
||||||
if request.param == "http":
|
if request.param == "http":
|
||||||
|
|||||||
@@ -116,6 +116,50 @@ space it leaves is too small for its next, larger version.
|
|||||||
**No journal.** A crash while an edit patches existing structures can leave
|
**No journal.** A crash while an edit patches existing structures can leave
|
||||||
the file inconsistent; see the `FileEditor` documentation.
|
the file inconsistent; see the `FileEditor` documentation.
|
||||||
|
|
||||||
|
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
|
||||||
|
|
||||||
|
**Status:** open (added 2026-09-27). The Python bindings edit through
|
||||||
|
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
|
||||||
|
before anything is written. On top of them:
|
||||||
|
|
||||||
|
- **No new or deleted objects:** `create_dataset`/`create_group` in an
|
||||||
|
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
|
||||||
|
`NotImplementedError` (the editor changes values, shapes and
|
||||||
|
attributes only). Mode `'a'` works on an existing file only.
|
||||||
|
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
|
||||||
|
whole elements of the same structured dtype are), HDF5 array-type
|
||||||
|
elements, variable-length data, strings padded with spaces or
|
||||||
|
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
|
||||||
|
ones, h5py's, are writable), compounds containing such strings, null
|
||||||
|
dataspaces, and index-list writes of more than 2²² elements (write them
|
||||||
|
in slices).
|
||||||
|
- **`str` attributes are fixed-length UTF-8**, where h5py writes
|
||||||
|
variable-length strings: h5py reads them back as `bytes`
|
||||||
|
(`numpy.bytes_`), not `str`.
|
||||||
|
- **Numeric conversion follows libhdf5's native-order results, not its
|
||||||
|
bugs.** Arrays are converted as libhdf5 converts them (integers
|
||||||
|
saturate, floats are truncated toward zero and clipped), checked value by
|
||||||
|
value against h5py 3.16 / HDF5 2.0 in
|
||||||
|
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
|
||||||
|
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
|
||||||
|
differs from h5py on purpose:
|
||||||
|
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
|
||||||
|
minimum or 2⁶³ depending on the type);
|
||||||
|
- when the dataset or the array is not in native byte order, libhdf5's
|
||||||
|
"soft" conversions store a float in (-1, 0) as the integer minimum and
|
||||||
|
wrap an unsigned value too large for the signed type of the same size
|
||||||
|
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
|
||||||
|
native order;
|
||||||
|
- libhdf5's native casts that are undefined in C: half floats into
|
||||||
|
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
|
||||||
|
integers (→ minimum), a float equal to the integer maximum rounded up
|
||||||
|
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
|
||||||
|
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
|
||||||
|
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
|
||||||
|
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
|
||||||
|
- **Each edit reopens the file** (a new memory map) so that reads see it;
|
||||||
|
reads from other threads wait while an edit is written.
|
||||||
|
|
||||||
## Selection reads that decode more than the selection
|
## Selection reads that decode more than the selection
|
||||||
|
|
||||||
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
|
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
|
||||||
|
|||||||
+4
-1
@@ -272,7 +272,10 @@ python_package() {
|
|||||||
-i "$PYTHON" \
|
-i "$PYTHON" \
|
||||||
--out "$out/wheel" || return 1
|
--out "$out/wheel" || return 1
|
||||||
"$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1
|
"$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1
|
||||||
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
|
# The editing tests run `h5rs check` on every file they edit.
|
||||||
|
cargo build -q -p clawhdf5-tools || return 1
|
||||||
|
CLAWHDF5_H5RS="${CARGO_TARGET_DIR:-$root/target}/debug/h5rs" \
|
||||||
|
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
|
||||||
"$root/crates/clawhdf5-py/tests"
|
"$root/crates/clawhdf5-py/tests"
|
||||||
}
|
}
|
||||||
if "$PYTHON" -m maturin --version >/dev/null 2>&1 && "$PYTHON" -c "import pytest" >/dev/null 2>&1; then
|
if "$PYTHON" -m maturin --version >/dev/null 2>&1 && "$PYTHON" -c "import pytest" >/dev/null 2>&1; then
|
||||||
|
|||||||
Reference in New Issue
Block a user