Files
clawhdf5/crates/clawhdf5-py/src/group.rs
T
osobhandClaude Opus 5.5 92fb0830e0 py: in-place editing (clawhdf5.File(path, 'r+')) through FileEditor
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():

- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
  slices and integers with extra leading 1-axes allowed; the exact shape
  for an index list, a scalar only where h5py expands it). Arrays are
  converted as libhdf5 converts them in native byte order (integers
  saturate, floats truncate toward zero and clip, integers go into h5py's
  bool enum by value); other values through
  numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
  dataset is a ValueError instead of libhdf5's arbitrary value. The value
  preparation is a small Python module compiled into the extension
  (src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
  attrs.modify: numeric, bool, complex, bytes and str data of any shape,
  with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
  cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.

Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.

Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.

Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 06:57:36 -05:00

403 lines
14 KiB
Rust

//! PyGroup — navigable HDF5 group with read and write support.
use std::collections::HashMap;
use std::sync::{Arc, Mutex, OnceLock};
use pyo3::exceptions::{PyIOError, PyKeyError, PyNotImplementedError, PyOSError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::PyList;
use crate::attrs::PyAttrs;
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
/// Shared state for a group being written.
pub(crate) struct WriteGroupState {
pub name: String,
pub datasets: Vec<DatasetSpec>,
pub attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
}
/// An HDF5 group.
///
/// In read mode it behaves like an h5py group: `grp['name']`,
/// `grp['sub/path']` and `grp['/absolute/path']`, `keys()`, `values()`,
/// `items()`, iteration, `len()`, `in`, `get()`, `name` and `attrs`.
/// In write mode, supports `create_dataset` and attribute setting.
#[pyclass(name = "Group")]
pub struct PyGroup {
inner: GroupInner,
}
enum GroupInner {
Read(ReadGroup),
Write(Arc<Mutex<WriteGroupState>>),
}
impl PyGroup {
pub(crate) fn from_read(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self {
inner: GroupInner::Read(ReadGroup::new(handle, path, addr)),
}
}
pub(crate) fn from_write(state: Arc<Mutex<WriteGroupState>>) -> Self {
Self {
inner: GroupInner::Write(state),
}
}
fn read_group(&self, what: &str) -> PyResult<&ReadGroup> {
match &self.inner {
GroupInner::Read(g) => Ok(g),
GroupInner::Write(_) => Err(PyIOError::new_err(format!(
"cannot {what} a group opened for writing"
))),
}
}
}
/// A group in a file opened for reading (a file is its root group, as in
/// h5py). It keeps its own address and, once listed, its links, so looking
/// up a child neither resolves the path from the root nor scans the group's
/// links again: visiting every member of a large group is linear, not
/// quadratic. (Edits never add or remove links, so these stay valid in a
/// file open for editing.)
pub(crate) struct ReadGroup {
pub handle: Arc<Handle>,
pub path: String,
pub addr: u64,
/// Link name -> object address (soft links resolved), filled on first use.
links: OnceLock<HashMap<String, u64>>,
/// Names of the datasets and subgroups, sorted (h5py's order).
members: OnceLock<Vec<String>>,
}
impl ReadGroup {
pub(crate) fn new(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self {
handle,
path,
addr,
links: OnceLock::new(),
members: OnceLock::new(),
}
}
fn links(&self, py: Python<'_>) -> PyResult<&HashMap<String, u64>> {
if let Some(links) = self.links.get() {
return Ok(links);
}
let (addr, path) = (self.addr, &self.path);
let entries = self.handle.with(py, |f| {
clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr)
.map_err(|e| node::format_err(path, e, PyValueError::new_err))
})?;
let map = entries
.into_iter()
.map(|e| (e.name, e.object_header_address))
.collect();
Ok(self.links.get_or_init(|| map))
}
/// The path and address of `key` (a name, a relative or an absolute path).
fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> {
let path = node::join(&self.path, key);
let rel = if self.path.is_empty() {
Some(path.as_str())
} else if path == self.path {
Some("")
} else {
path.strip_prefix(self.path.as_str())
.and_then(|r| r.strip_prefix('/'))
};
// A direct child: the link table, when it has the name.
if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/'))
&& let Some(&a) = self.links(py)?.get(name)
{
return Ok((path, a));
}
let addr = self.addr;
let found = self.handle.with(py, |f| match rel {
Some(rel) => node::resolve_from(f, addr, rel, &path),
None => node::address(f, &path),
})?;
Ok((path, found))
}
/// `group[key]`.
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
let (path, addr) = self.locate(py, key)?;
node::open(py, &self.handle, path, addr)
}
/// `group.get(key, default)`.
pub(crate) fn get(
&self,
py: Python<'_>,
key: &str,
default: Option<Py<PyAny>>,
) -> PyResult<Py<PyAny>> {
match self.get_item(py, key) {
Err(e) if e.is_instance_of::<PyKeyError>(py) => {
Ok(default.unwrap_or_else(|| py.None()))
}
other => other,
}
}
/// Names of the group's datasets and subgroups, sorted (h5py's order).
pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> {
if let Some(m) = self.members.get() {
return Ok(m);
}
let links = self.links(py)?;
let path = &self.path;
let mut names = self.handle.with(py, |f| {
let mut names = Vec::new();
for (name, &addr) in links {
if matches!(
node::kind_at(f, addr, &node::join(path, name))?,
Some(node::Kind::Dataset | node::Kind::Group)
) {
names.push(name.clone());
}
}
Ok(names)
})?;
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
Ok(self.members.get_or_init(|| names))
}
/// `key in group`: whether `key` names a dataset or group. A failed
/// read of the file (a network error) is raised, not `False`.
pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
let found = self
.locate(py, key)
.and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path)));
match found {
Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)),
Err(e) if e.is_instance_of::<PyOSError>(py) => Err(e),
Err(_) => Ok(false),
}
}
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
self.member_names(py)?
.iter()
.map(|n| self.get_item(py, n))
.collect()
}
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
self.member_names(py)?
.iter()
.map(|n| Ok((n.clone(), self.get_item(py, n)?)))
.collect()
}
pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
}
}
#[pymethods]
impl PyGroup {
/// Get a child object (dataset or subgroup) by name or path.
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
self.read_group("read children from")?.get_item(py, key)
}
/// `group.get(key, default=None)`.
#[pyo3(signature = (key, default=None))]
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
self.read_group("read children from")?.get(py, key, default)
}
/// List the names of all children (datasets and subgroups).
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
match &self.inner {
GroupInner::Read(g) => {
let list = PyList::new(py, g.member_names(py)?)?;
Ok(list.into_any().unbind())
}
GroupInner::Write(state) => {
let guard = state.lock().unwrap();
let names: Vec<&str> = guard.datasets.iter().map(|d| d.name.as_str()).collect();
let list = PyList::new(py, &names)?;
Ok(list.into_any().unbind())
}
}
}
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let g = self.read_group("read children from")?;
Ok(PyList::new(py, g.values(py)?)?.into_any().unbind())
}
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let g = self.read_group("read children from")?;
Ok(PyList::new(py, g.items(py)?)?.into_any().unbind())
}
fn __iter__(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
self.keys(py)?.call_method0(py, "__iter__")
}
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match &self.inner {
GroupInner::Read(g) => Ok(g.member_names(py)?.len()),
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
}
}
/// The group's full name, e.g. `/sensors`.
#[getter]
fn name(&self) -> String {
match &self.inner {
GroupInner::Read(g) => node::name(&g.path),
GroupInner::Write(state) => node::name(&state.lock().unwrap().name),
}
}
/// Create a dataset inside this group (write mode only).
///
/// Parameters:
/// name: dataset name
/// data: numpy array
/// chunks: optional chunk dimensions
/// compression: optional, only 'gzip' supported
/// compression_opts: gzip level (1-9)
#[pyo3(signature = (name, *, data, chunks=None, compression=None, compression_opts=None))]
fn create_dataset(
&self,
py: Python<'_>,
name: &str,
data: &Bound<'_, PyAny>,
chunks: Option<Vec<u64>>,
compression: Option<&str>,
compression_opts: Option<u32>,
) -> PyResult<()> {
match &self.inner {
GroupInner::Write(state) => {
let (dataset_data, shape) = extract_numpy_data(py, data)?;
let deflate_level = match compression {
Some("gzip") => Some(compression_opts.unwrap_or(4)),
Some(other) => {
return Err(PyErr::new::<pyo3::exceptions::PyValueError, _>(format!(
"unsupported compression: {other}; only 'gzip' is supported"
)));
}
None => None,
};
let spec = DatasetSpec {
name: name.to_string(),
data: dataset_data,
shape,
chunks,
deflate_level,
attrs: vec![],
};
state.lock().unwrap().datasets.push(spec);
Ok(())
}
GroupInner::Read(g) if g.handle.is_writable() => Err(PyNotImplementedError::new_err(
"creating datasets or groups in an existing file is not supported by \
clawhdf5's in-place editor (mode 'r+' changes values, shapes and attributes)",
)),
GroupInner::Read(_) => Err(PyIOError::new_err(
"cannot create datasets on a read-only group",
)),
}
}
/// Deleting objects is not supported (h5py's `del group[name]`).
fn __delitem__(&self, key: &str) -> PyResult<()> {
Err(PyNotImplementedError::new_err(format!(
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
)))
}
/// Attribute access.
#[getter]
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match &self.inner {
GroupInner::Read(g) => g.attrs(py),
GroupInner::Write(state) => {
let store = Arc::clone(&state.lock().unwrap().attrs);
Ok(PyAttrs::from_write(store))
}
}
}
fn __repr__(&self, py: Python<'_>) -> String {
match &self.inner {
GroupInner::Read(g) => {
let n = g.member_names(py).map_or(0, |m| m.len());
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
}
GroupInner::Write(state) => {
let name = &state.lock().unwrap().name;
format!("<HDF5 Group \"{name}\" (write)>")
}
}
}
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
match &self.inner {
GroupInner::Read(g) => g.contains(py, key),
GroupInner::Write(state) => {
let guard = state.lock().unwrap();
Ok(guard.datasets.iter().any(|d| d.name == key))
}
}
}
}
/// Finalize a write group into the file builder.
pub(crate) fn finalize_write_group(
builder: &mut clawhdf5_rs::FileBuilder,
state: &WriteGroupState,
) {
let mut gb = builder.create_group(&state.name);
for spec in &state.datasets {
let db = gb.create_dataset(&spec.name);
apply_dataset_spec(db, spec);
}
let attrs_guard = state.attrs.lock().unwrap();
for (name, val) in attrs_guard.iter() {
gb.set_attr(name, val.clone().into());
}
let finished = gb.finish();
builder.add_group(finished);
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn finalize_group() {
let state = WriteGroupState {
name: "mygroup".into(),
datasets: vec![DatasetSpec {
name: "vals".into(),
data: crate::DatasetData::F64(vec![1.0, 2.0]),
shape: vec![2],
chunks: None,
deflate_level: None,
attrs: vec![],
}],
attrs: Arc::new(Mutex::new(vec![("version".into(), OwnedAttrValue::I64(1))])),
};
let mut builder = clawhdf5_rs::FileBuilder::new();
// Need a root dataset for a valid file
builder.create_dataset("root_ds").with_f64_data(&[0.0]);
finalize_write_group(&mut builder, &state);
let bytes = builder.finish().unwrap();
let file = clawhdf5_rs::File::from_bytes(bytes).unwrap();
let ds = file.dataset("mygroup/vals").unwrap();
assert_eq!(ds.read_f64().unwrap(), vec![1.0, 2.0]);
}
}