Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python

# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
This commit is contained in:
osobh
2026-09-27 07:59:14 -05:00
33 changed files with 4057 additions and 393 deletions
+142
View File
@@ -72,6 +72,148 @@ Design: `docs/design/swmr.md`.
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing `HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
groups or attributes (a SWMR writer cannot add objects or attributes). groups or attributes (a SWMR writer cannot add objects or attributes).
### Correctness: edits planned from another file after a rename or `chdir` (2026-09-27)
- **`FileEditor` planned each edit by re-opening its path but wrote
through the file it held open** (fixed 2026-09-27; on main since PR #18,
no release). When the path came to name another file between edits — a
rename or replacement, or, for a relative path, a change of working
directory — an edit was laid out from the other file's metadata and
written into the held one, corrupting it (h5py: "invalid dataset size,
likely file corruption"). The Python `'r+'` handle had the same flaw in
its reads: it reopened the path after every edit, so reads came from the
other file. The editor now plans every edit from the file it holds, and
its path is canonicalised at open. New `FileEditor::reader()` opens the
held file anew for reading (on Linux through `/proc/self/fd`, so it
follows a renamed file; elsewhere by the path, refused when the path no
longer names the held file), without sharing the editor's lock; the
Python handle reads through it, and a `'w'` file is written at the
absolute path it was opened with. Tests: `edit_tests.rs`'s
`edits_go_to_the_file_held_not_the_path`; `test_edit.py`'s
`test_relative_path_and_chdir` and `test_path_replaced_between_edits`
(the review's repro).
### Correctness: zero extents in Fixed/Extensible Array chunk indexes (2026-09-27)
- **A chunked dataset whose maximum (or, with none recorded, current)
extent is 0 along a dimension made the reader divide by zero** (fixed
2026-09-27): `h5rs check` panicked ("attempt to divide by zero",
`chunk_grid.rs`) and the next `FileEditor::resize` failed with an
internal error. The unfixed editor produced such files by resizing a
clawhdf5-written dataset to a zero extent (12 of 30 extra random-edit
seeds on clawhdf5-written files). Such an index has no slot for any
chunk of the dataset, and `ChunkGrid::offsets` now says so instead of
dividing by the zero stride. Tests: `chunk_grid`'s
`zero_extent_has_no_chunks`, `edit_interop.rs`'s
`zero_extent_resizes_without_a_recorded_maximum` (a file the unfixed
editor left checks clean and resizes on; a 2.7.0-written file through
zero extents checks clean at each step), and `test_edit.py`'s random
edits on seeds 10 to 39 of a clawhdf5-written file.
### Correctness: resizing chunked datasets with no recorded maximum (2026-09-27)
- **`FileEditor::resize` scrambled the values of a chunked dataset whose
dataspace records no maximum dimensions when it shrank it** (fixed
2026-09-27). The editor shipped on main in PR #18 (a4c2ace) and reached
Python as `Dataset.resize` in `'r+'` files; no release has it. clawhdf5's
writer stores such a dataspace for every chunked dataset created without
a `maxshape`, with a Fixed Array (or Single Chunk) chunk index. The
editor patched only the current dimensions, and with no maximum
recorded the maximum is the current dimensions — which is also what the
Fixed Array linearises chunks by — so a shrink moved every chunk after
the first row: h5py, h5dump and our reader all read wrong values
without complaint (20x20, chunks 6x6, resized to 15x15: row 6 read
`0 0 0 0 0 0 120 ...`). After a shrink the dataset could not grow back
either (`3 exceeds the maximum 0`). libhdf5 itself never writes such a
dataspace (`H5S_set_extent_simple` records the maximum, equal to the
dimensions when none is given); reading one, `H5S_extent_get_dims`
reports the current dimensions as the maximum and `H5S_set_extent`
checks against no maximum at all, so libhdf5's own `H5Dset_extent`
scrambles such a file the same way (and lets it grow past its Fixed
Array). The editor now records the maximum libhdf5 would have written
— the dimensions before the first resize, the ones the index was built
with — then changes the current ones (the dataspace message grows by one
length per dimension and moves in the header when it has to). The
dataset then shrinks, grows back to that extent and refuses more, as
one libhdf5 wrote would. The writer (`FileBuilder`) now records the
maximum of every chunked dataset too, as libhdf5 does, so h5py can
resize what it writes (8 more bytes per dimension). Tests:
`crates/clawhdf5/tests/edit_resize_interop.rs` (a 2.7.0-written fixture,
new `FileBuilder` files and h5py files through shrinks, zero extents
and growth, checked against a model with our reader and h5py; h5py
resizing a `FileBuilder` file) and `test_edit.py`'s numpy-model checks.
### Python bindings: in-place editing (2026-09-27)
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
file for editing through `clawhdf5::FileEditor`, holding its exclusive
lock until `close()`:
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
one increasing index list) and broadcasting (numpy's rules for slices
and integers, allowing extra leading length-1 axes; the exact shape for
an index list, a scalar only where h5py expands it). A numpy array is
converted to the dataset's dtype as libhdf5 converts it in native byte
order (integers saturate; floats are truncated toward zero and clipped;
integers go into h5py's bool enum by value, as libhdf5 stores them);
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
does. NaN into an integer dataset is a `ValueError`.
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
and errors (`TypeError` for a dataset that is not chunked).
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
enum, complex as the `r`/`i` compound), except that `str` becomes
fixed-length UTF-8.
- Every edit is written and synced before it returns, then the file is
reopened: datasets and attrs objects taken earlier see new shapes and
attributes, and reads on other threads wait while an edit is written.
- What the editor cannot do raises `NotImplementedError` and writes
nothing (deleting attributes or objects, creating datasets or groups,
compound fields by name, variable-length data, ...;
`docs/known-issues.md`).
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
- Boolean-mask keys (`ds[mask]`, `ds[mask] = v`), which h5py supports,
raise `NotImplementedError` (they raised `TypeError`).
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
by h5py and by clawhdf5, and both read back through h5py after every
edit, on files h5py writes with `libver` earliest, v114 and latest and on
a clawhdf5-written one: a fixed sequence over every chunk index kind,
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
and compound types, and 16 random sequences of 40 edits (writes,
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
and leave its file unchanged. A matrix of every numeric source dtype into
every numeric dataset dtype at the edge values, dense attribute storage,
locking, objects seeing each other's edits, readers racing a writer
(never a partly written dataset). Every edited file must pass `h5dump`
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
a file opened `'r+'`.
### Python bindings: remote files (2026-09-27)
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
through the block cache, the whole read API (groups, attributes, every
dataset type and index the local reader handles). A URL is any
`scheme://…`; a remote file is read-only (another mode is a
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
`allow_full_download`, `max_full_download`, `require_validator`,
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
cache's counters. The default wheel builds plain HTTP only (no C: rustls
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
check now covers `clawhdf5-py`.
- **Every read parses through `File::storage()`** instead of
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
group listings, variable-length data through the global heap), inside one
shared file handle that releases the GIL for all file access, not only
dataset reads: a read waiting on the network lets other Python threads
run. A failed read of the storage (a network error, a file changed on
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
data; `key in group` raises it instead of answering `False`.
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
blocks and 1 KiB blocks) against a range-capable `http.server` in the
test process, plus `tests/test_remote.py`: request counts of a small
read, cache hits, a server without `Range` support (refused, or a
whole download when allowed), a file changed on the server, a server
that hangs up, 16 threads on one remote file, and a thread that keeps
running while a read waits on 0.2 s requests.
### Range reads, milestone M3: remote files (2026-09-26) ### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives - **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by a `clawhdf5::File` (through `File::open_storage`) that reads the file by
+30 -1
View File
@@ -574,8 +574,36 @@ with clawhdf5.File("data.h5", "r") as f:
records = f["table"] # compound -> numpy structured array records = f["table"] # compound -> numpy structured array
ids = records["id"] # one field ids = records["id"] # one field
# A file on a web server: range requests through a block cache, nothing
# downloaded up front; the same read API. The GIL is released while waiting.
with clawhdf5.File("http://data.example.org/run42.h5") as f:
first = f["group/temperatures"][0]
f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024,
headers={"Authorization": "Bearer ..."})
``` ```
An existing file opened with `"r+"` is edited in place (through
`clawhdf5::FileEditor`), with h5py's indexing, broadcasting and numeric
conversion; each edit is on disk when the statement returns:
```python
with clawhdf5.File("data.h5", "r+") as f:
f["group/temperatures"][100:200, ::4] = 0.0
f["series"].resize(5000, axis=0) # chunked datasets, within maxshape
f["series"][4000:] = new_values
f["group"].attrs["calibrated"] = True
```
Creating or deleting datasets, groups and attributes in an existing file is
not supported (`NotImplementedError`); limits are in
[known issues](docs/known-issues.md).
The default build reads `http://` URLs only; build with
`maturin develop --release --features https` (rustls with ring, which
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
object-store URLs.
Reads cover integers and IEEE floats of every width in either byte order, Reads cover integers and IEEE floats of every width in either byte order,
`bool`, enums, complex, fixed and variable-length strings, variable-length `bool`, enums, complex, fixed and variable-length strings, variable-length
sequences, opaque, HDF5 array types and compounds; other types (references, sequences, opaque, HDF5 array types and compounds; other types (references,
@@ -590,7 +618,8 @@ non-default fill value (`docs/known-issues.md`). An index list is read one
group of neighbouring chunks at a time. group of neighbouring chunks at a time.
Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`) Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`)
covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests
in `crates/clawhdf5-py/tests` compare every read with h5py; run them with in `crates/clawhdf5-py/tests` compare every read and every in-place edit
with h5py; run them with
`pip install pytest h5py && pytest crates/clawhdf5-py/tests`. `pip install pytest h5py && pytest crates/clawhdf5-py/tests`.
### Agent Memory ### Agent Memory
+28
View File
@@ -135,6 +135,12 @@ impl ChunkGrid {
let mut rem = index; let mut rem = index;
for p in 0..rank { for p in 0..rank {
let d = self.order[p]; let d = self.order[p];
// A zero stride: a later dimension has no chunks (its maximum,
// or with none recorded its current extent, is 0), so no slot of
// the index is a chunk of the dataset.
if self.down[p] == 0 {
return None;
}
let scaled = rem / self.down[p]; let scaled = rem / self.down[p];
rem %= self.down[p]; rem %= self.down[p];
if scaled >= self.cur_chunks[d] { if scaled >= self.cur_chunks[d] {
@@ -193,6 +199,28 @@ mod tests {
assert_eq!(g.offsets(11), Some(vec![2, 3])); assert_eq!(g.offsets(11), Some(vec![2, 3]));
} }
#[test]
fn zero_extent_has_no_chunks() {
// No maximum recorded and a zero current dimension: every stride
// before it is 0 (this divided by zero).
let g = ChunkGrid::fixed_array(&[1, 0], None, &[6, 6]).unwrap();
for i in 0..16 {
assert_eq!(g.offsets(i), None);
}
let g = ChunkGrid::fixed_array(&[0, 0, 3], Some(&[4, 0, 3]), &[2, 2, 3]).unwrap();
for i in 0..16 {
assert_eq!(g.offsets(i), None);
}
let g = ChunkGrid::extensible_array(&[0, 5], Some(&[u64::MAX, 0]), &[2, 2]).unwrap();
for i in 0..16 {
assert_eq!(g.offsets(i), None);
}
// A zero last dimension leaves the other strides alone.
let g = ChunkGrid::fixed_array(&[4, 0], Some(&[4, 6]), &[2, 3]).unwrap();
assert_eq!(g.offsets(0), None);
assert_eq!(g.linear_index(&[1, 1]), 3);
}
#[test] #[test]
fn rejects_two_unlimited_dims_after_the_first() { fn rejects_two_unlimited_dims_after_the_first() {
assert!(ChunkGrid::fixed_array(&[4, 6], Some(&[u64::MAX, u64::MAX]), &[2, 3]).is_err()); assert!(ChunkGrid::fixed_array(&[4, 6], Some(&[u64::MAX, u64::MAX]), &[2, 3]).is_err());
+16
View File
@@ -130,6 +130,22 @@ pub(crate) fn build_chunked_dataset_oh(
fill_message: &[u8], fill_message: &[u8],
refcount: u32, refcount: u32,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
// libhdf5 records every simple dataspace's maximum dimensions (the
// dimensions themselves when none are given, `H5S_set_extent_simple`).
// Without them libhdf5 takes the maximum to be the current dimensions,
// so a resize by libhdf5 (h5py's `Dataset.resize`) would also change
// the maximum a Fixed Array chunk index is laid out by, and move every
// chunk already written.
let recorded;
let ds = if ds.space_type == DataspaceType::Simple && ds.max_dimensions.is_none() {
recorded = Dataspace {
max_dimensions: Some(ds.dimensions.clone()),
..ds.clone()
};
&recorded
} else {
ds
};
let mut w = ObjectHeaderWriter::new(); let mut w = ObjectHeaderWriter::new();
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01); w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE)); w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
+17
View File
@@ -39,6 +39,23 @@ impl MmapReader {
Ok(Self { _file: file, mmap }) Ok(Self { _file: file, mmap })
} }
/// Memory-map a file that is already open (for reading).
///
/// The mapping references `file`'s open file description for as long as
/// it lives, so a `flock` taken through that description (or a
/// `try_clone` of it) is held until the reader is dropped.
///
/// # Safety
///
/// The same contract as [`open`](Self::open): the file must not be
/// modified while the mapping is active.
pub fn from_file(file: fs::File) -> io::Result<Self> {
// SAFETY: a read-only mapping; the caller keeps the file unmodified
// while it is alive.
let mmap = unsafe { Mmap::map(&file)? };
Ok(Self { _file: file, mmap })
}
/// Zero-copy access to the entire file contents. /// Zero-copy access to the entire file contents.
pub fn as_bytes(&self) -> &[u8] { pub fn as_bytes(&self) -> &[u8] {
&self.mmap &self.mmap
+9
View File
@@ -17,11 +17,20 @@ crate-type = ["cdylib", "rlib"]
[dependencies] [dependencies]
clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" } clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" }
clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" } clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
# Remote files (`clawhdf5.File(url)`): plain HTTP by default, which builds
# no C. HTTPS and the object stores are opt-in features below.
clawhdf5-remote = { path = "../clawhdf5-remote", version = "2.7.0" }
pyo3 = "0.29" pyo3 = "0.29"
numpy = "0.29" numpy = "0.29"
[features] [features]
extension-module = ["pyo3/extension-module"] extension-module = ["pyo3/extension-module"]
# https:// URLs (rustls with ring, which compiles C and assembly).
https = ["clawhdf5-remote/https"]
# s3://, gs://, az:// URLs (object_store; its cloud clients build aws-lc-rs, C).
s3 = ["clawhdf5-remote/s3"]
gcs = ["clawhdf5-remote/gcs"]
azure = ["clawhdf5-remote/azure"]
[package.metadata.docs.rs] [package.metadata.docs.rs]
features = [] features = []
+74 -3
View File
@@ -43,8 +43,9 @@ with clawhdf5.File("data.h5", "r") as f:
Other types raise `TypeError`. Other types raise `TypeError`.
- Keys are h5py's: integers, slices with a positive step, `...`, one - Keys are h5py's: integers, slices with a positive step, `...`, one
increasing list of integers, compound field names. Each maps onto a increasing list of integers, compound field names. Each maps onto a
hyperslab selection. `None`, negative steps and boolean masks are refused hyperslab selection. `None` and negative steps are refused
with h5py's errors. with h5py's errors; boolean masks (which h5py supports) raise
`NotImplementedError`, for reads and writes.
- What is read from the file: a selection whose bounding box covers at - What is read from the file: a selection whose bounding box covers at
most half the dataset decodes only the chunks (or contiguous rows) the box most half the dataset decodes only the chunks (or contiguous rows) the box
overlaps. The library decodes the whole dataset for a larger box overlaps. The library decodes the whole dataset for a larger box
@@ -61,6 +62,36 @@ with clawhdf5.File("data.h5", "r") as f:
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null - Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
dataspace (h5py's `Empty`). dataspace (h5py's `Empty`).
## Remote files
A URL instead of a path reads the file where it is, through
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
```python
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
```
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise `OSError` instead of mixing versions.
Network failures are `OSError` too.
- Remote files are read-only.
- Schemes: the default build (no C) reads `http://`. `https://` needs
`maturin develop --release --features https` (rustls with ring, which
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
`azure` features (credentials from the environment; aws-lc-rs, C).
## Writing ## Writing
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array, `clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
@@ -68,6 +99,41 @@ chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
written on `close()`. written on `close()`.
## Editing a file in place
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
where it is, through clawhdf5's `FileEditor`; the file is locked until
`close()`, and every edit is written and synced before the statement
returns.
```python
with clawhdf5.File("data.h5", "r+") as f:
ds = f["grid"]
ds[10:20, ::2] = 0 # h5py keys and broadcasting
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
f["series"].attrs["units"] = "K"
f.attrs.create("version", 2, dtype="u1")
```
- Values: a numpy array is converted to the dataset's dtype as libhdf5
converts it (integers saturate at the target's limits; floats are
truncated toward zero and clipped); anything else goes through
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
value). A few libhdf5 edge cases differ on purpose; see
`docs/known-issues.md`.
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
see its result.
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
`str` is stored as a fixed-length UTF-8 string (h5py stores a
variable-length one), so h5py reads it back as `bytes`.
- Not supported (`NotImplementedError`, nothing written): creating or
deleting datasets, groups and attributes, writing compound fields by
name, variable-length data, HDF5 array types, and whatever
`FileEditor` refuses (listed in `docs/known-issues.md`).
## Tests ## Tests
```bash ```bash
@@ -76,7 +142,12 @@ pytest crates/clawhdf5-py/tests
``` ```
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py `tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
writes. `scripts/ci-test.sh` builds the wheel and runs these in CI. writes, opened locally and over HTTP (an in-process range server,
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
the wheel and runs these in CI.
## License ## License
+166 -65
View File
@@ -1,22 +1,62 @@
//! PyAttrs — dict-like access to HDF5 attributes. //! PyAttrs — dict-like access to HDF5 attributes.
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex, PoisonError};
use clawhdf5_format::attribute::AttributeMessage; use clawhdf5_format::attribute::AttributeMessage;
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError}; use pyo3::exceptions::{PyKeyError, PyNotImplementedError, PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple}; use pyo3::types::{PyList, PyTuple};
use crate::convert::{Converter, Elements, resolve_vl}; use crate::convert::{Converter, Elements, resolve_vl};
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value}; use crate::handle::Handle;
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, edit, node, py_to_attr_value};
/// The attributes of an object in a file opened for reading (or editing).
struct ReadAttrs {
handle: Arc<Handle>,
addr: u64,
path: String,
/// Sorted by name, with the file generation they were read at: an edit
/// (`attrs[name] = value`, here or through another handle on the same
/// object) makes them re-read.
cache: Mutex<(u64, Arc<Vec<AttributeMessage>>)>,
}
impl ReadAttrs {
fn current(&self, py: Python<'_>) -> PyResult<Arc<Vec<AttributeMessage>>> {
let generation = self.handle.generation();
{
let cached = self.cache.lock().unwrap_or_else(PoisonError::into_inner);
if cached.0 == generation {
return Ok(Arc::clone(&cached.1));
}
}
let (addr, path) = (self.addr, &self.path);
let attrs = Arc::new(self.handle.with(py, |f| node::attributes(f, addr, path))?);
*self.cache.lock().unwrap_or_else(PoisonError::into_inner) =
(generation, Arc::clone(&attrs));
Ok(attrs)
}
fn check_writable(&self) -> PyResult<()> {
if self.handle.is_writable() {
return Ok(());
}
Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"cannot set attributes on a read-only file (open it with mode 'r+')",
))
}
fn set(&self, py: Python<'_>, name: &str, value: clawhdf5_rs::AttrValue) -> PyResult<()> {
self.check_writable()?;
let path = node::name(&self.path);
self.handle.edit(py, |ed| ed.set_attr(&path, name, &value))
}
}
/// Backing storage for attributes. /// Backing storage for attributes.
enum AttrsInner { enum AttrsInner {
/// Attributes of an object in a file opened for reading, sorted by name. Read(ReadAttrs),
Read {
file: Arc<clawhdf5_rs::File>,
attrs: Vec<AttributeMessage>,
},
/// Writable attribute list shared with a parent (PyFile or PyGroup). /// Writable attribute list shared with a parent (PyFile or PyGroup).
Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>), Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>),
} }
@@ -26,8 +66,11 @@ enum AttrsInner {
/// In read mode, values are what h5py returns: numpy scalars for scalar /// In read mode, values are what h5py returns: numpy scalars for scalar
/// attributes, numpy arrays otherwise, `str` for variable-length strings, /// attributes, numpy arrays otherwise, `str` for variable-length strings,
/// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace. /// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace.
/// In write mode, attributes set here are accumulated and written when /// In a file opened with `'r+'`, `attrs[name] = value` adds or replaces an
/// the parent file is closed. /// attribute in the file at once (as h5py stores it, except that `str`
/// values become fixed-length UTF-8 strings). In write mode (`'w'`),
/// attributes set here are accumulated and written when the parent file is
/// closed.
#[pyclass(name = "Attrs")] #[pyclass(name = "Attrs")]
pub struct PyAttrs { pub struct PyAttrs {
inner: AttrsInner, inner: AttrsInner,
@@ -36,10 +79,21 @@ pub struct PyAttrs {
impl PyAttrs { impl PyAttrs {
/// The attributes of the object at `addr` (whose path is `path`) in a /// The attributes of the object at `addr` (whose path is `path`) in a
/// file opened for reading. /// file opened for reading.
pub(crate) fn read(file: Arc<clawhdf5_rs::File>, addr: u64, path: &str) -> PyResult<Self> { pub(crate) fn read(
let attrs = node::attributes(&file, addr, path)?; py: Python<'_>,
handle: Arc<Handle>,
addr: u64,
path: &str,
) -> PyResult<Self> {
let generation = handle.generation();
let attrs = Arc::new(handle.with(py, |f| node::attributes(f, addr, path))?);
Ok(Self { Ok(Self {
inner: AttrsInner::Read { file, attrs }, inner: AttrsInner::Read(ReadAttrs {
handle,
addr,
path: path.to_string(),
cache: Mutex::new((generation, attrs)),
}),
}) })
} }
@@ -49,14 +103,48 @@ impl PyAttrs {
inner: AttrsInner::Write(store), inner: AttrsInner::Write(store),
} }
} }
fn set_value(
&self,
py: Python<'_>,
key: &str,
value: &Bound<'_, PyAny>,
dtype: Option<&Bound<'_, PyAny>>,
shape: Option<&Bound<'_, PyAny>>,
) -> PyResult<()> {
match &self.inner {
AttrsInner::Read(r) => {
r.check_writable()?;
let value = edit::attr_value(py, value, dtype, shape)?;
r.set(py, key, value)
}
AttrsInner::Write(store) => {
if dtype.is_some() || shape.is_some() {
return Err(PyNotImplementedError::new_err(
"attrs.create with a dtype or shape is only supported in a file opened \
with 'r+'",
));
}
let owned = py_to_attr_value(value)?;
let mut guard = store.lock().unwrap();
// Replace existing key if present.
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
entry.1 = owned;
} else {
guard.push((key.to_string(), owned));
}
Ok(())
}
}
}
} }
#[pymethods] #[pymethods]
impl PyAttrs { impl PyAttrs {
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> { fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
match &self.inner { match &self.inner {
AttrsInner::Read { file, attrs } => match attrs.iter().find(|a| a.name == key) { AttrsInner::Read(r) => match r.current(py)?.iter().find(|a| a.name == key) {
Some(attr) => Ok(attr_to_py(py, file, attr)?.unbind()), Some(attr) => Ok(attr_to_py(py, &r.handle, attr)?.unbind()),
None => Err(PyKeyError::new_err(format!( None => Err(PyKeyError::new_err(format!(
"Can't open attribute (can't locate attribute: '{key}')" "Can't open attribute (can't locate attribute: '{key}')"
))), ))),
@@ -74,36 +162,63 @@ impl PyAttrs {
} }
} }
fn __setitem__(&self, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> { /// `attrs[name] = value`. In a file opened with `'r+'` this writes the
/// attribute (numeric, bool, complex, bytes and str data, any shape)
/// into the file before returning; see the class docs.
fn __setitem__(&self, py: Python<'_>, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
self.set_value(py, key, value, None, None)
}
/// Deleting attributes is not supported in a file (the in-place editor
/// cannot remove them); in write mode it removes a pending attribute.
fn __delitem__(&self, key: &str) -> PyResult<()> {
match &self.inner { match &self.inner {
AttrsInner::Read { .. } => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>( AttrsInner::Read(_) => Err(PyNotImplementedError::new_err(format!(
"cannot set attributes on a read-only file", "cannot delete attribute '{key}': deleting attributes is not supported by \
)), clawhdf5's in-place editor"
))),
AttrsInner::Write(store) => { AttrsInner::Write(store) => {
let owned = py_to_attr_value(value)?;
let mut guard = store.lock().unwrap(); let mut guard = store.lock().unwrap();
// Replace existing key if present. let before = guard.len();
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) { guard.retain(|(k, _)| k != key);
entry.1 = owned; if guard.len() == before {
} else { return Err(PyKeyError::new_err(key.to_string()));
guard.push((key.to_string(), owned));
} }
Ok(()) Ok(())
} }
} }
} }
fn __len__(&self) -> usize { /// h5py's `attrs.create(name, data, shape=None, dtype=None)`: `data`
/// converted to `dtype` and reshaped to `shape` first.
#[pyo3(signature = (name, data, shape=None, dtype=None))]
fn create(
&self,
py: Python<'_>,
name: &str,
data: &Bound<'_, PyAny>,
shape: Option<&Bound<'_, PyAny>>,
dtype: Option<&Bound<'_, PyAny>>,
) -> PyResult<()> {
self.set_value(py, name, data, dtype, shape)
}
/// h5py's `attrs.modify(name, value)`: same as `attrs[name] = value`.
fn modify(&self, py: Python<'_>, name: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
self.set_value(py, name, value, None, None)
}
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match &self.inner { match &self.inner {
AttrsInner::Read { attrs, .. } => attrs.len(), AttrsInner::Read(r) => Ok(r.current(py)?.len()),
AttrsInner::Write(store) => store.lock().unwrap().len(), AttrsInner::Write(store) => Ok(store.lock().unwrap().len()),
} }
} }
fn __contains__(&self, key: &str) -> bool { fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
match &self.inner { match &self.inner {
AttrsInner::Read { attrs, .. } => attrs.iter().any(|a| a.name == key), AttrsInner::Read(r) => Ok(r.current(py)?.iter().any(|a| a.name == key)),
AttrsInner::Write(store) => store.lock().unwrap().iter().any(|(k, _)| k == key), AttrsInner::Write(store) => Ok(store.lock().unwrap().iter().any(|(k, _)| k == key)),
} }
} }
@@ -113,15 +228,17 @@ impl PyAttrs {
Ok(iter) Ok(iter)
} }
fn __repr__(&self) -> String { fn __repr__(&self, py: Python<'_>) -> String {
let n = self.__len__(); match self.__len__(py) {
format!("<HDF5 Attrs ({n} members)>") Ok(n) => format!("<HDF5 Attrs ({n} members)>"),
Err(_) => "<HDF5 Attrs>".to_string(),
}
} }
/// The value of `key`, or `default` if there is no such attribute. /// The value of `key`, or `default` if there is no such attribute.
#[pyo3(signature = (key, default=None))] #[pyo3(signature = (key, default=None))]
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> { fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
if self.__contains__(key) { if self.__contains__(py, key)? {
self.__getitem__(py, key) self.__getitem__(py, key)
} else { } else {
Ok(default.unwrap_or_else(|| py.None())) Ok(default.unwrap_or_else(|| py.None()))
@@ -131,7 +248,7 @@ impl PyAttrs {
/// Return attribute names as a list. /// Return attribute names as a list.
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let names: Vec<String> = match &self.inner { let names: Vec<String> = match &self.inner {
AttrsInner::Read { attrs, .. } => attrs.iter().map(|a| a.name.clone()).collect(), AttrsInner::Read(r) => r.current(py)?.iter().map(|a| a.name.clone()).collect(),
AttrsInner::Write(store) => store AttrsInner::Write(store) => store
.lock() .lock()
.unwrap() .unwrap()
@@ -146,9 +263,10 @@ impl PyAttrs {
/// Return attribute values as a list. /// Return attribute values as a list.
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let vals: Vec<Py<PyAny>> = match &self.inner { let vals: Vec<Py<PyAny>> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs AttrsInner::Read(r) => r
.current(py)?
.iter() .iter()
.map(|a| attr_to_py(py, file, a).map(Bound::unbind)) .map(|a| attr_to_py(py, &r.handle, a).map(Bound::unbind))
.collect::<PyResult<_>>()?, .collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store AttrsInner::Write(store) => store
.lock() .lock()
@@ -167,9 +285,10 @@ impl PyAttrs {
/// Return attribute (key, value) pairs as a list of tuples. /// Return attribute (key, value) pairs as a list of tuples.
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner { let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs AttrsInner::Read(r) => r
.current(py)?
.iter() .iter()
.map(|a| Ok((a.name.clone(), attr_to_py(py, file, a)?.unbind()))) .map(|a| Ok((a.name.clone(), attr_to_py(py, &r.handle, a)?.unbind())))
.collect::<PyResult<_>>()?, .collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store AttrsInner::Write(store) => store
.lock() .lock()
@@ -189,12 +308,11 @@ impl PyAttrs {
/// An attribute's value as h5py returns it. /// An attribute's value as h5py returns it.
fn attr_to_py<'py>( fn attr_to_py<'py>(
py: Python<'py>, py: Python<'py>,
file: &clawhdf5_rs::File, handle: &Handle,
attr: &AttributeMessage, attr: &AttributeMessage,
) -> PyResult<Bound<'py, PyAny>> { ) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| { crate::no_panic(|| {
let sb = file.superblock(); let conv = Converter::new(py, &attr.datatype, handle.offset_size)
let conv = Converter::new(py, &attr.datatype, sb.offset_size)
.map_err(|e| prefix_err(py, &attr.name, e))?; .map_err(|e| prefix_err(py, &attr.name, e))?;
if node::is_null(&attr.dataspace) { if node::is_null(&attr.dataspace) {
return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any()); return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any());
@@ -216,12 +334,11 @@ fn attr_to_py<'py>(
))); )));
} }
let raw = &attr.raw_data[..want]; let raw = &attr.raw_data[..want];
let file_data = file.as_bytes(); let (osz, lsz, unit) = (handle.offset_size, handle.length_size, conv.vl_unit);
let (osz, lsz, unit) = (sb.offset_size, sb.length_size, conv.vl_unit); let what = format!("attribute {}", attr.name);
Elements::Vl( Elements::Vl(handle.with(py, |f| {
py.detach(|| resolve_vl(file_data, raw, n, osz, lsz, unit)) resolve_vl(f.storage(), raw, n, osz, lsz, unit).map_err(|e| e.into_py(&what))
.map_err(|e| PyValueError::new_err(format!("attribute {}: {e}", attr.name)))?, })?)
)
} else { } else {
Elements::Bytes(attr.raw_data.clone()) Elements::Bytes(attr.raw_data.clone())
}; };
@@ -244,19 +361,3 @@ fn prefix_err(py: Python<'_>, name: &str, e: PyErr) -> PyErr {
PyValueError::new_err(msg) PyValueError::new_err(msg)
} }
} }
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn write_attrs_len() {
let store = Arc::new(Mutex::new(Vec::new()));
store
.lock()
.unwrap()
.push(("key".into(), OwnedAttrValue::I64(99)));
let attrs = PyAttrs::from_write(store);
assert_eq!(attrs.__len__(), 1);
}
}
+49 -21
View File
@@ -17,6 +17,7 @@ use std::collections::HashMap;
use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder}; use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder};
use clawhdf5_format::global_heap::GlobalHeapCollection; use clawhdf5_format::global_heap::GlobalHeapCollection;
use clawhdf5_format::storage::Storage;
use numpy::PyArray1; use numpy::PyArray1;
use pyo3::exceptions::{PyTypeError, PyValueError}; use pyo3::exceptions::{PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
@@ -538,16 +539,16 @@ fn object_array<'py>(
/// their bytes: each element's stored length times `unit` (1 for strings, /// their bytes: each element's stored length times `unit` (1 for strings,
/// the base type's size for sequences). Pure Rust, so it runs without the /// the base type's size for sequences). Pure Rust, so it runs without the
/// GIL. /// GIL.
pub(crate) fn resolve_vl( pub(crate) fn resolve_vl<S: Storage + ?Sized>(
file_data: &[u8], file: &S,
raw: &[u8], raw: &[u8],
count: usize, count: usize,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
unit: usize, unit: usize,
) -> Result<Vec<Vec<u8>>, String> { ) -> Result<Vec<Vec<u8>>, VlError> {
let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size) let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size)
.map_err(|e| e.to_string())?; .map_err(|e| VlError::Invalid(e.to_string()))?;
let undefined = match offset_size { let undefined = match offset_size {
2 => 0xFFFF, 2 => 0xFFFF,
4 => 0xFFFF_FFFF, 4 => 0xFFFF_FFFF,
@@ -558,43 +559,70 @@ pub(crate) fn resolve_vl(
for vl in &refs { for vl in &refs {
if vl.collection_address == 0 || vl.collection_address == undefined { if vl.collection_address == 0 || vl.collection_address == undefined {
if vl.length != 0 { if vl.length != 0 {
return Err(format!( return Err(VlError::Invalid(format!(
"variable-length element of length {} has no heap address", "variable-length element of length {} has no heap address",
vl.length vl.length
)); )));
} }
out.push(Vec::new()); out.push(Vec::new());
continue; continue;
} }
let coll = match collections.entry(vl.collection_address) { let coll = match collections.entry(vl.collection_address) {
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(), std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
std::collections::hash_map::Entry::Vacant(e) => { std::collections::hash_map::Entry::Vacant(e) => e.insert(
let addr = usize::try_from(vl.collection_address) GlobalHeapCollection::parse_in(file, vl.collection_address, length_size)
.map_err(|_| "global heap address out of range".to_string())?; .map_err(VlError::from_format)?,
e.insert( ),
GlobalHeapCollection::parse(file_data, addr, length_size)
.map_err(|e| e.to_string())?,
)
}
}; };
let index = u16::try_from(vl.object_index) let index = u16::try_from(vl.object_index).map_err(|_| {
.map_err(|_| format!("global heap object index {} out of range", vl.object_index))?; VlError::Invalid(format!(
"global heap object index {} out of range",
vl.object_index
))
})?;
let obj = coll.get_object(index).ok_or_else(|| { let obj = coll.get_object(index).ok_or_else(|| {
format!( VlError::Invalid(format!(
"global heap object {index} not found in the collection at {}", "global heap object {index} not found in the collection at {}",
vl.collection_address vl.collection_address
) ))
})?; })?;
let need = (vl.length as usize) let need = (vl.length as usize)
.checked_mul(unit) .checked_mul(unit)
.ok_or("variable-length element too long")?; .ok_or_else(|| VlError::Invalid("variable-length element too long".into()))?;
if need > obj.data.len() { if need > obj.data.len() {
return Err(format!( return Err(VlError::Invalid(format!(
"variable-length element of {need} bytes in a {}-byte heap object", "variable-length element of {need} bytes in a {}-byte heap object",
obj.data.len() obj.data.len()
)); )));
} }
out.push(obj.data[..need].to_vec()); out.push(obj.data[..need].to_vec());
} }
Ok(out) Ok(out)
} }
/// Why variable-length elements could not be resolved.
#[derive(Debug)]
pub(crate) enum VlError {
/// Reading the file failed (a network error on a remote file).
Storage(String),
/// The references or the heap are not valid.
Invalid(String),
}
impl VlError {
fn from_format(e: clawhdf5_format::error::FormatError) -> Self {
match e {
clawhdf5_format::error::FormatError::Storage(_) => VlError::Storage(e.to_string()),
e => VlError::Invalid(e.to_string()),
}
}
/// As a Python exception, the message prefixed with `what`: a storage
/// failure is an `OSError`, anything else a `ValueError`.
pub(crate) fn into_py(self, what: &str) -> PyErr {
match self {
VlError::Storage(m) => pyo3::exceptions::PyOSError::new_err(format!("{what}: {m}")),
VlError::Invalid(m) => PyValueError::new_err(format!("{what}: {m}")),
}
}
}
+269 -81
View File
@@ -6,20 +6,52 @@
//! whole dataset instead); the //! whole dataset instead); the
//! bytes it returns become the numpy array's buffer without a copy (see //! bytes it returns become the numpy array's buffer without a copy (see
//! `convert`). All file access and decoding runs with the GIL released, so //! `convert`). All file access and decoding runs with the GIL released, so
//! Python threads reading the same or different datasets run in parallel. //! Python threads reading the same or different datasets run in parallel,
//! and a remote file's network reads never hold the GIL.
use std::sync::Arc; use std::sync::{Arc, Mutex, PoisonError};
use clawhdf5_format::datatype::Datatype; use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::object_header::ObjectHeader; use clawhdf5_format::object_header::ObjectHeader;
use pyo3::exceptions::{PyTypeError, PyValueError}; use clawhdf5_rs::File;
use pyo3::exceptions::{PyNotImplementedError, PyOSError, PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple}; use pyo3::types::{PyList, PyTuple};
use crate::attrs::PyAttrs; use crate::attrs::PyAttrs;
use crate::convert::{Converter, Elements, resolve_vl}; use crate::convert::{Converter, Elements, VlError, resolve_vl};
use crate::handle::Handle;
use crate::select::{self, Plan}; use crate::select::{self, Plan};
use crate::{PyEmpty, node, to_py_err}; use crate::{PyEmpty, edit, node, to_py_err};
/// What opening a dataset reads from the file (without the GIL).
pub(crate) struct DatasetMeta {
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
shape: Option<Vec<u64>>,
chunks: Option<Vec<u64>>,
datatype: Datatype,
}
impl DatasetMeta {
pub(crate) fn load(f: &File, addr: u64, hdr: &ObjectHeader, path: &str) -> PyResult<Self> {
let null = node::is_null(&node::dataspace(f, hdr, path)?);
let ds = f.dataset_at(addr).map_err(to_py_err)?;
let shape = if null {
None
} else {
Some(ds.shape().map_err(to_py_err)?)
};
let datatype = ds.raw_datatype().map_err(to_py_err)?;
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(f, hdr, s.len()));
Ok(Self {
shape,
chunks,
datatype,
})
}
}
/// A dataset in a file opened for reading. /// A dataset in a file opened for reading.
/// ///
@@ -30,13 +62,14 @@ use crate::{PyEmpty, node, to_py_err};
/// ``` /// ```
#[pyclass(name = "Dataset")] #[pyclass(name = "Dataset")]
pub struct PyDataset { pub struct PyDataset {
file: Arc<clawhdf5_rs::File>, handle: Arc<Handle>,
path: String, path: String,
/// Where the dataset's object header is: reads open it from here rather /// Where the dataset's object header is: reads open it from here rather
/// than resolve `path` again. /// than resolve `path` again.
addr: u64, addr: u64,
/// `None` for a dataset with a null dataspace (h5py's `Empty`). /// The shape (`None` for a null dataspace, h5py's `Empty`), with the
shape: Option<Vec<u64>>, /// file generation it was read at: an edit (a resize) may change it.
shape: Mutex<(u64, Option<Vec<u64>>)>,
/// The chunk shape, for a chunked dataset. /// The chunk shape, for a chunked dataset.
chunks: Option<Vec<u64>>, chunks: Option<Vec<u64>>,
datatype: Datatype, datatype: Datatype,
@@ -45,39 +78,25 @@ pub struct PyDataset {
} }
impl PyDataset { impl PyDataset {
pub(crate) fn open( pub(crate) fn new(
py: Python<'_>, py: Python<'_>,
file: Arc<clawhdf5_rs::File>, handle: Arc<Handle>,
path: String, path: String,
addr: u64, addr: u64,
hdr: &ObjectHeader, meta: DatasetMeta,
) -> PyResult<Self> { ) -> Self {
crate::no_panic(|| { let generation = handle.generation();
let null = node::is_null(&node::dataspace(&file, hdr)?); let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
let (shape, datatype) = { .map_err(|e| e.value(py).to_string());
let ds = file.dataset_at(addr).map_err(to_py_err)?; Self {
let shape = if null { handle,
None path,
} else { addr,
Some(ds.shape().map_err(to_py_err)?) shape: Mutex::new((generation, meta.shape)),
}; chunks: meta.chunks,
(shape, ds.raw_datatype().map_err(to_py_err)?) datatype: meta.datatype,
}; conv,
let conv = Converter::new(py, &datatype, file.superblock().offset_size) }
.map_err(|e| e.value(py).to_string());
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(&file, hdr, s.len()));
Ok(Self {
file,
path,
addr,
shape,
chunks,
datatype,
conv,
})
})
} }
fn converter(&self) -> PyResult<&Converter> { fn converter(&self) -> PyResult<&Converter> {
@@ -86,10 +105,54 @@ impl PyDataset {
.map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path)))) .map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path))))
} }
/// The current shape: the one read at open, or re-read after an edit.
fn dims(&self, py: Python<'_>) -> PyResult<Option<Vec<u64>>> {
let generation = self.handle.generation();
{
let cached = self.shape.lock().unwrap_or_else(PoisonError::into_inner);
if cached.0 == generation {
return Ok(cached.1.clone());
}
}
let addr = self.addr;
let null = self
.shape
.lock()
.unwrap_or_else(PoisonError::into_inner)
.1
.is_none();
let shape = if null {
None
} else {
Some(self.handle.with(py, |f| {
f.dataset_at(addr)
.and_then(|ds| ds.shape())
.map_err(to_py_err)
})?)
};
*self.shape.lock().unwrap_or_else(PoisonError::into_inner) = (generation, shape.clone());
Ok(shape)
}
fn check_writable(&self) -> PyResult<()> {
if self.handle.is_writable() {
Ok(())
} else {
Err(PyOSError::new_err(format!(
"{}: the file is open read-only; open it with mode 'r+' to change it",
node::name(&self.path)
)))
}
}
/// Read the selection described by `plan` into a numpy array. /// Read the selection described by `plan` into a numpy array.
fn read_plan<'py>(&self, py: Python<'py>, plan: &Plan) -> PyResult<Bound<'py, PyAny>> { fn read_plan<'py>(
&self,
py: Python<'py>,
plan: &Plan,
dims: &[u64],
) -> PyResult<Bound<'py, PyAny>> {
let conv = self.converter()?; let conv = self.converter()?;
let dims = self.shape.as_deref().unwrap_or(&[]);
let out_shape = plan.out_shape(); let out_shape = plan.out_shape();
let arr = if plan.is_empty() { let arr = if plan.is_empty() {
@@ -103,10 +166,10 @@ impl PyDataset {
}; };
let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size); let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size);
let read_shape = plan.read_shape(); let read_shape = plan.read_shape();
let file = &*self.file; let handle = &*self.handle;
let addr = self.addr; let addr = self.addr;
// Everything below touches only Rust data: release the GIL. // Everything below touches only Rust data: release the GIL.
let read = || -> Result<Elements, ReadError> { let read = |file: &File| -> Result<Elements, ReadError> {
let ds = file.dataset_at(addr)?; let ds = file.dataset_at(addr)?;
let mut blocks = Vec::with_capacity(reads.len()); let mut blocks = Vec::with_capacity(reads.len());
for read in reads { for read in reads {
@@ -147,7 +210,7 @@ impl PyDataset {
let sb = file.superblock(); let sb = file.superblock();
let n = read_shape.iter().product(); let n = read_shape.iter().product();
resolve_vl( resolve_vl(
file.as_bytes(), file.storage(),
&raw, &raw,
n, n,
sb.offset_size, sb.offset_size,
@@ -155,12 +218,20 @@ impl PyDataset {
unit, unit,
) )
.map(Elements::Vl) .map(Elements::Vl)
.map_err(ReadError::Other) .map_err(ReadError::Vl)
}; };
let data = py let data = py
.detach(|| { .detach(|| {
std::panic::catch_unwind(std::panic::AssertUnwindSafe(read)) handle
.unwrap_or_else(|p| Err(ReadError::Panic(crate::panic_text(&*p)))) .with_detached(|f| {
Ok(
std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| read(f)))
.unwrap_or_else(|p| {
Err(ReadError::Panic(crate::panic_text(&*p)))
}),
)
})
.unwrap_or_else(|e| Err(ReadError::Py(e)))
}) })
.map_err(|e| e.into_py(&self.path))?; .map_err(|e| e.into_py(&self.path))?;
let joined = conv.to_array(py, data, &read_shape, false)?; let joined = conv.to_array(py, data, &read_shape, false)?;
@@ -183,6 +254,8 @@ impl PyDataset {
/// An error from the read closure, turned into a Python error with the GIL. /// An error from the read closure, turned into a Python error with the GIL.
enum ReadError { enum ReadError {
Lib(clawhdf5_rs::Error), Lib(clawhdf5_rs::Error),
Vl(VlError),
Py(PyErr),
Other(String), Other(String),
Panic(String), Panic(String),
} }
@@ -197,6 +270,8 @@ impl ReadError {
fn into_py(self, path: &str) -> PyErr { fn into_py(self, path: &str) -> PyErr {
match self { match self {
ReadError::Lib(e) => to_py_err(e), ReadError::Lib(e) => to_py_err(e),
ReadError::Vl(e) => e.into_py(&node::name(path)),
ReadError::Py(e) => e,
ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))), ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))),
ReadError::Panic(msg) => crate::InternalError::new_err(format!( ReadError::Panic(msg) => crate::InternalError::new_err(format!(
"{}: clawhdf5 internal error (please report it): {msg}", "{}: clawhdf5 internal error (please report it): {msg}",
@@ -243,7 +318,7 @@ impl PyDataset {
/// The shape of the dataset (`None` for an empty/null dataspace). /// The shape of the dataset (`None` for an empty/null dataspace).
#[getter] #[getter]
fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> { fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
match &self.shape { match self.dims(py)? {
Some(s) => Ok(PyTuple::new(py, s)?.into_any()), Some(s) => Ok(PyTuple::new(py, s)?.into_any()),
None => Ok(py.None().into_bound(py)), None => Ok(py.None().into_bound(py)),
} }
@@ -252,22 +327,32 @@ impl PyDataset {
/// The maximum shape (`None` per unlimited dimension), like h5py. /// The maximum shape (`None` per unlimited dimension), like h5py.
#[getter] #[getter]
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> { fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| { let Some(shape) = self.dims(py)? else {
let Some(shape) = &self.shape else { return Ok(py.None().into_bound(py));
return Ok(py.None().into_bound(py)); };
}; let addr = self.addr;
let max = self let max = self
.file .handle
.dataset_at(self.addr) .with(py, |f| {
.and_then(|ds| ds.max_dimensions()) f.dataset_at(addr)
.map_err(to_py_err)? .and_then(|ds| ds.max_dimensions())
.unwrap_or_else(|| shape.clone()); .map_err(to_py_err)
let items: Vec<Option<u64>> = max })?
.into_iter() .unwrap_or(shape);
.map(|d| (d != u64::MAX).then_some(d)) let items: Vec<Option<u64>> = max
.collect(); .into_iter()
Ok(PyTuple::new(py, items)?.into_any()) .map(|d| (d != u64::MAX).then_some(d))
}) .collect();
Ok(PyTuple::new(py, items)?.into_any())
}
/// The chunk shape, or `None` for a dataset that is not chunked.
#[getter]
fn chunks<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
match &self.chunks {
Some(c) => Ok(PyTuple::new(py, c)?.into_any()),
None => Ok(py.None().into_bound(py)),
}
} }
/// The dataset's numpy dtype, as h5py reports it. /// The dataset's numpy dtype, as h5py reports it.
@@ -277,14 +362,14 @@ impl PyDataset {
} }
#[getter] #[getter]
fn ndim(&self) -> usize { fn ndim(&self, py: Python<'_>) -> PyResult<usize> {
self.shape.as_ref().map_or(0, Vec::len) Ok(self.dims(py)?.map_or(0, |s| s.len()))
} }
/// Number of elements (`None` for an empty/null dataspace, as h5py). /// Number of elements (`None` for an empty/null dataspace, as h5py).
#[getter] #[getter]
fn size(&self) -> Option<u64> { fn size(&self, py: Python<'_>) -> PyResult<Option<u64>> {
self.shape.as_ref().map(|s| s.iter().product()) Ok(self.dims(py)?.map(|s| s.iter().product()))
} }
/// The dataset's full name, e.g. `/group/data`. /// The dataset's full name, e.g. `/group/data`.
@@ -293,10 +378,11 @@ impl PyDataset {
node::name(&self.path) node::name(&self.path)
} }
/// The dataset's attributes (read-only, dict-like). /// The dataset's attributes (dict-like; writable in a file opened with
/// `'r+'`).
#[getter] #[getter]
fn attrs(&self) -> PyResult<PyAttrs> { fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path) PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
} }
/// Read with h5py indexing: integers, slices with positive steps, /// Read with h5py indexing: integers, slices with positive steps,
@@ -308,7 +394,7 @@ impl PyDataset {
py: Python<'py>, py: Python<'py>,
key: &Bound<'py, PyAny>, key: &Bound<'py, PyAny>,
) -> PyResult<Bound<'py, PyAny>> { ) -> PyResult<Bound<'py, PyAny>> {
let Some(dims) = &self.shape else { let Some(dims) = self.dims(py)? else {
let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty()); let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty());
let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>(); let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>();
if is_empty_tuple || is_ellipsis { if is_empty_tuple || is_ellipsis {
@@ -317,8 +403,109 @@ impl PyDataset {
} }
return Err(PyValueError::new_err("Empty datasets cannot be sliced")); return Err(PyValueError::new_err("Empty datasets cannot be sliced"));
}; };
let plan = select::parse(key, dims)?; let plan = select::parse(key, &dims)?;
self.read_plan(py, &plan) self.read_plan(py, &plan, &dims)
}
/// Write with h5py indexing (file opened with `'r+'`): `ds[key] = value`.
///
/// The key is what `ds[key]` reads (without compound field names). The
/// value is converted to the dataset's dtype as h5py converts it (a
/// numpy array as libhdf5 does, clipping out-of-range numbers; anything
/// else through `numpy.asarray(value, dtype=ds.dtype)`), and broadcast
/// to the selection as h5py broadcasts. The edit is written and synced
/// before this returns; what the in-place editor cannot write raises
/// `NotImplementedError` and leaves the file as it was.
fn __setitem__(
&self,
py: Python<'_>,
key: &Bound<'_, PyAny>,
value: &Bound<'_, PyAny>,
) -> PyResult<()> {
self.check_writable()?;
let Some(dims) = self.dims(py)? else {
return Err(PyNotImplementedError::new_err(
"writing to an empty (null dataspace) dataset is not supported",
));
};
let plan = select::parse(key, &dims)?;
if !plan.fields.is_empty() {
return Err(PyNotImplementedError::new_err(
"writing compound fields by name is not supported by clawhdf5's in-place editor; \
write whole elements",
));
}
let category = edit::category(&self.datatype)?;
let conv = self.converter()?;
let bytes = edit::dataset_bytes(
py,
value,
conv.dtype.bind(py),
category,
&plan,
self.chunks.as_deref(),
)?;
if plan.is_empty() {
return Ok(());
}
let sel = edit::selection(&plan, &dims)?;
let path = node::name(&self.path);
self.handle
.edit(py, |ed| ed.write_selection(&path, &sel, &bytes))
}
/// Change the dataset's shape (file opened with `'r+'`), as h5py's
/// `Dataset.resize`: `ds.resize((100, 20))`, or `ds.resize(100, axis=0)`.
/// Only chunked datasets, within their maximum shape; new elements read
/// as the fill value.
#[pyo3(signature = (size, axis=None))]
fn resize(&self, py: Python<'_>, size: &Bound<'_, PyAny>, axis: Option<isize>) -> PyResult<()> {
self.check_writable()?;
let Some(dims) = self.dims(py)? else {
return Err(PyTypeError::new_err("Empty datasets cannot be resized"));
};
if self.chunks.is_none() {
return Err(PyTypeError::new_err("Only chunked datasets can be resized"));
}
let shape: Vec<u64> = match axis {
Some(axis) => {
let rank = dims.len();
let a = usize::try_from(axis)
.ok()
.filter(|&a| a < rank)
.ok_or_else(|| {
PyValueError::new_err(format!(
"Invalid axis (0 to {} allowed)",
rank.saturating_sub(1)
))
})?;
let n: u64 = size.extract().map_err(|_| {
PyTypeError::new_err("Argument must be a single int if axis is specified")
})?;
let mut s = dims.clone();
s[a] = n;
s
}
// As h5py: without `axis` the size is a sequence (`tuple(size)`).
None => size.extract().map_err(|_| {
PyTypeError::new_err(format!(
"'{}' object is not iterable",
size.get_type()
.name()
.map(|n| n.to_string())
.unwrap_or_default()
))
})?,
};
if shape.len() != dims.len() {
return Err(PyValueError::new_err(format!(
"new shape {shape:?} has {} dimensions, the dataset {}",
shape.len(),
dims.len()
)));
}
let path = node::name(&self.path);
self.handle.edit(py, |ed| ed.resize(&path, &shape))
} }
/// `numpy.asarray(ds)` reads the whole dataset. /// `numpy.asarray(ds)` reads the whole dataset.
@@ -330,20 +517,20 @@ impl PyDataset {
copy: Option<bool>, copy: Option<bool>,
) -> PyResult<Bound<'py, PyAny>> { ) -> PyResult<Bound<'py, PyAny>> {
let _ = copy; // every read is a fresh array let _ = copy; // every read is a fresh array
let Some(dims) = &self.shape else { let Some(dims) = self.dims(py)? else {
return Err(PyValueError::new_err("an empty dataset has no array value")); return Err(PyValueError::new_err("an empty dataset has no array value"));
}; };
let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any(); let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any();
let plan = select::parse(&ellipsis, dims)?; let plan = select::parse(&ellipsis, &dims)?;
let arr = self.read_plan(py, &plan)?; let arr = self.read_plan(py, &plan, &dims)?;
match dtype { match dtype {
Some(dt) => arr.call_method1("astype", (dt,)), Some(dt) => arr.call_method1("astype", (dt,)),
None => Ok(arr), None => Ok(arr),
} }
} }
fn __len__(&self) -> PyResult<usize> { fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match self.shape.as_deref() { match self.dims(py)?.as_deref() {
Some([first, ..]) => Ok(*first as usize), Some([first, ..]) => Ok(*first as usize),
_ => Err(PyTypeError::new_err( _ => Err(PyTypeError::new_err(
"Attempt to take len() of scalar dataset", "Attempt to take len() of scalar dataset",
@@ -361,9 +548,10 @@ impl PyDataset {
.unwrap_or_default(), .unwrap_or_default(),
Err(_) => format!("{:?}", self.datatype), Err(_) => format!("{:?}", self.datatype),
}; };
let shape = match &self.shape { let shape = match self.dims(py) {
Some(s) => format!("{s:?}"), Ok(Some(s)) => format!("{s:?}"),
None => "None".to_string(), Ok(None) => "None".to_string(),
Err(_) => "?".to_string(),
}; };
format!( format!(
"<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">", "<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">",
+387
View File
@@ -0,0 +1,387 @@
//! In-place editing (`clawhdf5.File(path, 'r+')`) through `FileEditor`:
//! turning what Python assigns into the bytes, selections and attribute
//! values the editor takes.
//!
//! Value conversion follows h5py (see `edit_helpers.py`, run inside the
//! extension module); what `FileEditor` cannot do is `NotImplementedError`
//! before anything is written.
use std::ffi::CString;
use clawhdf5_format::datatype::{
CharacterSet, CompoundMember, Datatype, DatatypeByteOrder, EnumMember, StringPadding,
};
use clawhdf5_format::selection::Selection;
use clawhdf5_rs::AttrValue;
use pyo3::exceptions::{PyNotImplementedError, PyTypeError};
use pyo3::prelude::*;
use pyo3::sync::PyOnceLock;
use pyo3::types::{PyBytes, PyModule, PyTuple};
use crate::select::{Axis, Plan};
/// The helper module, compiled once.
pub(crate) fn helpers(py: Python<'_>) -> PyResult<&Bound<'_, PyModule>> {
static HELPERS: PyOnceLock<Py<PyModule>> = PyOnceLock::new();
let module = HELPERS.get_or_try_init(py, || -> PyResult<Py<PyModule>> {
let code = CString::new(include_str!("edit_helpers.py"))
.map_err(|e| PyTypeError::new_err(e.to_string()))?;
Ok(PyModule::from_code(
py,
&code,
c"clawhdf5/edit_helpers.py",
c"clawhdf5._edit_helpers",
)?
.unbind())
})?;
Ok(module.bind(py))
}
fn not_implemented(what: impl std::fmt::Display) -> PyErr {
PyNotImplementedError::new_err(format!(
"{what} is not supported by clawhdf5's in-place editor"
))
}
/// How values for a dataset of type `dt` are converted (a category of
/// `edit_helpers._convert_array`), or why they cannot be written.
pub(crate) fn category(dt: &Datatype) -> PyResult<&'static str> {
match dt {
Datatype::FixedPoint { .. } => Ok("int"),
Datatype::FloatingPoint { .. } => Ok("float"),
Datatype::Enumeration {
base_type, members, ..
} => {
let is_bool = base_type.type_size() == 1
&& members.len() == 2
&& members
.iter()
.any(|m| m.name == "FALSE" && m.value.first() == Some(&0))
&& members
.iter()
.any(|m| m.name == "TRUE" && m.value.first() == Some(&1));
Ok(if is_bool { "bool" } else { "enum" })
}
Datatype::String {
padding: StringPadding::NullPad,
..
} => Ok("string"),
Datatype::String { padding, .. } => Err(not_implemented(format!(
"writing fixed-length strings padded {padding:?} (libhdf5 converts them \
differently from numpy)"
))),
Datatype::Compound { size, members } => {
if is_complex(*size, members) {
return Ok("complex");
}
check_exact(dt)?;
Ok("exact")
}
Datatype::Opaque { .. } => Ok("exact"),
Datatype::Array { .. } => Err(not_implemented("writing HDF5 array-type elements")),
Datatype::VariableLength { .. } => Err(not_implemented("writing variable-length data")),
Datatype::Reference { .. } => Err(not_implemented("writing references")),
Datatype::BitField { .. } => Err(not_implemented("writing bitfields")),
Datatype::Time { .. } => Err(not_implemented("writing time values")),
}
}
/// h5py's complex numbers: a compound of two identical floats `r`, `i`.
fn is_complex(size: u32, members: &[CompoundMember]) -> bool {
matches!(members, [r, i] if r.name == "r" && i.name == "i"
&& r.datatype == i.datatype
&& matches!(r.datatype, Datatype::FloatingPoint { size: fs, .. }
if r.byte_offset == 0 && i.byte_offset == u64::from(fs) && size == 2 * fs))
}
/// Compound members written byte for byte from the same numpy dtype: fine
/// unless libhdf5 would convert them on the way (strings padded other than
/// with NULs), or the editor cannot write them at all.
fn check_exact(dt: &Datatype) -> PyResult<()> {
match dt {
Datatype::Compound { members, .. } => {
members.iter().try_for_each(|m| check_exact(&m.datatype))
}
Datatype::Array { base_type, .. } => check_exact(base_type),
Datatype::String {
padding: StringPadding::NullPad,
..
}
| Datatype::FixedPoint { .. }
| Datatype::FloatingPoint { .. }
| Datatype::Enumeration { .. }
| Datatype::Opaque { .. }
| Datatype::BitField { .. } => Ok(()),
Datatype::String { .. } => Err(not_implemented(
"writing compounds with strings not padded with NULs",
)),
Datatype::VariableLength { .. } => Err(not_implemented(
"writing compounds with variable-length members",
)),
Datatype::Reference { .. } => Err(not_implemented("writing references")),
Datatype::Time { .. } => Err(not_implemented("writing time values")),
}
}
/// Largest point selection an index-list write builds (one coordinate
/// vector per element).
const MAX_POINTS: usize = 1 << 22;
/// The selection `plan` writes, whose elements are numbered as the value's
/// (row-major over the selection's shape).
pub(crate) fn selection(plan: &Plan, dims: &[u64]) -> PyResult<Selection> {
if plan.axes.is_empty() {
return Ok(Selection::All);
}
if plan.list_axis().is_none() {
let (reads, _) = plan.reads(dims, None, 1);
return match <[_; 1]>::try_from(reads) {
Ok([read]) => Ok(read.sel),
Err(_) => Err(PyTypeError::new_err("internal error: several hyperslabs")),
};
}
// An index list: the points, in the value's order.
let per_axis: Vec<Vec<u64>> = plan
.axes
.iter()
.map(|a| match a {
Axis::Index(i) => vec![*i],
Axis::Slice { start, step, count } => (0..*count).map(|k| start + k * step).collect(),
Axis::List(v) => v.clone(),
})
.collect();
let n = per_axis
.iter()
.try_fold(1usize, |acc, v| acc.checked_mul(v.len()))
.filter(|&n| n <= MAX_POINTS)
.ok_or_else(|| {
not_implemented(format!(
"an index-list write of more than {MAX_POINTS} elements (write it in slices)"
))
})?;
let mut points = Vec::with_capacity(n);
let mut at = vec![0usize; per_axis.len()];
for _ in 0..n {
points.push(at.iter().zip(&per_axis).map(|(&i, v)| v[i]).collect());
for d in (0..at.len()).rev() {
at[d] += 1;
if at[d] < per_axis[d].len() {
break;
}
at[d] = 0;
}
}
Ok(Selection::Points(points))
}
/// The bytes to write for `value` under `plan`, in the dataset's dtype.
pub(crate) fn dataset_bytes(
py: Python<'_>,
value: &Bound<'_, PyAny>,
dtype: &Bound<'_, PyAny>,
category: &str,
plan: &Plan,
chunks: Option<&[u64]>,
) -> PyResult<Vec<u8>> {
let shape = PyTuple::new(py, plan.out_shape())?;
let fancy = plan.list_axis().is_some();
let chunk_elems = chunks.map_or(0, |c| c.iter().fold(1u64, |a, &d| a.saturating_mul(d)));
let bytes = helpers(py)?.call_method1(
"dataset_values",
(value, dtype, category, shape, fancy, chunk_elems),
)?;
Ok(bytes.cast::<PyBytes>()?.as_bytes().to_vec())
}
fn ieee_float(size: u32, byte_order: DatatypeByteOrder) -> Option<Datatype> {
let (exponent_location, exponent_size, mantissa_size, exponent_bias) = match size {
2 => (10, 5, 10, 15),
4 => (23, 8, 23, 127),
8 => (52, 11, 52, 1023),
_ => return None,
};
Some(Datatype::FloatingPoint {
size,
byte_order,
bit_offset: 0,
bit_precision: (size * 8) as u16,
exponent_location,
exponent_size,
mantissa_location: 0,
mantissa_size,
exponent_bias,
})
}
/// The HDF5 datatype h5py writes for a numpy dtype string (`'<i4'`,
/// `'|b1'`, `'>f8'`, `'<c16'`, `'|S5'`).
fn datatype_of(dtype: &str) -> Option<Datatype> {
let order = match dtype.as_bytes().first()? {
b'<' | b'|' | b'=' => DatatypeByteOrder::LittleEndian,
b'>' => DatatypeByteOrder::BigEndian,
_ => return None,
};
let kind = dtype.as_bytes().get(1)?;
let size: u32 = dtype.get(2..)?.parse().ok()?;
match kind {
b'b' if size == 1 => Some(Datatype::Enumeration {
size: 1,
base_type: Box::new(Datatype::FixedPoint {
size: 1,
byte_order: DatatypeByteOrder::LittleEndian,
signed: true,
bit_offset: 0,
bit_precision: 8,
}),
members: vec![
EnumMember {
name: "FALSE".into(),
value: vec![0],
},
EnumMember {
name: "TRUE".into(),
value: vec![1],
},
],
}),
b'i' | b'u' if matches!(size, 1 | 2 | 4 | 8) => Some(Datatype::FixedPoint {
size,
byte_order: order,
signed: *kind == b'i',
bit_offset: 0,
bit_precision: (size * 8) as u16,
}),
b'f' => ieee_float(size, order),
b'c' => {
let part = ieee_float(size / 2, order)?;
Some(Datatype::Compound {
size,
members: vec![
CompoundMember {
name: "r".into(),
byte_offset: 0,
datatype: part.clone(),
},
CompoundMember {
name: "i".into(),
byte_offset: u64::from(size / 2),
datatype: part,
},
],
})
}
b'S' if size > 0 => Some(Datatype::String {
size,
padding: StringPadding::NullPad,
charset: CharacterSet::Ascii,
}),
_ => None,
}
}
/// An attribute value as h5py would store it (`attrs[name] = value`, or
/// `attrs.create(name, data, shape, dtype)`), except that `str` data is
/// stored as fixed-length UTF-8 strings (h5py stores variable-length ones,
/// which the editor cannot write).
pub(crate) fn attr_value(
py: Python<'_>,
value: &Bound<'_, PyAny>,
dtype: Option<&Bound<'_, PyAny>>,
shape: Option<&Bound<'_, PyAny>>,
) -> PyResult<AttrValue> {
if value.is_instance_of::<crate::PyEmpty>() {
return Err(not_implemented(
"writing an empty (null dataspace) attribute",
));
}
let (kind, dt, dims, data): (String, String, Vec<u64>, Vec<u8>) = helpers(py)?
.call_method1("attr_value", (value, dtype, shape))?
.extract()?;
let datatype = if kind == "str" {
let size: u32 = dt
.parse()
.map_err(|_| PyTypeError::new_err("bad string size"))?;
Datatype::String {
size,
padding: StringPadding::NullPad,
charset: CharacterSet::Utf8,
}
} else {
datatype_of(&dt).ok_or_else(|| not_implemented(format!("an attribute of dtype {dt}")))?
};
Ok(AttrValue::Raw {
datatype,
shape: dims,
data,
})
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn numpy_dtypes_map_to_h5py_types() {
assert!(matches!(
datatype_of("<i4"),
Some(Datatype::FixedPoint {
size: 4,
signed: true,
byte_order: DatatypeByteOrder::LittleEndian,
..
})
));
assert!(matches!(
datatype_of(">u2"),
Some(Datatype::FixedPoint {
size: 2,
signed: false,
byte_order: DatatypeByteOrder::BigEndian,
..
})
));
assert!(matches!(
datatype_of("<f2"),
Some(Datatype::FloatingPoint { size: 2, .. })
));
let c = datatype_of("<c16").unwrap();
assert!(matches!(&c, Datatype::Compound { size: 16, members } if is_complex(16, members)));
assert_eq!(category(&c).unwrap(), "complex");
let b = datatype_of("|b1").unwrap();
assert_eq!(category(&b).unwrap(), "bool");
assert!(matches!(
datatype_of("|S5"),
Some(Datatype::String { size: 5, .. })
));
assert!(datatype_of("<f16").is_none());
assert!(datatype_of("<M8").is_none());
assert!(datatype_of("|S0").is_none());
}
#[test]
fn list_writes_become_points_in_value_order() {
let plan = Plan {
axes: vec![
Axis::List(vec![1, 4]),
Axis::Slice {
start: 0,
step: 2,
count: 2,
},
Axis::Index(3),
],
fields: vec![],
scalar: false,
};
let sel = selection(&plan, &[5, 4, 4]).unwrap();
assert_eq!(
sel,
Selection::Points(vec![
vec![1, 0, 3],
vec![1, 2, 3],
vec![4, 0, 3],
vec![4, 2, 3]
])
);
}
}
+175
View File
@@ -0,0 +1,175 @@
"""Values for in-place writes (clawhdf5.File(path, 'r+')), prepared the way
h5py prepares them, so `ds[key] = value` stores what h5py would store.
Loaded by the extension module (src/edit.rs); not a public API.
h5py converts in two ways, and so does this module:
- a value that is not a numpy array (a list, a Python or numpy scalar) is
converted by numpy straight to the dataset's dtype
(`numpy.asarray(value, dtype=ds.dtype)`), with numpy's rules and errors;
- a numpy array is converted by libhdf5, whose numeric conversions clip to
the target's range instead of wrapping: integers saturate, floats are
truncated toward zero and clipped, a double too large for a float becomes
infinity. That is what `_convert_array` reproduces. Where libhdf5 has no
meaningful answer — NaN into an integer, for which it writes a different
arbitrary value per type — this raises ValueError instead of guessing.
"""
import numpy as np
def _no_path(src, dst):
return TypeError(f"No conversion path for dtype: {src!r} -> {dst!r}")
def _to_int(arr, dtype):
"""Integer target: libhdf5's saturating conversion."""
info = np.iinfo(dtype)
kind = arr.dtype.kind
if kind == "b":
return arr.astype(dtype)
if kind in "iu":
src = np.iinfo(arr.dtype)
lo = max(info.min, src.min)
hi = min(info.max, src.max)
clipped = np.clip(arr, np.array(lo, arr.dtype), np.array(hi, arr.dtype))
return clipped.astype(dtype)
if kind == "f":
if np.isnan(arr).any():
raise ValueError(
"cannot write NaN to an integer dataset (libhdf5 would store an arbitrary value)"
)
t = np.trunc(arr.astype(np.float64))
# info.max + 1 and info.min are powers of two: exact as floats.
over = t >= float(info.max + 1)
under = t < float(info.min)
out = np.where(over | under, 0.0, t).astype(dtype)
out[over] = info.max
out[under] = info.min
return out
raise _no_path(arr.dtype, dtype)
def _convert_array(arr, dtype, category):
kind = arr.dtype.kind
if category in ("int", "enum"):
if arr.dtype == dtype and kind in "iu":
return arr
if category == "enum" and kind not in "iu":
raise _no_path(arr.dtype, dtype)
return _to_int(arr, dtype)
if category == "bool":
if kind == "b":
return arr.astype(dtype)
if kind in "iu":
# h5py's bool is an enum over int8; libhdf5 converts integers
# into it by value (saturating), not to FALSE/TRUE, so 3 is
# stored as 3. Keep those bytes: a view, not a cast.
return _to_int(arr, np.dtype("i1")).view(dtype)
raise _no_path(arr.dtype, dtype)
if category == "float":
if kind not in "biuf":
raise _no_path(arr.dtype, dtype)
with np.errstate(over="ignore", invalid="ignore"):
return arr.astype(dtype)
if category == "complex":
if kind != "c":
raise _no_path(arr.dtype, dtype)
with np.errstate(over="ignore", invalid="ignore"):
return arr.astype(dtype)
if category == "string":
if kind != "S":
raise _no_path(arr.dtype, dtype)
return arr.astype(dtype)
# "exact": compound and opaque types, written only from the same dtype.
if arr.dtype == dtype:
return arr
raise _no_path(arr.dtype, dtype)
def _convert_other(value, dtype, category):
if category == "string":
items = np.asarray(value, dtype=object)
if any(isinstance(x, str) for x in items.flat):
meta = dtype.metadata or {}
if meta.get("h5py_encoding") == "utf-8":
enc = [x.encode("utf-8") if isinstance(x, str) else x for x in items.flat]
return np.array(enc, dtype=dtype).reshape(items.shape)
return np.asarray(value, dtype=dtype)
def _broadcast(arr, shape, fancy, chunk_elems):
"""h5py's broadcasting: numpy's rules against the selection's shape
(extra leading length-1 axes allowed) for slices and integers. For an
index list, the exact shape; a scalar only where h5py expands it to the
whole selection (a chunked dataset whose chunk holds at least as many
elements as the selection)."""
if arr.shape == shape:
return arr
if fancy:
size = int(np.prod(shape))
if arr.ndim == 0 and ((chunk_elems > 0 and size <= chunk_elems) or len(shape) == 1):
return np.broadcast_to(arr, shape)
raise TypeError("Broadcasting is not supported for complex selections")
if arr.ndim == 0:
return np.broadcast_to(arr, shape)
err = TypeError(f"Can't broadcast {arr.shape} -> {shape}")
src = arr.shape
while len(src) > len(shape) and src[0] == 1:
src = src[1:]
if len(src) > len(shape):
raise err
try:
return np.broadcast_to(arr.reshape(src), shape)
except ValueError:
raise err from None
def dataset_values(value, dtype, category, shape, fancy, chunk_elems):
"""The bytes to write for `value` under a selection of `shape`, as a
C-ordered array of the dataset's dtype."""
if isinstance(value, np.ndarray):
arr = _convert_array(value, dtype, category)
else:
arr = _convert_other(value, dtype, category)
arr = _broadcast(arr, tuple(shape), fancy, chunk_elems)
return np.ascontiguousarray(arr, dtype=dtype).tobytes()
def attr_value(value, dtype=None, shape=None):
"""(kind, dtype string, shape, bytes) for an attribute value, h5py's
`attrs[name] = value` / `attrs.create(name, data, shape, dtype)`:
- "str": `str` data (h5py would store a variable-length string; this
stores a fixed-length UTF-8 string, which clawhdf5 can write);
the dtype string is the byte length of the longest element;
- "raw": a numeric, bool or bytes array, as numpy lays it out.
"""
if dtype is not None:
arr = np.asarray(value, dtype=dtype, order="C")
else:
arr = np.asarray(value, order="C")
if shape is not None:
arr = arr.reshape(shape)
kind = arr.dtype.kind
if kind == "O":
if arr.size and all(isinstance(x, str) for x in arr.flat):
kind = "U"
elif arr.size and all(isinstance(x, bytes) for x in arr.flat):
arr = arr.astype(bytes)
kind = "S"
else:
raise TypeError(
f"clawhdf5 cannot write an attribute of Python objects ({value!r:.60})"
)
if kind == "U":
enc = [str(x).encode("utf-8") for x in arr.flat]
size = max([len(b) for b in enc] + [1])
data = np.array(enc, dtype=f"S{size}").reshape(arr.shape)
return ("str", str(size), arr.shape, data.tobytes())
if kind in "biufcS":
return ("raw", arr.dtype.str, arr.shape, np.ascontiguousarray(arr).tobytes())
raise NotImplementedError(
f"clawhdf5 cannot write an attribute of dtype {arr.dtype} in place"
)
+228 -32
View File
@@ -1,13 +1,17 @@
//! PyFile — the main entry point for opening and creating HDF5 files. //! PyFile — the main entry point for opening and creating HDF5 files.
use std::collections::HashMap;
use std::path::PathBuf; use std::path::PathBuf;
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use std::time::Duration;
use pyo3::exceptions::{PyNotImplementedError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::PyList; use pyo3::types::{PyDict, PyList};
use crate::attrs::PyAttrs; use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group}; use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err}; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
/// Internal state for write mode. /// Internal state for write mode.
@@ -23,8 +27,10 @@ struct WriteState {
/// Mirrors the h5py.File interface: /// Mirrors the h5py.File interface:
/// ///
/// ```python /// ```python
/// # Reading /// # Reading, a local file or a URL (range requests, nothing downloaded
/// # up front)
/// f = clawhdf5.File('data.h5', 'r') /// f = clawhdf5.File('data.h5', 'r')
/// f = clawhdf5.File('https://example.org/data.h5')
/// ds = f['dataset'] /// ds = f['dataset']
/// f.close() /// f.close()
/// ///
@@ -44,49 +50,200 @@ enum FileInner {
Write(WriteState), Write(WriteState),
} }
/// Whether `s` is a URL (`scheme://…`) rather than a path: the scheme is a
/// letter followed by letters, digits, `+`, `-` or `.` (RFC 3986).
fn is_url(s: &str) -> bool {
let Some((scheme, _)) = s.split_once("://") else {
return false;
};
let mut chars = scheme.chars();
chars.next().is_some_and(|c| c.is_ascii_alphabetic())
&& chars.all(|c| c.is_ascii_alphanumeric() || matches!(c, '+' | '-' | '.'))
}
impl PyFile {
fn from_handle(handle: Arc<Handle>, filename: String) -> Self {
let root = handle.root;
Self {
inner: Some(FileInner::Read(ReadGroup::new(handle, String::new(), root))),
filename,
}
}
}
#[pymethods] #[pymethods]
impl PyFile { impl PyFile {
/// Open or create an HDF5 file. /// Open or create an HDF5 file.
/// ///
/// Parameters: /// Parameters:
/// path: file path /// path: file path, or a URL (`http://`, `https://`, `s3://`, `gs://`,
/// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write /// mode: 'r' for read (default), 'w' for write
#[new] #[new]
#[pyo3(signature = (path, mode="r"))] #[pyo3(signature = (path, mode="r"))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> { fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
let filename = path.to_string(); let filename = path.to_string();
match mode { if is_url(path) {
"r" => { if mode != "r" {
let file = py.detach(|| { return Err(PyValueError::new_err(format!(
crate::no_panic(|| clawhdf5_rs::File::open(path).map_err(to_py_err)) "remote files are read-only: mode '{mode}' is not supported for a URL"
})?; )));
Ok(Self {
inner: Some(FileInner::Read(root_group(Arc::new(file)))),
filename,
})
} }
let handle = Handle::open_url(py, path, &clawhdf5_remote::Options::default())?;
return Ok(Self::from_handle(handle, filename));
}
match mode {
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
"r+" => Ok(Self::from_handle(
Handle::open_editable(py, path)?,
filename,
)),
"a" if std::path::Path::new(path).exists() => Ok(Self::from_handle(
Handle::open_editable(py, path)?,
filename,
)),
"a" => Err(PyNotImplementedError::new_err(format!(
"mode 'a' on {path}, which does not exist: clawhdf5 can only edit an existing \
file in place; create a new one with mode 'w'"
))),
"w" => Ok(Self { "w" => Ok(Self {
filename, filename,
inner: Some(FileInner::Write(WriteState { inner: Some(FileInner::Write(WriteState {
path: PathBuf::from(path), // Absolute now: the file is written at close, possibly
// after the working directory changed.
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
root_datasets: Vec::new(), root_datasets: Vec::new(),
root_attrs: Arc::new(Mutex::new(Vec::new())), root_attrs: Arc::new(Mutex::new(Vec::new())),
groups: Vec::new(), groups: Vec::new(),
})), })),
}), }),
other => Err(PyErr::new::<pyo3::exceptions::PyValueError, _>(format!( other => Err(PyValueError::new_err(format!(
"unsupported mode '{other}'; expected 'r' or 'w'" "unsupported mode '{other}'; expected 'r', 'r+', 'a' or 'w'"
))), ))),
} }
} }
/// Open a remote file for reading, with options.
///
/// The file is read through a block cache with range requests: opening
/// costs one request (it also fetches the first block), and a read
/// fetches only the blocks it needs. The GIL is released while waiting
/// on the network.
///
/// Parameters (all optional):
/// block_size: bytes per cached block (default 1 MiB)
/// cache_size: byte budget of the block cache (default 64 MiB)
/// headers: dict of extra HTTP headers (e.g. Authorization), sent only
/// to the URL's own origin
/// retries: retries of a request that failed transiently (default 3)
/// timeout: seconds to connect and receive response headers (default 30)
/// allow_full_download: when the server ignores Range requests,
/// download the whole file once instead of failing (default False)
/// max_full_download: largest file such a download may fetch
/// (default 1 GiB)
/// require_validator: refuse a server that sends neither ETag nor
/// Last-Modified (default False)
/// max_redirects: redirects followed per request (default 5)
/// max_parallel: requests of one read in flight at once (default 8)
#[staticmethod]
#[allow(clippy::too_many_arguments)]
#[pyo3(signature = (url, *, block_size=None, cache_size=None, headers=None, retries=None,
timeout=None, allow_full_download=None, max_full_download=None,
require_validator=None, max_redirects=None, max_parallel=None))]
fn open_url(
py: Python<'_>,
url: &str,
block_size: Option<u64>,
cache_size: Option<u64>,
headers: Option<HashMap<String, String>>,
retries: Option<u32>,
timeout: Option<f64>,
allow_full_download: Option<bool>,
max_full_download: Option<u64>,
require_validator: Option<bool>,
max_redirects: Option<u32>,
max_parallel: Option<usize>,
) -> PyResult<Self> {
let mut options = clawhdf5_remote::Options::default();
if let Some(b) = block_size {
if b == 0 {
return Err(PyValueError::new_err("block_size must be positive"));
}
options.cache.block_size = b;
options.cache.coalesce_gap = b;
// The opening request fetches the first block, not 1 MiB.
options.http.first_request = b;
}
if let Some(c) = cache_size {
options.cache.capacity = c;
}
let http = &mut options.http;
if let Some(h) = headers {
http.headers = h.into_iter().collect();
}
if let Some(r) = retries {
http.retries = r;
}
if let Some(t) = timeout {
if !(t.is_finite() && t > 0.0) {
return Err(PyValueError::new_err("timeout must be a positive number"));
}
http.timeout = Duration::from_secs_f64(t);
}
if let Some(a) = allow_full_download {
http.allow_full_download = a;
}
if let Some(m) = max_full_download {
http.max_full_download = m;
}
if let Some(v) = require_validator {
http.require_validator = v;
}
if let Some(r) = max_redirects {
http.max_redirects = r;
}
if let Some(p) = max_parallel {
if p == 0 {
return Err(PyValueError::new_err("max_parallel must be positive"));
}
http.max_parallel = p;
}
let handle = Handle::open_url(py, url, &options)?;
Ok(Self::from_handle(handle, url.to_string()))
}
/// For a remote file, what its block cache has done so far (reads,
/// hits, misses, requests, bytes fetched, ...); `None` for a local file.
#[getter]
fn remote_stats<'py>(&self, py: Python<'py>) -> PyResult<Option<Bound<'py, PyDict>>> {
let Some(storage) = self.read_file()?.handle.remote_storage() else {
return Ok(None);
};
let s = storage.stats();
let d = PyDict::new(py);
d.set_item("reads", s.reads)?;
d.set_item("hits", s.hits)?;
d.set_item("misses", s.misses)?;
d.set_item("waits", s.waits)?;
d.set_item("requests", s.requests)?;
d.set_item("fetch_calls", s.fetch_calls)?;
d.set_item("bytes_fetched", s.bytes_fetched)?;
d.set_item("evictions", s.evictions)?;
d.set_item("cached_bytes", s.cached_bytes)?;
Ok(Some(d))
}
/// Close the file. In write mode, this finalizes and writes the file. /// Close the file. In write mode, this finalizes and writes the file.
fn close(&mut self) -> PyResult<()> { fn close(&mut self) -> PyResult<()> {
let inner = self.inner.take().ok_or_else(|| { let inner = self.inner.take().ok_or_else(|| {
PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed") PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed")
})?; })?;
match inner { match inner {
FileInner::Read(_) => Ok(()), FileInner::Read(root) => {
root.handle.close();
Ok(())
}
FileInner::Write(state) => finalize_write(state), FileInner::Write(state) => finalize_write(state),
} }
} }
@@ -121,7 +278,7 @@ impl PyFile {
/// List the names of all children in the root group. /// List the names of all children in the root group.
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let names = self.read_file()?.member_names()?; let names = self.read_file()?.member_names(py)?;
Ok(PyList::new(py, names)?.into_any().unbind()) Ok(PyList::new(py, names)?.into_any().unbind())
} }
@@ -139,8 +296,8 @@ impl PyFile {
self.keys(py)?.call_method0(py, "__iter__") self.keys(py)?.call_method0(py, "__iter__")
} }
fn __len__(&self) -> PyResult<usize> { fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
Ok(self.read_file()?.member_names()?.len()) Ok(self.read_file()?.member_names(py)?.len())
} }
/// The root group's name, `/`. /// The root group's name, `/`.
@@ -149,7 +306,31 @@ impl PyFile {
"/" "/"
} }
/// The path the file was opened with. /// `'r'` for a file opened read-only (a local file or a URL), `'r+'`
/// for one open for editing or writing, as h5py reports it.
#[getter]
fn mode(&self) -> PyResult<&'static str> {
match &self.inner {
Some(FileInner::Read(root)) if !root.handle.is_writable() => Ok("r"),
Some(_) => Ok("r+"),
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"file is closed",
)),
}
}
/// Nothing to do: every edit is written and synced when it is made, and
/// a file opened with 'w' is written on `close()`.
fn flush(&self) {}
/// Deleting objects is not supported (h5py's `del f[name]`).
fn __delitem__(&self, key: &str) -> PyResult<()> {
Err(PyNotImplementedError::new_err(format!(
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
)))
}
/// The path (or URL) the file was opened with.
#[getter] #[getter]
fn filename(&self) -> &str { fn filename(&self) -> &str {
&self.filename &self.filename
@@ -204,9 +385,9 @@ impl PyFile {
/// Attribute access. In read mode, returns attributes of the root group. /// Attribute access. In read mode, returns attributes of the root group.
/// In write mode, returns a writable attrs handle. /// In write mode, returns a writable attrs handle.
#[getter] #[getter]
fn attrs(&self) -> PyResult<PyAttrs> { fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match self.inner.as_ref() { match self.inner.as_ref() {
Some(FileInner::Read(root)) => root.attrs(), Some(FileInner::Read(root)) => root.attrs(py),
Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))), Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))),
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>( None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"file is closed", "file is closed",
@@ -216,9 +397,10 @@ impl PyFile {
fn __repr__(&self) -> String { fn __repr__(&self) -> String {
match &self.inner { match &self.inner {
Some(FileInner::Read(root)) => { Some(FileInner::Read(root)) => match root.handle.redacted_url() {
format!("<HDF5 File (read, {} bytes)>", root.file.as_bytes().len()) Some(url) => format!("<HDF5 File (read, \"{url}\")>"),
} None => format!("<HDF5 File (read, \"{}\")>", self.filename),
},
Some(FileInner::Write(s)) => { Some(FileInner::Write(s)) => {
format!("<HDF5 File (write, \"{}\")>", s.path.display()) format!("<HDF5 File (write, \"{}\")>", s.path.display())
} }
@@ -226,8 +408,8 @@ impl PyFile {
} }
} }
fn __contains__(&self, key: &str) -> PyResult<bool> { fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
Ok(self.read_file()?.contains(key)) self.read_file()?.contains(py, key)
} }
} }
@@ -248,6 +430,13 @@ impl PyFile {
fn write_state_mut(&mut self) -> PyResult<&mut WriteState> { fn write_state_mut(&mut self) -> PyResult<&mut WriteState> {
match &mut self.inner { match &mut self.inner {
Some(FileInner::Write(s)) => Ok(s), Some(FileInner::Write(s)) => Ok(s),
Some(FileInner::Read(root)) if root.handle.is_writable() => {
Err(PyNotImplementedError::new_err(
"creating datasets or groups in an existing file is not supported by \
clawhdf5's in-place editor (mode 'r+' changes values, shapes and \
attributes)",
))
}
Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>( Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"cannot write to a file opened for reading", "cannot write to a file opened for reading",
)), )),
@@ -271,11 +460,6 @@ fn parse_compression(
} }
} }
fn root_group(file: Arc<clawhdf5_rs::File>) -> ReadGroup {
let root = file.superblock().root_group_address;
ReadGroup::new(file, String::new(), root)
}
/// Build and write the HDF5 file from accumulated write state. /// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> { fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| { crate::no_panic(|| {
@@ -309,6 +493,18 @@ fn finalize_write(state: WriteState) -> PyResult<()> {
mod tests { mod tests {
use super::*; use super::*;
#[test]
fn urls_and_paths() {
assert!(is_url("http://h/f.h5"));
assert!(is_url("s3://bucket/key.h5"));
assert!(is_url("git+https://x"));
assert!(!is_url("data.h5"));
assert!(!is_url("/tmp/a://b.h5"));
assert!(!is_url("dir/x://y"));
assert!(!is_url("1http://x"));
assert!(!is_url("://x"));
}
#[test] #[test]
fn parse_gzip_compression() { fn parse_gzip_compression() {
assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6)); assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6));
+79 -85
View File
@@ -3,11 +3,12 @@
use std::collections::HashMap; use std::collections::HashMap;
use std::sync::{Arc, Mutex, OnceLock}; use std::sync::{Arc, Mutex, OnceLock};
use pyo3::exceptions::{PyIOError, PyKeyError, PyValueError}; use pyo3::exceptions::{PyIOError, PyKeyError, PyNotImplementedError, PyOSError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::PyList; use pyo3::types::PyList;
use crate::attrs::PyAttrs; use crate::attrs::PyAttrs;
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node}; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
/// Shared state for a group being written. /// Shared state for a group being written.
@@ -34,9 +35,9 @@ enum GroupInner {
} }
impl PyGroup { impl PyGroup {
pub(crate) fn from_read(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self { pub(crate) fn from_read(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self { Self {
inner: GroupInner::Read(ReadGroup::new(file, path, addr)), inner: GroupInner::Read(ReadGroup::new(handle, path, addr)),
} }
} }
@@ -60,9 +61,10 @@ impl PyGroup {
/// h5py). It keeps its own address and, once listed, its links, so looking /// h5py). It keeps its own address and, once listed, its links, so looking
/// up a child neither resolves the path from the root nor scans the group's /// up a child neither resolves the path from the root nor scans the group's
/// links again: visiting every member of a large group is linear, not /// links again: visiting every member of a large group is linear, not
/// quadratic. /// quadratic. (Edits never add or remove links, so these stay valid in a
/// file open for editing.)
pub(crate) struct ReadGroup { pub(crate) struct ReadGroup {
pub file: Arc<clawhdf5_rs::File>, pub handle: Arc<Handle>,
pub path: String, pub path: String,
pub addr: u64, pub addr: u64,
/// Link name -> object address (soft links resolved), filled on first use. /// Link name -> object address (soft links resolved), filled on first use.
@@ -72,9 +74,9 @@ pub(crate) struct ReadGroup {
} }
impl ReadGroup { impl ReadGroup {
pub(crate) fn new(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self { pub(crate) fn new(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self { Self {
file, handle,
path, path,
addr, addr,
links: OnceLock::new(), links: OnceLock::new(),
@@ -82,17 +84,14 @@ impl ReadGroup {
} }
} }
fn links(&self) -> PyResult<&HashMap<String, u64>> { fn links(&self, py: Python<'_>) -> PyResult<&HashMap<String, u64>> {
if let Some(links) = self.links.get() { if let Some(links) = self.links.get() {
return Ok(links); return Ok(links);
} }
let entries = crate::no_panic(|| { let (addr, path) = (self.addr, &self.path);
clawhdf5_format::group_v2::resolve_group_children( let entries = self.handle.with(py, |f| {
self.file.as_bytes(), clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr)
self.file.superblock(), .map_err(|e| node::format_err(path, e, PyValueError::new_err))
self.addr,
)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", node::name(&self.path))))
})?; })?;
let map = entries let map = entries
.into_iter() .into_iter()
@@ -102,7 +101,7 @@ impl ReadGroup {
} }
/// The path and address of `key` (a name, a relative or an absolute path). /// The path and address of `key` (a name, a relative or an absolute path).
fn locate(&self, key: &str) -> PyResult<(String, u64)> { fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> {
let path = node::join(&self.path, key); let path = node::join(&self.path, key);
let rel = if self.path.is_empty() { let rel = if self.path.is_empty() {
Some(path.as_str()) Some(path.as_str())
@@ -112,24 +111,24 @@ impl ReadGroup {
path.strip_prefix(self.path.as_str()) path.strip_prefix(self.path.as_str())
.and_then(|r| r.strip_prefix('/')) .and_then(|r| r.strip_prefix('/'))
}; };
let addr = match rel { // A direct child: the link table, when it has the name.
// A direct child: the link table, when it has the name. if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/'))
Some(name) if !name.is_empty() && !name.contains('/') => { && let Some(&a) = self.links(py)?.get(name)
match self.links()?.get(name) { {
Some(&a) => a, return Ok((path, a));
None => node::resolve_from(&self.file, self.addr, name, &path)?, }
} let addr = self.addr;
} let found = self.handle.with(py, |f| match rel {
Some(rel) => node::resolve_from(&self.file, self.addr, rel, &path)?, Some(rel) => node::resolve_from(f, addr, rel, &path),
None => node::address(&self.file, &path)?, None => node::address(f, &path),
}; })?;
Ok((path, addr)) Ok((path, found))
} }
/// `group[key]`. /// `group[key]`.
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> { pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
let (path, addr) = self.locate(key)?; let (path, addr) = self.locate(py, key)?;
node::open(py, &self.file, path, addr) node::open(py, &self.handle, path, addr)
} }
/// `group.get(key, default)`. /// `group.get(key, default)`.
@@ -148,48 +147,57 @@ impl ReadGroup {
} }
/// Names of the group's datasets and subgroups, sorted (h5py's order). /// Names of the group's datasets and subgroups, sorted (h5py's order).
pub(crate) fn member_names(&self) -> PyResult<&[String]> { pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> {
if let Some(m) = self.members.get() { if let Some(m) = self.members.get() {
return Ok(m); return Ok(m);
} }
let mut names = Vec::new(); let links = self.links(py)?;
for (name, &addr) in self.links()? { let path = &self.path;
let hdr = node::header_at(&self.file, addr, &node::join(&self.path, name))?; let mut names = self.handle.with(py, |f| {
if matches!( let mut names = Vec::new();
node::kind(&hdr), for (name, &addr) in links {
Some(node::Kind::Dataset | node::Kind::Group) if matches!(
) { node::kind_at(f, addr, &node::join(path, name))?,
names.push(name.clone()); Some(node::Kind::Dataset | node::Kind::Group)
) {
names.push(name.clone());
}
} }
} Ok(names)
})?;
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes())); names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
Ok(self.members.get_or_init(|| names)) Ok(self.members.get_or_init(|| names))
} }
pub(crate) fn contains(&self, key: &str) -> bool { /// `key in group`: whether `key` names a dataset or group. A failed
self.locate(key) /// read of the file (a network error) is raised, not `False`.
.and_then(|(path, addr)| node::header_at(&self.file, addr, &path)) pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
.ok() let found = self
.and_then(|h| node::kind(&h)) .locate(py, key)
.is_some_and(|k| k != node::Kind::Datatype) .and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path)));
match found {
Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)),
Err(e) if e.is_instance_of::<PyOSError>(py) => Err(e),
Err(_) => Ok(false),
}
} }
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> { pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
self.member_names()? self.member_names(py)?
.iter() .iter()
.map(|n| self.get_item(py, n)) .map(|n| self.get_item(py, n))
.collect() .collect()
} }
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> { pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
self.member_names()? self.member_names(py)?
.iter() .iter()
.map(|n| Ok((n.clone(), self.get_item(py, n)?))) .map(|n| Ok((n.clone(), self.get_item(py, n)?)))
.collect() .collect()
} }
pub(crate) fn attrs(&self) -> PyResult<PyAttrs> { pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path) PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
} }
} }
@@ -210,7 +218,7 @@ impl PyGroup {
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => { GroupInner::Read(g) => {
let list = PyList::new(py, g.member_names()?)?; let list = PyList::new(py, g.member_names(py)?)?;
Ok(list.into_any().unbind()) Ok(list.into_any().unbind())
} }
GroupInner::Write(state) => { GroupInner::Write(state) => {
@@ -236,9 +244,9 @@ impl PyGroup {
self.keys(py)?.call_method0(py, "__iter__") self.keys(py)?.call_method0(py, "__iter__")
} }
fn __len__(&self) -> PyResult<usize> { fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => Ok(g.member_names()?.len()), GroupInner::Read(g) => Ok(g.member_names(py)?.len()),
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()), GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
} }
} }
@@ -293,17 +301,28 @@ impl PyGroup {
state.lock().unwrap().datasets.push(spec); state.lock().unwrap().datasets.push(spec);
Ok(()) Ok(())
} }
GroupInner::Read { .. } => Err(PyIOError::new_err( GroupInner::Read(g) if g.handle.is_writable() => Err(PyNotImplementedError::new_err(
"creating datasets or groups in an existing file is not supported by \
clawhdf5's in-place editor (mode 'r+' changes values, shapes and attributes)",
)),
GroupInner::Read(_) => Err(PyIOError::new_err(
"cannot create datasets on a read-only group", "cannot create datasets on a read-only group",
)), )),
} }
} }
/// Deleting objects is not supported (h5py's `del group[name]`).
fn __delitem__(&self, key: &str) -> PyResult<()> {
Err(PyNotImplementedError::new_err(format!(
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
)))
}
/// Attribute access. /// Attribute access.
#[getter] #[getter]
fn attrs(&self) -> PyResult<PyAttrs> { fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => g.attrs(), GroupInner::Read(g) => g.attrs(py),
GroupInner::Write(state) => { GroupInner::Write(state) => {
let store = Arc::clone(&state.lock().unwrap().attrs); let store = Arc::clone(&state.lock().unwrap().attrs);
Ok(PyAttrs::from_write(store)) Ok(PyAttrs::from_write(store))
@@ -311,10 +330,10 @@ impl PyGroup {
} }
} }
fn __repr__(&self) -> String { fn __repr__(&self, py: Python<'_>) -> String {
match &self.inner { match &self.inner {
GroupInner::Read(g) => { GroupInner::Read(g) => {
let n = g.member_names().map_or(0, |m| m.len()); let n = g.member_names(py).map_or(0, |m| m.len());
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path)) format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
} }
GroupInner::Write(state) => { GroupInner::Write(state) => {
@@ -324,9 +343,9 @@ impl PyGroup {
} }
} }
fn __contains__(&self, key: &str) -> PyResult<bool> { fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => Ok(g.contains(key)), GroupInner::Read(g) => g.contains(py, key),
GroupInner::Write(state) => { GroupInner::Write(state) => {
let guard = state.lock().unwrap(); let guard = state.lock().unwrap();
Ok(guard.datasets.iter().any(|d| d.name == key)) Ok(guard.datasets.iter().any(|d| d.name == key))
@@ -357,31 +376,6 @@ pub(crate) fn finalize_write_group(
mod tests { mod tests {
use super::*; use super::*;
#[test]
fn member_names_are_sorted() {
let mut b = clawhdf5_rs::FileBuilder::new();
b.create_dataset("zeta").with_f64_data(&[1.0]);
b.create_dataset("alpha").with_f64_data(&[1.0]);
let mut g = b.create_group("mid");
g.create_dataset("x").with_f64_data(&[1.0]);
let finished = g.finish();
b.add_group(finished);
let bytes = b.finish().unwrap();
let file = Arc::new(clawhdf5_rs::File::from_bytes(bytes).unwrap());
let root = file.superblock().root_group_address;
let top = ReadGroup::new(Arc::clone(&file), String::new(), root);
assert_eq!(top.member_names().unwrap(), ["alpha", "mid", "zeta"]);
let (path, addr) = top.locate("mid").unwrap();
assert_eq!(path, "mid");
let mid = ReadGroup::new(Arc::clone(&file), path, addr);
assert_eq!(mid.member_names().unwrap(), ["x"]);
assert!(top.contains("mid/x"));
assert!(mid.contains("/alpha"));
assert!(mid.contains("x") && mid.contains("./x"));
assert!(!top.contains("nope"));
assert!(!mid.contains("alpha"));
}
#[test] #[test]
fn finalize_group() { fn finalize_group() {
let state = WriteGroupState { let state = WriteGroupState {
+230
View File
@@ -0,0 +1,230 @@
//! The open file every object of a `File` shares.
//!
//! Every read goes through [`Handle::with`], which releases the GIL and
//! parses through `File::storage()`, so the same code serves a local file
//! (memory-mapped), a remote one (`clawhdf5-remote`: range requests through
//! a block cache, so a network read never holds the GIL) and a file open
//! for editing.
//!
//! A file opened with `'r+'` also holds a [`FileEditor`]. An edit takes the
//! file's write lock, so no read runs while the file changes underneath it,
//! and reopens the file afterwards, through the editor's own open file
//! rather than its path: reads after an edit see the new bytes
//! (a grown file, a new dataspace), never a stale mapping or chunk cache.
//! Objects that cache something an edit can change compare
//! [`Handle::generation`] with the value they cached it at.
//!
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
//! with the GIL released, and code that holds it never touches Python.
use std::sync::atomic::{AtomicU64, Ordering};
use std::sync::{Arc, Mutex, PoisonError, RwLock};
use clawhdf5_rs::{File, FileEditor};
use pyo3::exceptions::PyOSError;
use pyo3::prelude::*;
use crate::{panic_text, to_py_err};
/// Where the file's bytes come from.
pub(crate) enum Source {
/// A local file (memory-mapped). Its path is not kept: nothing reopens
/// it by path (see `open_editable`).
Local,
/// A URL, read through `clawhdf5-remote`'s block cache.
Remote {
url: String,
storage: Arc<clawhdf5_remote::RemoteStorage>,
},
}
pub(crate) struct Handle {
/// The file as last opened; `None` if reopening it after an edit failed
/// (every read is then an error rather than a read of stale bytes).
file: RwLock<Option<File>>,
/// For `'r+'`: the editor, until the file is closed.
editor: Option<Mutex<Option<FileEditor>>>,
source: Source,
/// Bumped by every edit.
generation: AtomicU64,
pub offset_size: u8,
pub length_size: u8,
pub root: u64,
}
fn closed_after_failed_reopen() -> PyErr {
PyOSError::new_err("the file could not be reopened after an edit; open it again")
}
impl Handle {
fn new(file: File, source: Source, editor: Option<FileEditor>) -> Arc<Self> {
let sb = file.superblock();
let (offset_size, length_size, root) =
(sb.offset_size, sb.length_size, sb.root_group_address);
Arc::new(Self {
file: RwLock::new(Some(file)),
editor: editor.map(|e| Mutex::new(Some(e))),
source,
generation: AtomicU64::new(0),
offset_size,
length_size,
root,
})
}
/// A local file, read-only.
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
Ok(Self::new(file, Source::Local, None))
}
/// A local file, open for in-place editing (`'r+'`): the editor takes
/// the file's exclusive lock and checks that it can edit the file, then
/// the file is read through the editor's own file — never by path
/// again, so a later `os.chdir` or a rename or replacement of the path
/// cannot make reads (or the editor's plans) come from another file.
pub(crate) fn open_editable(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
let (file, editor) = py.detach(|| {
crate::no_panic(|| {
let editor = FileEditor::open(path).map_err(to_py_err)?;
let file = editor.reader().map_err(to_py_err)?;
Ok((file, editor))
})
})?;
Ok(Self::new(file, Source::Local, Some(editor)))
}
/// A remote file (`http(s)://`, `s3://`, ...).
pub(crate) fn open_url(
py: Python<'_>,
url: &str,
options: &clawhdf5_remote::Options,
) -> PyResult<Arc<Self>> {
let (file, storage) = py.detach(|| {
crate::no_panic(|| {
let storage = clawhdf5_remote::storage_for_url(url, options).map_err(remote_err)?;
let file = File::open_storage(storage.clone()).map_err(to_py_err)?;
Ok((file, storage))
})
})?;
Ok(Self::new(
file,
Source::Remote {
url: url.to_string(),
storage,
},
None,
))
}
/// Run `f` on the file with the GIL released (a remote read may wait
/// on the network; other Python threads run meanwhile). `f` must not
/// touch Python.
pub(crate) fn with<R: Send>(
&self,
py: Python<'_>,
f: impl FnOnce(&File) -> PyResult<R> + Send,
) -> PyResult<R> {
py.detach(|| self.with_detached(f))
}
/// [`with`](Self::with) for code that already runs without the GIL.
pub(crate) fn with_detached<R>(&self, f: impl FnOnce(&File) -> PyResult<R>) -> PyResult<R> {
crate::no_panic(|| {
let guard = self.file.read().unwrap_or_else(PoisonError::into_inner);
let file = guard.as_ref().ok_or_else(closed_after_failed_reopen)?;
f(file)
})
}
/// Edits so far: objects that cache something an edit can change (a
/// dataset's shape, an object's attributes) re-read it when this moved.
pub(crate) fn generation(&self) -> u64 {
self.generation.load(Ordering::Acquire)
}
/// Whether the file was opened for editing (`'r+'`), even if closed since.
pub(crate) fn is_writable(&self) -> bool {
self.editor.is_some()
}
/// The remote file's block cache.
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
match &self.source {
Source::Remote { storage, .. } => Some(storage),
Source::Local => None,
}
}
/// The URL of a remote file, credentials and query values redacted.
pub(crate) fn redacted_url(&self) -> Option<String> {
match &self.source {
Source::Remote { url, .. } => Some(clawhdf5_remote::redact_url(url)),
Source::Local => None,
}
}
/// Release the editor, and with it the file's lock. Objects still
/// open keep reading the file as it was last written; an edit through
/// them is an error.
pub(crate) fn close(&self) {
if let Some(ed) = &self.editor {
ed.lock().unwrap_or_else(PoisonError::into_inner).take();
}
}
/// Apply one edit with the GIL released. No read runs while it writes,
/// and the file is reopened afterwards — also after a failed edit, since
/// a commit that failed part-way may have changed the file.
pub(crate) fn edit<R: Send>(
&self,
py: Python<'_>,
f: impl FnOnce(&mut FileEditor) -> Result<R, clawhdf5_rs::Error> + Send,
) -> PyResult<R> {
let Some(editor) = &self.editor else {
return Err(PyOSError::new_err(match self.source {
Source::Remote { .. } => "remote files are read-only",
Source::Local => "the file is open read-only; open it with mode 'r+' to change it",
}));
};
if matches!(self.source, Source::Remote { .. }) {
return Err(PyOSError::new_err("remote files are read-only"));
}
py.detach(|| {
let mut ed = editor.lock().unwrap_or_else(PoisonError::into_inner);
let ed = ed
.as_mut()
.ok_or_else(|| PyOSError::new_err("the file is closed"))?;
let mut file = self.file.write().unwrap_or_else(PoisonError::into_inner);
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| f(ed)));
// Drop the old mapping and its chunk cache before reopening.
*file = None;
// Through the editor's file, not the path (see `open_editable`).
let reopened = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| ed.reader()));
self.generation.fetch_add(1, Ordering::AcqRel);
match reopened {
Ok(Ok(f)) => *file = Some(f),
Ok(Err(e)) => return Err(to_py_err(e)),
Err(_) => return Err(closed_after_failed_reopen()),
}
drop(file);
match result {
Ok(r) => r.map_err(to_py_err),
Err(p) => Err(crate::InternalError::new_err(format!(
"clawhdf5 internal error (please report it): {}",
panic_text(&*p)
))),
}
})
}
}
/// A `clawhdf5_remote::Error` as a Python exception: the network side
/// (unreachable, a status, no range support, a changed file) is `OSError`,
/// a file that is not HDF5 is what `to_py_err` makes of it.
pub(crate) fn remote_err(e: clawhdf5_remote::Error) -> PyErr {
match e {
clawhdf5_remote::Error::Hdf5(e) => to_py_err(e),
other => PyOSError::new_err(other.to_string()),
}
}
+14 -1
View File
@@ -7,13 +7,22 @@
//! //!
//! with clawhdf5.File('data.h5', 'r') as f: //! with clawhdf5.File('data.h5', 'r') as f:
//! data = f['dataset_name'][:] //! data = f['dataset_name'][:]
//!
//! with clawhdf5.File('http://host/data.h5') as f: # range requests
//! block = f['dataset_name'][10:20]
//!
//! with clawhdf5.File('data.h5', 'r+') as f: # in-place edits
//! f['dataset_name'][0] = 1.5
//! f.attrs['note'] = 'edited'
//! ``` //! ```
mod attrs; mod attrs;
mod convert; mod convert;
mod dataset; mod dataset;
mod edit;
mod file; mod file;
mod group; mod group;
mod handle;
mod node; mod node;
mod select; mod select;
@@ -63,7 +72,7 @@ fn _panic_for_test() -> PyResult<()> {
/// Convert a `clawhdf5_rs::Error` into a `PyErr`. /// Convert a `clawhdf5_rs::Error` into a `PyErr`.
/// ///
/// Maps different error variants to more specific Python exception types: /// Maps different error variants to more specific Python exception types:
/// - I/O errors -> `PyIOError` /// - I/O errors, and failed reads of a remote file -> `PyIOError`/`PyOSError`
/// - Format/parsing errors -> `PyValueError` /// - Format/parsing errors -> `PyValueError`
/// - Missing dataset/path errors -> `PyKeyError` /// - Missing dataset/path errors -> `PyKeyError`
/// - Invalid arguments -> `PyValueError` /// - Invalid arguments -> `PyValueError`
@@ -73,6 +82,10 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr {
use clawhdf5_rs::Error; use clawhdf5_rs::Error;
match &e { match &e {
Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()), Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()),
// A failed read of the storage: a network error on a remote file.
Error::Format(clawhdf5_format::error::FormatError::Storage(_)) => {
PyErr::new::<pyo3::exceptions::PyOSError, _>(e.to_string())
}
Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()), Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()),
Error::NotADataset(_) | Error::MissingMessage(_) => { Error::NotADataset(_) | Error::MissingMessage(_) => {
PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string()) PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string())
+93 -72
View File
@@ -1,16 +1,24 @@
//! Resolving paths to objects in a file opened for reading. //! Resolving paths to objects in a file opened for reading.
//!
//! Everything here parses through `File::storage()` (the `clawhdf5_format`
//! `*_in` functions), never `File::as_bytes()`, so it works the same on a
//! memory-mapped local file and on a remote one; and it runs inside
//! `Handle::with`, without the GIL.
use std::sync::Arc; use std::sync::Arc;
use clawhdf5_format::attribute::AttributeMessage; use clawhdf5_format::attribute::AttributeMessage;
use clawhdf5_format::dataspace::{Dataspace, DataspaceType}; use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
use clawhdf5_format::error::FormatError;
use clawhdf5_format::message_type::MessageType; use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader; use clawhdf5_format::object_header::ObjectHeader;
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError}; use clawhdf5_rs::File;
use pyo3::exceptions::{PyKeyError, PyOSError, PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use crate::dataset::PyDataset; use crate::dataset::{DatasetMeta, PyDataset};
use crate::group::PyGroup; use crate::group::PyGroup;
use crate::handle::Handle;
/// Join `key` onto the group path `base` the way h5py does: an absolute key /// Join `key` onto the group path `base` the way h5py does: an absolute key
/// starts from the root, a relative one from `base`. Paths are kept without /// starts from the root, a relative one from `base`. Paths are kept without
@@ -33,42 +41,46 @@ pub(crate) fn name(path: &str) -> String {
format!("/{path}") format!("/{path}")
} }
/// A format error met at `path`: a failed read of the storage (a network
/// error on a remote file) is an `OSError`, anything else `other(message)`.
pub(crate) fn format_err(path: &str, e: FormatError, other: fn(String) -> PyErr) -> PyErr {
let msg = format!("{}: {e}", name(path));
match e {
FormatError::Storage(_) => PyOSError::new_err(msg),
_ => other(msg),
}
}
fn value_err(msg: String) -> PyErr {
PyValueError::new_err(msg)
}
/// The address of the object at `path`, resolved from the root group. /// The address of the object at `path`, resolved from the root group.
pub(crate) fn address(file: &clawhdf5_rs::File, path: &str) -> PyResult<u64> { pub(crate) fn address(file: &File, path: &str) -> PyResult<u64> {
resolve_from(file, file.superblock().root_group_address, path, path) resolve_from(file, file.superblock().root_group_address, path, path)
} }
/// The address of `rel` resolved from the group at `group` (`full` is the /// The address of `rel` resolved from the group at `group` (`full` is the
/// resulting path, for the error message). /// resulting path, for the error message).
pub(crate) fn resolve_from( pub(crate) fn resolve_from(file: &File, group: u64, rel: &str, full: &str) -> PyResult<u64> {
file: &clawhdf5_rs::File,
group: u64,
rel: &str,
full: &str,
) -> PyResult<u64> {
if rel.is_empty() { if rel.is_empty() {
return Ok(group); return Ok(group);
} }
crate::no_panic(|| { clawhdf5_format::group_v2::resolve_path_from_in(file.storage(), file.superblock(), group, rel)
clawhdf5_format::group_v2::resolve_path_from(file.as_bytes(), file.superblock(), group, rel) .map_err(|e| match e {
.map_err(|e| { FormatError::Storage(_) => format_err(full, e, value_err),
PyKeyError::new_err(format!( e => PyKeyError::new_err(format!(
"Unable to open object (object '{}' doesn't exist): {e}", "Unable to open object (object '{}' doesn't exist): {e}",
name(full) name(full)
)) )),
}) })
})
} }
/// The object header at `addr` (the object at `path`). /// The object header at `addr` (the object at `path`).
pub(crate) fn header_at(file: &clawhdf5_rs::File, addr: u64, path: &str) -> PyResult<ObjectHeader> { pub(crate) fn header_at(file: &File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
crate::no_panic(|| { let sb = file.superblock();
let sb = file.superblock(); ObjectHeader::parse_in(file.storage(), addr, sb.offset_size, sb.length_size)
let at = usize::try_from(addr) .map_err(|e| format_err(path, e, value_err))
.map_err(|_| PyValueError::new_err(format!("{}: address out of range", name(path))))?;
ObjectHeader::parse(file.as_bytes(), at, sb.offset_size, sb.length_size)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))
})
} }
/// What an object header describes. /// What an object header describes.
@@ -96,29 +108,50 @@ pub(crate) fn kind(hdr: &ObjectHeader) -> Option<Kind> {
} }
} }
/// The kind of the object at `addr`, from its header.
pub(crate) fn kind_at(file: &File, addr: u64, path: &str) -> PyResult<Option<Kind>> {
Ok(kind(&header_at(file, addr, path)?))
}
/// What opening an object found, read without the GIL.
enum Found {
Dataset(DatasetMeta),
Group,
Datatype,
Other,
}
/// Open the object at `addr` (whose path is `path`) as a `Dataset` or /// Open the object at `addr` (whose path is `path`) as a `Dataset` or
/// `Group`. Both keep the address, so later reads resolve nothing. /// `Group`. Both keep the address, so later reads resolve nothing.
pub(crate) fn open( pub(crate) fn open(
py: Python<'_>, py: Python<'_>,
file: &Arc<clawhdf5_rs::File>, handle: &Arc<Handle>,
path: String, path: String,
addr: u64, addr: u64,
) -> PyResult<Py<PyAny>> { ) -> PyResult<Py<PyAny>> {
let hdr = header_at(file, addr, &path)?; let found = handle.with(py, |f| {
match kind(&hdr) { let hdr = header_at(f, addr, &path)?;
Some(Kind::Dataset) => Ok(PyDataset::open(py, Arc::clone(file), path, addr, &hdr)? Ok(match kind(&hdr) {
Some(Kind::Dataset) => Found::Dataset(DatasetMeta::load(f, addr, &hdr, &path)?),
Some(Kind::Group) => Found::Group,
Some(Kind::Datatype) => Found::Datatype,
None => Found::Other,
})
})?;
match found {
Found::Dataset(meta) => Ok(PyDataset::new(py, Arc::clone(handle), path, addr, meta)
.into_pyobject(py)? .into_pyobject(py)?
.into_any() .into_any()
.unbind()), .unbind()),
Some(Kind::Group) => Ok(PyGroup::from_read(Arc::clone(file), path, addr) Found::Group => Ok(PyGroup::from_read(Arc::clone(handle), path, addr)
.into_pyobject(py)? .into_pyobject(py)?
.into_any() .into_any()
.unbind()), .unbind()),
Some(Kind::Datatype) => Err(PyTypeError::new_err(format!( Found::Datatype => Err(PyTypeError::new_err(format!(
"{}: committed (named) datatypes are not supported by clawhdf5", "{}: committed (named) datatypes are not supported by clawhdf5",
name(&path) name(&path)
))), ))),
None => Err(PyValueError::new_err(format!( Found::Other => Err(PyValueError::new_err(format!(
"{}: not a dataset, group or datatype", "{}: not a dataset, group or datatype",
name(&path) name(&path)
))), ))),
@@ -126,32 +159,26 @@ pub(crate) fn open(
} }
/// The dataspace message of an object header. /// The dataspace message of an object header.
pub(crate) fn dataspace(file: &clawhdf5_rs::File, hdr: &ObjectHeader) -> PyResult<Dataspace> { pub(crate) fn dataspace(file: &File, hdr: &ObjectHeader, path: &str) -> PyResult<Dataspace> {
crate::no_panic(|| { let sb = file.superblock();
let sb = file.superblock(); let msg = hdr
let msg = hdr .messages
.messages .iter()
.iter() .find(|m| m.msg_type == MessageType::Dataspace)
.find(|m| m.msg_type == MessageType::Dataspace) .ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?; let data = clawhdf5_format::shared_message::message_data_in(
let data = clawhdf5_format::shared_message::message_data( file.storage(),
file.as_bytes(), msg,
msg, sb.offset_size,
sb.offset_size, sb.length_size,
sb.length_size, )
) .map_err(|e| format_err(path, e, value_err))?;
.map_err(|e| PyValueError::new_err(e.to_string()))?; Dataspace::parse(&data, sb.length_size).map_err(|e| format_err(path, e, value_err))
Dataspace::parse(&data, sb.length_size).map_err(|e| PyValueError::new_err(e.to_string()))
})
} }
/// The chunk shape of a chunked dataset (one entry per dataset dimension), /// The chunk shape of a chunked dataset (one entry per dataset dimension),
/// or `None` for other layouts or a layout message that does not parse. /// or `None` for other layouts or a layout message that does not parse.
pub(crate) fn chunk_shape( pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Option<Vec<u64>> {
file: &clawhdf5_rs::File,
hdr: &ObjectHeader,
rank: usize,
) -> Option<Vec<u64>> {
let sb = file.superblock(); let sb = file.superblock();
let msg = hdr let msg = hdr
.messages .messages
@@ -179,24 +206,18 @@ pub(crate) fn is_null(space: &Dataspace) -> bool {
/// The attributes of the object at `addr` (whose path is `path`), sorted by /// The attributes of the object at `addr` (whose path is `path`), sorted by
/// name (h5py's order). Attributes whose messages cannot be parsed are left /// name (h5py's order). Attributes whose messages cannot be parsed are left
/// out, as the facade's `attrs()` does. /// out, as the facade's `attrs()` does.
pub(crate) fn attributes( pub(crate) fn attributes(file: &File, addr: u64, path: &str) -> PyResult<Vec<AttributeMessage>> {
file: &clawhdf5_rs::File,
addr: u64,
path: &str,
) -> PyResult<Vec<AttributeMessage>> {
let hdr = header_at(file, addr, path)?; let hdr = header_at(file, addr, path)?;
crate::no_panic(|| { let sb = file.superblock();
let sb = file.superblock(); let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant( file.storage(),
file.as_bytes(), &hdr,
&hdr, sb.offset_size,
sb.offset_size, sb.length_size,
sb.length_size, )
) .map_err(|e| format_err(path, e, value_err))?;
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))?; attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes())); Ok(attrs)
Ok(attrs)
})
} }
#[cfg(test)] #[cfg(test)]
+27 -4
View File
@@ -7,11 +7,12 @@
//! (negative from the end) drop their axis, slices must have a positive //! (negative from the end) drop their axis, slices must have a positive
//! step, one `Ellipsis` fills the unmentioned axes, a single increasing list //! step, one `Ellipsis` fills the unmentioned axes, a single increasing list
//! of integers may index one axis, and strings name compound fields. //! of integers may index one axis, and strings name compound fields.
//! Everything else (`None`/`np.newaxis`, boolean masks, several index lists) //! Everything else (`None`/`np.newaxis`, several index lists) is refused
//! is refused with the error h5py gives. //! with the error h5py gives; boolean masks, which h5py supports, raise
//! `NotImplementedError`.
use clawhdf5_format::selection::Selection; use clawhdf5_format::selection::Selection;
use pyo3::exceptions::{PyIndexError, PyTypeError, PyValueError}; use pyo3::exceptions::{PyIndexError, PyNotImplementedError, PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::{PyEllipsis, PySlice, PyString, PyTuple}; use pyo3::types::{PyEllipsis, PySlice, PyString, PyTuple};
@@ -254,6 +255,17 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
} }
} }
// A mask of the dataset's whole shape (`ds[ds[()] > 0]`).
if let [a] = args.as_slice() {
let np = key.py().import("numpy")?;
if a.is_instance(&np.getattr("ndarray")?)?
&& a.getattr("dtype")?.getattr("kind")?.extract::<String>()? == "b"
&& a.getattr("ndim")?.extract::<usize>()? > 1
&& a.getattr("shape")?.extract::<Vec<u64>>()? == dims
{
return Err(mask_unsupported());
}
}
if args.iter().any(|a| a.is_none()) { if args.iter().any(|a| a.is_none()) {
return Err(PyTypeError::new_err( return Err(PyTypeError::new_err(
"Indexing with None (or np.newaxis) is not supported", "Indexing with None (or np.newaxis) is not supported",
@@ -332,6 +344,10 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
}) })
} }
fn mask_unsupported() -> PyErr {
PyNotImplementedError::new_err("boolean mask indexing is not supported by clawhdf5")
}
fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> { fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
if a.is_none() { if a.is_none() {
return Err(PyTypeError::new_err( return Err(PyTypeError::new_err(
@@ -379,8 +395,15 @@ fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
let arr = np.call_method1("asarray", (a,))?; let arr = np.call_method1("asarray", (a,))?;
let kind: String = arr.getattr("dtype")?.getattr("kind")?.extract()?; let kind: String = arr.getattr("dtype")?.getattr("kind")?.extract()?;
if kind == "b" { if kind == "b" {
// A mask along this axis (h5py supports them; clawhdf5 does
// not, for reads or writes: an unsupported operation). A mask
// of any other shape is a wrong key, as in h5py.
let shape: Vec<u64> = arr.getattr("shape")?.extract()?;
if shape == [n] {
return Err(mask_unsupported());
}
return Err(PyTypeError::new_err( return Err(PyTypeError::new_err(
"Boolean mask indexing is not supported by clawhdf5", "Boolean indexing array has incompatible shape",
)); ));
} }
let ndim: usize = arr.getattr("ndim")?.extract()?; let ndim: usize = arr.getattr("ndim")?.extract()?;
+135
View File
@@ -1,6 +1,10 @@
"""Shared fixtures for the clawhdf5 Python binding tests.""" """Shared fixtures for the clawhdf5 Python binding tests."""
import os import os
import re
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import pytest import pytest
@@ -16,3 +20,134 @@ def h5py():
pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable") pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable")
pytest.skip("h5py not installed") pytest.skip("h5py not installed")
return mod return mod
# ---------------------------------------------------------------------------
# An HTTP server for remote reads
# ---------------------------------------------------------------------------
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
class RangeServer:
"""A static file server on 127.0.0.1, in a thread of this process, that
answers `Range: bytes=a-b` with 206 and `Content-Range` (the way S3 and
common web servers do), sends an ETag and honours `If-Match`.
- `ranges=False`: ignores `Range` and answers 200 with the whole file,
like a server without range support.
- `down` (set by `close()`): hang up on every request.
- `delay`: seconds to wait before answering each request after the
first `delay_after` ones (a slow network).
- `log`: every request as `(method, path, range header)`.
"""
def __init__(self, root, ranges=True):
self.root = str(root)
self.ranges = ranges
self.delay = 0.0
self.delay_after = 0
self.down = False
self.log = []
self._lock = threading.Lock()
server = self
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def log_message(self, *args): # quiet
pass
def do_HEAD(self):
self._serve(body=False)
def do_GET(self):
self._serve(body=True)
def _serve(self, body):
with server._lock:
server.log.append((self.command, self.path, self.headers.get("Range")))
n = len(server.log)
if server.delay and n > server.delay_after:
time.sleep(server.delay)
if server.down:
# Hang up without an answer (open keep-alive
# connections outlive shutdown(), so close() sets this).
self.close_connection = True
return
path = os.path.join(server.root, self.path.lstrip("/").split("?")[0])
if not os.path.isfile(path):
self.send_response(404)
self.send_header("Content-Length", "0")
self.end_headers()
return
with open(path, "rb") as fh:
data = fh.read()
st = os.stat(path)
etag = f'"{st.st_mtime_ns:x}-{st.st_size:x}"'
want = self.headers.get("If-Match")
if want is not None and want != etag and want != "*":
self.send_response(412)
self.send_header("Content-Length", "0")
self.end_headers()
return
rng = self.headers.get("Range") if server.ranges else None
m = _RANGE.match(rng.strip()) if rng else None
if m and (m.group(1) or m.group(2)):
size = len(data)
if m.group(1):
start = int(m.group(1))
end = int(m.group(2)) if m.group(2) else size - 1
else:
start = max(0, size - int(m.group(2)))
end = size - 1
if start >= size:
self.send_response(416)
self.send_header("Content-Range", f"bytes */{size}")
self.send_header("Content-Length", "0")
self.end_headers()
return
end = min(end, size - 1)
part = data[start : end + 1]
self.send_response(206)
self.send_header("Content-Range", f"bytes {start}-{end}/{size}")
else:
part = data
self.send_response(200)
if server.ranges:
self.send_header("Accept-Ranges", "bytes")
self.send_header("ETag", etag)
self.send_header("Content-Length", str(len(part)))
self.send_header("Content-Type", "application/x-hdf5")
self.end_headers()
if body:
try:
self.wfile.write(part)
except (BrokenPipeError, ConnectionResetError):
pass
self.httpd = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
self.httpd.daemon_threads = True
self.port = self.httpd.server_address[1]
self.thread = threading.Thread(target=self.httpd.serve_forever, daemon=True)
self.thread.start()
def url(self, name):
return f"http://127.0.0.1:{self.port}/{name}"
def requests(self):
with self._lock:
return len(self.log)
def close(self):
self.down = True
self.httpd.shutdown()
self.httpd.server_close()
@pytest.fixture
def range_server(tmp_path):
"""A range-capable server over `tmp_path`."""
server = RangeServer(tmp_path)
yield server
server.close()
+940
View File
@@ -0,0 +1,940 @@
"""In-place editing: clawhdf5.File(path, 'r+') against h5py.
Every edit is applied twice, to two copies of the same file: once through
h5py (libhdf5) and once through clawhdf5 (FileEditor). After every edit both
files are read back with h5py and must hold the same shapes, values and
attributes; clawhdf5's own view must agree; when h5py refuses an edit,
clawhdf5 must refuse it too and leave its file as it was. Files are written
by h5py (libver earliest and latest, so every chunk index kind) and by
clawhdf5; `h5dump` must read every result."""
import io
import os
import shutil
import subprocess
import threading
import numpy as np
import pytest
import clawhdf5
# ---------------------------------------------------------------------------
# Files
# ---------------------------------------------------------------------------
ENUM = {"RED": 0, "GREEN": 1, "BLUE": 7}
def _h5py_file(h5py, path, libver):
rng = np.random.default_rng(1)
with h5py.File(path, "w", libver=libver) as f:
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
f.create_dataset("be_i2", data=np.arange(24, dtype=">i2").reshape(4, 6))
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
f.create_dataset("f2", data=rng.standard_normal(12).astype("<f2"))
f.create_dataset("f4_2d", data=rng.standard_normal((8, 9)).astype("<f4"))
f.create_dataset("f8_3d", data=rng.standard_normal((4, 5, 6)))
f.create_dataset("u8", data=np.arange(10, dtype="<u8"))
f.create_dataset("c8", data=(np.arange(6) + 1j * np.arange(6)).astype("<c8"))
f.create_dataset("bool", data=np.array([True, False, True, True]))
f.create_dataset("enum", data=np.array([0, 1, 7, 0], dtype="i1"),
dtype=h5py.enum_dtype(ENUM, basetype="i1"))
f.create_dataset("s5", data=np.array([b"ab", b"cdefg", b""], dtype="S5"))
cmp_dt = np.dtype([("id", "<i4"), ("x", "<f8"), ("tag", "S3")])
f.create_dataset("cmp", data=np.array([(i, i / 2, b"t%d" % i) for i in range(5)], dtype=cmp_dt))
f.create_dataset("scalar", data=np.float64(3.5))
# Chunked: fixed maxshape (v4 fixed array under latest), one
# unlimited dimension (extensible array), two (v2 B-tree), one chunk.
f.create_dataset("chunk_fixed", data=np.arange(100, dtype="<i8").reshape(10, 10), chunks=(3, 4))
f.create_dataset("chunk_ext", data=rng.standard_normal((12, 7)), chunks=(5, 7), maxshape=(None, 7))
f.create_dataset("chunk_bt2", data=np.arange(30, dtype="<i4").reshape(5, 6), chunks=(2, 2),
maxshape=(None, None))
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20), chunks=(6, 6),
compression="gzip", maxshape=(40, 40))
f.create_dataset("chunk_single", data=np.arange(12, dtype="<u2").reshape(3, 4), chunks=(3, 4),
maxshape=(3, 4))
f.create_dataset("chunk_fill", shape=(8,), dtype="<i4", chunks=(3,), maxshape=(20,), fillvalue=-1)
f.create_dataset("vlen", data=["a", "bb"], dtype=h5py.string_dtype())
# Compact layout (low-level API).
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
dcpl.set_layout(h5py.h5d.COMPACT)
space = h5py.h5s.create_simple((7,))
dsid = h5py.h5d.create(f.id, b"compact", h5py.h5t.STD_I32LE, space, dcpl=dcpl)
dsid.write(h5py.h5s.ALL, h5py.h5s.ALL, np.arange(7, dtype="<i4"))
g = f.create_group("grp")
g.create_dataset("leaf", data=np.arange(5.0))
g.attrs["units"] = "m"
f.attrs["version"] = np.int32(1)
def _clawhdf5_file(path):
with clawhdf5.File(str(path), "w") as f:
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
f.create_dataset("f8", data=np.linspace(0, 1, 30).reshape(5, 6))
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20),
chunks=[6, 6], compression="gzip")
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
g = f.create_group("grp")
g.create_dataset("leaf", data=np.arange(5.0))
g.attrs["units"] = "m"
f.attrs["version"] = 1
# h5py's libver: "earliest" (v1 B-tree chunk indexes), "v114" (the 1.10+
# indexes: fixed and extensible arrays, v2 B-trees, single chunk) and
# "latest" (HDF5 2.0's newest format, which h5dump 1.14 cannot read).
SOURCES = ["h5py-earliest", "h5py-v114", "h5py-latest", "clawhdf5"]
def _make(h5py, tmp_path, source):
base = tmp_path / f"base-{source}.h5"
if source == "clawhdf5":
_clawhdf5_file(base)
else:
_h5py_file(h5py, str(base), source.split("-")[1])
theirs = tmp_path / f"theirs-{source}.h5"
ours = tmp_path / f"ours-{source}.h5"
shutil.copy(base, theirs)
shutil.copy(base, ours)
return str(theirs), str(ours), str(base)
# ---------------------------------------------------------------------------
# Comparing files through h5py
# ---------------------------------------------------------------------------
def _norm_attr(v):
"""Attribute values comparable across the two writers: clawhdf5 stores
`str` as fixed-length UTF-8 (h5py reads bytes), h5py as variable-length
(h5py reads str)."""
if isinstance(v, bytes):
return ("str", v.decode("utf-8"))
if isinstance(v, str):
return ("str", v)
arr = np.asarray(v)
if arr.dtype.kind in "SO":
return ("strs", [x.decode() if isinstance(x, bytes) else x for x in arr.ravel().tolist()], arr.shape)
return (arr.dtype.str, arr.shape, arr.tobytes())
def snapshot(h5py, path):
"""What h5py sees in the file: every dataset's shape, dtype, bytes and
attributes (read without locking: clawhdf5 may hold the file open)."""
out = {}
with h5py.File(path, "r", locking=False) as f:
def visit(name, obj):
attrs = {k: _norm_attr(obj.attrs[k]) for k in obj.attrs}
if isinstance(obj, h5py.Dataset):
if obj.dtype.kind == "O":
data = [x for x in obj[...].ravel().tolist()]
else:
data = obj[()].tobytes() if obj.shape is not None else None
out[name] = (obj.shape, obj.dtype.str, obj.maxshape, data, attrs)
else:
out[name] = ("group", attrs)
visit("/", f)
f.visititems(visit)
return out
def assert_same_files(h5py, theirs, ours, what):
a, b = snapshot(h5py, theirs), snapshot(h5py, ours)
assert a.keys() == b.keys(), what
for k in a:
assert a[k] == b[k], f"{what}: {k} differs\n h5py: {a[k]}\n clawhdf5: {b[k]}"
def assert_ours_reads_like_h5py(h5py, f, path, what):
"""clawhdf5's own view of the file it is editing matches h5py's."""
with h5py.File(path, "r", locking=False) as t:
for name in ["i4", "chunk_ext", "chunk_bt2", "chunk_gzip", "f8_3d", "cmp", "bool", "enum", "scalar"]:
if name not in t:
continue
o = f[name]
assert o.shape == t[name].shape, f"{what}: {name} shape"
assert o.maxshape == t[name].maxshape, f"{what}: {name} maxshape"
np.testing.assert_array_equal(o[()], t[name][()], err_msg=f"{what}: {name}")
for obj in ["/", "grp"]:
assert sorted(f[obj].attrs.keys()) == sorted(t[obj].attrs.keys()), what
for k in t[obj].attrs:
assert _norm_attr(f[obj].attrs[k]) == _norm_attr(t[obj].attrs[k]), f"{what}: {obj}.attrs[{k}]"
def h5dump_reads(path, base=None):
"""h5dump (libhdf5 1.14) reads every object and value of `path` — when it
reads the unedited `base` (it cannot read HDF5 2.0's newest format)."""
exe = shutil.which("h5dump")
if exe is None:
if os.environ.get("CLAWHDF5_REQUIRE_INTEROP") == "1":
pytest.fail("h5dump is required (CLAWHDF5_REQUIRE_INTEROP=1)")
return
h5rs = os.environ.get("CLAWHDF5_H5RS")
if h5rs:
# clawhdf5's structural and checksum validator (scripts/ci-test.sh
# points this at the h5rs it built).
r = subprocess.run([h5rs, "check", path], capture_output=True, text=True)
assert r.returncode == 0, (r.stdout + r.stderr)[-2000:]
if base is not None and subprocess.run([exe, "-H", base], capture_output=True).returncode != 0:
return
r = subprocess.run([exe, path], capture_output=True, text=True)
assert r.returncode == 0, r.stderr[-2000:]
# ---------------------------------------------------------------------------
# Applying one edit both ways
# ---------------------------------------------------------------------------
def _native_conversion(h5py, value, ds_dtype):
"""`value` as libhdf5 converts it to `ds_dtype` in native byte order.
libhdf5 2.0 (h5py 3.16) converts numbers differently when either side is
not in native byte order (its "soft" conversions): a float in (-1, 0)
becomes the integer type's minimum instead of 0, and an unsigned integer
too large for the signed type of the same size wraps instead of
saturating. clawhdf5 applies the native-order results to every byte
order, so the reference is h5py converting into a native dataset of the
same kind; the result then reaches the real dataset by a plain byte
swap."""
if not isinstance(value, np.ndarray):
return value
if value.dtype.kind not in "biuf" or ds_dtype.kind not in "biuf":
return value
if value.dtype.isnative and ds_dtype.isnative:
return value
with h5py.File(io.BytesIO(), "w") as tmp:
d = tmp.create_dataset("t", shape=value.shape, dtype=ds_dtype.newbyteorder("="))
d[...] = value.astype(value.dtype.newbyteorder("="))
return np.asarray(d[()])
def _apply(f, op, h5py=None):
"""Apply `op` to `f`; with `h5py`, `f` is an h5py file and a numpy array
value is first converted as libhdf5 converts in native byte order (see
`_native_conversion`)."""
kind = op[0]
if kind == "set":
_, name, key, value = op
if h5py is not None:
value = _native_conversion(h5py, value, f[name].dtype)
f[name][key] = value
elif kind == "resize":
_, name, size, axis = op
if axis is None:
f[name].resize(size)
else:
f[name].resize(size, axis=axis)
elif kind == "attr":
_, obj, name, value = op
f[obj].attrs[name] = value
else:
raise AssertionError(op)
def edit_both(h5py, theirs, ours_path, ours, op):
"""Apply `op` with h5py and with clawhdf5 (`ours`, open 'r+'); the two
files must then read the same through h5py. If h5py refuses, clawhdf5
must refuse and its file must be unchanged. Returns h5py's error."""
before = snapshot(h5py, ours_path)
try:
with h5py.File(theirs, "r+") as t:
_apply(t, op, h5py)
except Exception as e: # noqa: BLE001 - h5py refuses: so must we
try:
_apply(ours, op)
except Exception: # noqa: BLE001
pass
else:
pytest.fail(f"{op!r:.300}: h5py refused ({type(e).__name__}: {e}), clawhdf5 did not")
assert snapshot(h5py, ours_path) == before, f"{op!r}: clawhdf5 changed the file while failing"
return e
_apply(ours, op)
assert_same_files(h5py, theirs, ours_path, repr(op)[:200])
return None
# ---------------------------------------------------------------------------
# Tests
# ---------------------------------------------------------------------------
@pytest.mark.parametrize("source", SOURCES)
def test_edit_sequence_matches_h5py(h5py, tmp_path, source):
theirs, ours_path, base = _make(h5py, tmp_path, source)
ops = [
("set", "i4", 0, 99),
("set", "i4", (slice(1, 5, 2), slice(None, None, 3)), np.array([[1.5, -2.5, 1e12, -1e12]])),
("set", "i4", (slice(None), 2), np.arange(6, dtype="<i8") * 1000),
("set", "i4", ([0, 2, 5], slice(4, 6)), np.array([[1, 2], [3, 4], [5, 6]])),
("set", "i4", (Ellipsis, -1), np.int8(-7)),
("set", "i4", (3, 3), 12345.9),
("set", "u1", slice(2, 8), np.array([-5, 0, 300, 255, 256, 1], dtype="<i4")),
("set", "u1", slice(0, 3), [1, 2, 3]),
("set", "chunk_gzip", (slice(0, 20, 7), slice(3, 17)), 42.25),
("set", "chunk_gzip", (slice(5, 11), slice(5, 11)), np.ones((6, 6), dtype="<f8") * np.pi),
("set", "grp/leaf", slice(None), np.array([1, 2, 3, 4, 5], dtype="<u2")),
("attr", "/", "version", np.int32(2)),
("attr", "/", "count", 7),
("attr", "grp", "scale", np.array([0.5, 0.25], dtype="<f4")),
("attr", "grp", "matrix", np.arange(6, dtype=">i8").reshape(2, 3)),
("attr", "i4", "flag", np.bool_(True)),
("attr", "i4", "z", np.complex64(1 - 2j)),
("attr", "i4", "raw", np.bytes_(b"abc")),
("set", "i4", slice(0, 2), np.zeros((3, 10))), # shape mismatch: refused by both
]
if source != "clawhdf5":
ops += [
("set", "be_i2", (slice(None), slice(1, 3)), np.array([70000, -70000], dtype="<i8")),
("set", "f2", slice(None, None, 4), np.array([1e6, -3.25, 0.1])),
("set", "f4_2d", (2, slice(None)), np.linspace(-1, 1, 9)),
("set", "f8_3d", (slice(1, 3), 2, slice(None, None, 2)), np.arange(3, dtype="<i2")),
("set", "u8", slice(None), np.array([-1, 0, 2**63, 1e30, -1e30, 5.5, 2, 3, 4, 5])),
("set", "c8", slice(1, 3), np.array([1 + 1j, 2 - 2j], dtype="<c16")),
("set", "c8", 0, np.float64(1.0)), # h5py: no conversion path
("set", "bool", slice(None), np.array([0, 3, 0, -1], dtype="<i4")),
("set", "bool", 1, np.array(True)),
("set", "enum", slice(0, 2), np.array([7, 1], dtype="<i4")),
("set", "s5", 0, np.bytes_(b"xyzuvw")),
("set", "s5", slice(1, 3), [b"q", b"rs"]),
("set", "s5", 2, np.array("uni")), # h5py: no conversion from 'U'
("set", "cmp", 2, np.array((9, 9.5, b"zz"), dtype=[("id", "<i4"), ("x", "<f8"), ("tag", "S3")])),
("set", "cmp", slice(3, 5), [(1, 0.5, b"a"), (2, 1.5, b"b")]),
("set", "scalar", (), 7.25),
("set", "scalar", Ellipsis, np.float32(-1.5)),
("set", "compact", slice(1, 6, 2), np.array([10, 20, 30])),
("set", "chunk_fixed", (slice(2, 9), slice(1, 10, 4)), np.arange(21).reshape(7, 3)),
("set", "chunk_single", (1, slice(None)), np.array([9, 8, 7, 6])),
("resize", "chunk_ext", (20, 7), None),
("set", "chunk_ext", slice(12, 20), np.full((8, 7), 2.5)),
("resize", "chunk_ext", 9, 0),
("resize", "chunk_ext", 16, 0),
("resize", "chunk_bt2", (9, 11), None),
("set", "chunk_bt2", (slice(4, 9), slice(5, 11)), np.arange(30).reshape(5, 6)),
("resize", "chunk_bt2", (3, 3), None),
("resize", "chunk_bt2", (7, 8), None),
("resize", "chunk_gzip", (40, 25), None),
("resize", "chunk_gzip", (41, 25), None), # beyond maxshape: refused
("resize", "chunk_fill", 15, None), # h5py: a size without axis must be a tuple
("resize", "chunk_fill", (15,), None),
("set", "chunk_fill", slice(10, 12), [1, 2]),
("resize", "chunk_fill", (4,), None),
("resize", "chunk_fill", (20,), None),
("resize", "i4", (7, 10), None), # not chunked: refused
]
with clawhdf5.File(ours_path, "r+") as ours:
assert ours.mode == "r+"
for i, op in enumerate(ops):
edit_both(h5py, theirs, ours_path, ours, op)
if i % 5 == 0:
assert_ours_reads_like_h5py(h5py, ours, ours_path, repr(op))
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
h5dump_reads(ours_path, base)
# Reopened, clawhdf5 reads what h5py reads.
with clawhdf5.File(ours_path, "r") as f:
assert f.mode == "r"
assert_ours_reads_like_h5py(h5py, f, ours_path, "reopened")
def _random_key(rng, shape):
key = []
for n in shape:
r = rng.random()
if n == 0 or r < 0.3:
start = int(rng.integers(0, n + 1)) if n else 0
stop = int(rng.integers(start, n + 1)) if n else 0
step = int(rng.integers(1, 4))
key.append(slice(start, stop, step))
elif r < 0.55:
key.append(int(rng.integers(-n, n)))
elif r < 0.7:
k = int(rng.integers(1, min(n, 4) + 1))
key.append(sorted(rng.choice(n, size=k, replace=False).tolist()))
else:
key.append(slice(None))
# Only one index list per key.
lists = [i for i, k in enumerate(key) if isinstance(k, list)]
for i in lists[1:]:
key[i] = slice(None)
return tuple(key)
def _selection_shape(key, shape):
out = []
fancy = False
for k, n in zip(key, shape):
if isinstance(k, slice):
out.append(len(range(*k.indices(n))))
elif isinstance(k, list):
out.append(len(k))
fancy = True
return tuple(out), fancy
def _random_value(rng, sel_shape, fancy, dtype):
r = rng.random()
if r < 0.2 or not sel_shape:
v = rng.standard_normal() * 1000
return np.float64(v) if rng.random() < 0.5 else int(v)
shape = list(sel_shape)
if not fancy and r < 0.35 and shape:
shape[0] = 1 # broadcast along the first axis
kind = rng.choice(["same", "f8", "i8", "u1", "f4"])
if kind == "same" and np.dtype(dtype).kind in "iuf":
dt = np.dtype(dtype)
else:
dt = np.dtype(str(kind) if kind != "same" else "f8")
base = rng.standard_normal(size=shape) * (10 ** rng.integers(0, 6))
with np.errstate(all="ignore"):
return base.astype(dt)
@pytest.mark.parametrize("source", SOURCES)
@pytest.mark.parametrize("seed", [0, 1, 2, 3])
def test_random_edits_match_h5py(h5py, tmp_path, source, seed):
"""Random writes (slices, steps, integers, index lists, broadcasts,
other dtypes and out-of-range values), resizes and attributes, each
compared with h5py applying the same edit."""
rng = np.random.default_rng(1000 * seed + SOURCES.index(source))
theirs, ours_path, base = _make(h5py, tmp_path, source)
with h5py.File(theirs, "r") as t:
names = [n for n in ["i4", "u1", "f8", "f4_2d", "f8_3d", "be_i2", "chunk_fixed", "chunk_ext",
"chunk_bt2", "chunk_gzip", "chunk_fill", "compact", "grp/leaf"] if n in t]
resizable = [n for n in ["chunk_ext", "chunk_bt2", "chunk_gzip", "chunk_fill"] if n in names]
refused = 0
with clawhdf5.File(ours_path, "r+") as ours:
for step in range(40):
r = rng.random()
if r < 0.15 and resizable:
name = str(rng.choice(resizable))
with h5py.File(theirs, "r") as t:
maxshape = t[name].maxshape
shape = t[name].shape
new = tuple(int(rng.integers(0, (m if m is not None else s + 10) + 1)) for m, s in zip(maxshape, shape))
op = ("resize", name, new, None)
elif r < 0.25:
obj = str(rng.choice(["/", "grp", names[0]]))
choices = [np.int16(rng.integers(-100, 100)), rng.standard_normal(3),
np.arange(int(rng.integers(1, 5)), dtype=">u4"), np.float32(0.5)]
value = choices[int(rng.integers(0, len(choices)))]
op = ("attr", obj, f"a{int(rng.integers(0, 4))}", value)
else:
name = str(rng.choice(names))
with h5py.File(theirs, "r") as t:
shape, dtype = t[name].shape, t[name].dtype
key = _random_key(rng, shape)
sel_shape, fancy = _selection_shape(key, shape)
op = ("set", name, key, _random_value(rng, sel_shape, fancy, dtype))
before = None
if op[0] == "resize":
with h5py.File(ours_path, "r", locking=False) as o:
before, fill = o[op[1]][()], o[op[1]].fillvalue
if edit_both(h5py, theirs, ours_path, ours, op) is not None:
refused += 1
elif before is not None:
# Independently of h5py (which a wrong index layout fools
# the same way): the kept elements keep their values.
want = resized_model(before, op[2], fill)
np.testing.assert_array_equal(ours[op[1]][()], want, err_msg=repr(op))
with h5py.File(ours_path, "r", locking=False) as o:
np.testing.assert_array_equal(o[op[1]][()], want, err_msg=repr(op))
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
assert refused < 30
h5dump_reads(ours_path, base)
@pytest.mark.parametrize("seed", range(10, 40))
def test_random_edits_on_clawhdf5_files(h5py, tmp_path, seed):
"""More random sequences on a clawhdf5-written file, whose resizes to
zero extents once left files `h5rs check` could not read."""
test_random_edits_match_h5py(h5py, tmp_path, "clawhdf5", seed)
def resized_model(before, shape, fill):
"""`before` resized to `shape` as HDF5 resizes: elements inside both
extents keep their values, the others read as the fill value."""
out = np.full(shape, fill, dtype=before.dtype)
common = tuple(slice(0, min(a, b)) for a, b in zip(before.shape, shape))
out[common] = before[common]
return out
RESIZES = [(15, 15), (3, 2), (20, 20), (1, 1), (1, 0), (0, 0), (7, 20), (20, 13), (20, 20)]
def _check_resizes(h5py, path, name, orig, maxshape=(20, 20), base=None):
"""Resize `name` through RESIZES in 'r+', each checked against a numpy
model with clawhdf5 and h5py and the file with `h5rs check`."""
model = orig
with clawhdf5.File(path, "r+") as f:
ds = f[name]
for shape in RESIZES:
ds.resize(shape)
model = resized_model(model, shape, 0)
np.testing.assert_array_equal(ds[()], model, err_msg=f"{name} {shape}")
with h5py.File(path, "r", locking=False) as t:
np.testing.assert_array_equal(t[name][()], model, err_msg=f"h5py: {name} {shape}")
assert t[name].maxshape == maxshape
h5rs = os.environ.get("CLAWHDF5_H5RS")
if h5rs: # h5dump would wait for the editor's lock
r = subprocess.run([h5rs, "check", "--data", path], capture_output=True, text=True)
assert r.returncode == 0, f"{name} {shape}: " + (r.stdout + r.stderr)[-2000:]
with pytest.raises(ValueError):
ds.resize((maxshape[0] + 1, 20))
# Written values survive a shrink.
ds[...] = orig
ds.resize((15, 15))
with h5py.File(path, "r") as t:
np.testing.assert_array_equal(t[name][()], orig[:15, :15])
h5dump_reads(path, base)
@pytest.mark.parametrize("source", SOURCES)
def test_resizes_keep_values(h5py, tmp_path, source):
"""Shrinking, zero extents and growing back keep the values a numpy
model keeps, on every source (on clawhdf5's own files a shrink once
moved every chunk: h5py read the same wrong values)."""
_, ours, base = _make(h5py, tmp_path, source)
with clawhdf5.File(ours, "r+") as f:
f["chunk_gzip"].resize((20, 20))
maxshape = (20, 20) if source == "clawhdf5" else (40, 40)
_check_resizes(h5py, ours, "chunk_gzip", np.arange(400, dtype="<f4").reshape(20, 20), maxshape, base)
FIXTURES = os.path.join(os.path.dirname(__file__), "..", "..", "clawhdf5", "tests", "fixtures")
@pytest.mark.parametrize("name", ["d", "z"])
def test_resizes_of_a_file_with_no_recorded_maxshape(h5py, tmp_path, name):
"""A file clawhdf5 2.7.0 wrote records no maximum dimensions for chunked
datasets; resizing it must keep the Fixed Array index's layout."""
path = str(tmp_path / "old.h5")
shutil.copy(os.path.join(FIXTURES, "chunked_no_maxshape_v2_7_0.h5"), path)
with h5py.File(path, "r") as t:
orig = t[name][()]
_check_resizes(h5py, path, name, orig)
NUMERIC = ["<i1", "<u1", "<i2", ">u2", "<i4", "<u4", ">i8", "<u8", "<f2", "<f4", ">f8"]
def _quiet(f):
with np.errstate(all="ignore"):
return f()
def _libhdf5_undefined(vals, target):
"""The values whose conversion to `target` libhdf5 2.0 (h5py 3.16) gets
wrong even in native byte order, where its C casts are undefined
behaviour; clawhdf5 saturates them as libhdf5's range handling
intends (docs/known-issues.md):
- half floats into unsigned integers: negatives wrap (-1 -> 65535) and
+inf becomes 0; into signed integers, +-inf becomes the minimum;
- a float equal to the integer maximum rounded up in the float's
precision (float32(2**31 - 1) == 2**31 -> int32, float64(2**64 - 1)
-> uint64) becomes the minimum (or 0);
- a double between 65504 and 65520 into a half float becomes infinity
(IEEE rounds it down to 65504, as numpy does).
"""
bad = np.zeros(vals.shape, dtype=bool)
if vals.dtype.kind != "f":
return bad
if target.kind in "iu":
if vals.dtype.itemsize == 2:
bad |= np.isinf(vals)
if target.kind == "u":
bad |= vals <= -1
top = _quiet(lambda: np.array(np.iinfo(target).max).astype(vals.dtype))
if float(top) > np.iinfo(target).max:
bad |= vals == top
if target.kind == "f" and target.itemsize == 2 and vals.dtype.itemsize > 2:
bad |= (np.abs(vals) > 65504) & (np.abs(vals) < 65520)
return bad
def test_numeric_conversions_match_h5py(h5py, tmp_path):
"""Every numeric source dtype into every numeric dataset dtype, with
values at and beyond the targets' limits, as libhdf5 converts them."""
edge = np.array([0, 1, -1, -0.3, 2.5, -2.5, 3.7, -3.7, 127.9, -128.9, 200.5, 255.5, 256, -129,
32767.5, 40000, 65504, 70000, -70000, 2**31 - 1, 2**31, -2**31 - 1,
4e9, 1e15, -1e15, 1e19, 1e300, -1e300, np.inf, -np.inf])
sources = {
"f8": edge,
"f4": _quiet(lambda: edge.astype("<f4")),
"f2": np.array([0, 1, -1, 2.5, -3.5, 65504, -65504, np.inf, -np.inf, 100.5], dtype="<f2"),
"i8": np.array([0, 1, -1, 127, 128, -129, 255, 256, 32768, -32769, 65536, 2**31, -2**31 - 1,
2**32, 2**62, -2**63, 2**63 - 1], dtype="<i8"),
"u8": np.array([0, 1, 127, 128, 255, 256, 65535, 65536, 2**31, 2**32, 2**63, 2**64 - 1], dtype="<u8"),
"i1": np.array([-128, -1, 0, 1, 127], dtype="i1"),
"u2": np.array([0, 255, 256, 65535], dtype=">u2"),
"b": np.array([True, False, True]),
}
for target in NUMERIC:
for sname, src in sources.items():
vals = src[~_libhdf5_undefined(src, np.dtype(target))]
path_t = str(tmp_path / f"t_{target[1:]}_{sname}.h5")
path_o = str(tmp_path / f"o_{target[1:]}_{sname}.h5")
with h5py.File(path_t, "w") as f:
f.create_dataset("d", shape=vals.shape, dtype=target)
shutil.copy(path_t, path_o)
what = f"{vals.dtype} -> {target}"
try:
with h5py.File(path_t, "r+") as f:
f["d"][...] = _native_conversion(h5py, vals, f["d"].dtype)
except Exception: # noqa: BLE001
with clawhdf5.File(path_o, "r+") as f, pytest.raises(Exception):
f["d"][...] = vals
continue
with clawhdf5.File(path_o, "r+") as f:
f["d"][...] = vals
with h5py.File(path_t, "r") as a, h5py.File(path_o, "r") as b:
assert a["d"][...].tobytes() == b["d"][...].tobytes(), (
f"{what}: h5py {a['d'][...].tolist()} clawhdf5 {b['d'][...].tolist()}"
)
def test_nan_into_an_integer_dataset_is_refused(h5py, tmp_path):
"""libhdf5 stores NaN as an arbitrary integer (0, the minimum or 2**63,
depending on the type); clawhdf5 refuses and writes nothing."""
path = str(tmp_path / "nan.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
with clawhdf5.File(path, "r+") as f:
with pytest.raises(ValueError, match="NaN"):
f["d"][...] = np.array([1.0, np.nan, 2.0, 3.0])
# A Python list goes through numpy, which refuses NaN too.
with pytest.raises(ValueError):
f["d"][0:2] = [np.nan, 1.0]
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["d"][...], np.arange(4))
def test_unsupported_edits_are_clear_errors(h5py, tmp_path):
path = str(tmp_path / "u.h5")
_h5py_file(h5py, path, "earliest")
before = snapshot(h5py, path)
with clawhdf5.File(path, "r+") as f:
with pytest.raises(NotImplementedError, match="delet"):
del f.attrs["version"]
with pytest.raises(NotImplementedError, match="delet"):
del f["i4"]
with pytest.raises(NotImplementedError, match="delet"):
del f["grp"]["leaf"]
with pytest.raises(NotImplementedError):
f.create_dataset("new", data=np.arange(3.0))
with pytest.raises(NotImplementedError):
f.create_group("newgrp")
with pytest.raises(NotImplementedError):
f["grp"].create_dataset("new", data=np.arange(3.0))
with pytest.raises(NotImplementedError, match="variable-length"):
f["vlen"][0] = "x"
with pytest.raises(NotImplementedError, match="field"):
f["cmp"]["id"] = np.arange(5)
with pytest.raises(NotImplementedError):
f.attrs["empty"] = clawhdf5.Empty("f8")
with pytest.raises(TypeError, match="chunked"):
f["i4"].resize((7, 10))
with pytest.raises(ValueError):
f["chunk_gzip"].resize((41, 20))
with pytest.raises(ValueError, match="axis"):
f["chunk_ext"].resize(3, axis=2)
with pytest.raises(TypeError):
f["i4"][0] = np.array(["a"] * 10)
# h5py writes through boolean masks; clawhdf5 does not.
with pytest.raises(NotImplementedError, match="mask"):
f["u1"][np.arange(16) % 2 == 0] = 5
with pytest.raises(NotImplementedError, match="mask"):
f["i4"][f["i4"][()] > 30] = 0
with pytest.raises(NotImplementedError, match="mask"):
f["i4"][np.ones(6, dtype=bool), 2] = 0
assert snapshot(h5py, path) == before
h5dump_reads(path)
def test_read_only_files_and_modes(h5py, tmp_path):
path = str(tmp_path / "m.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"), chunks=(2,), maxshape=(None,))
with clawhdf5.File(path, "r") as f:
with pytest.raises(OSError, match="r\\+"):
f["d"][0] = 1
with pytest.raises(OSError):
f["d"].resize((8,))
with pytest.raises(OSError):
f.attrs["x"] = 1
with pytest.raises(NotImplementedError, match="does not exist"):
clawhdf5.File(str(tmp_path / "missing.h5"), "a")
with pytest.raises(ValueError, match="mode"):
clawhdf5.File(path, "rw")
with clawhdf5.File(path, "a") as f:
assert f.mode == "r+"
f["d"][1] = 10
with h5py.File(path, "r") as f:
assert f["d"][1] == 10
def test_the_file_is_locked_while_open_for_editing(h5py, tmp_path):
path = str(tmp_path / "lock.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
f = clawhdf5.File(path, "r+")
with pytest.raises(OSError):
clawhdf5.File(path, "r+")
with pytest.raises(OSError):
h5py.File(path, "r+")
f.close()
with h5py.File(path, "r+") as g:
g["d"][0] = 5
with clawhdf5.File(path, "r+") as g:
g["d"][1] = 6
with h5py.File(path, "r") as g:
np.testing.assert_array_equal(g["d"][...], [5, 6, 2, 3])
def test_objects_see_edits_made_through_others(h5py, tmp_path):
"""A dataset or attrs object taken before an edit reports the file as
it is after it: the new shape, the new attribute."""
path = str(tmp_path / "live.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(6.0), chunks=(4,), maxshape=(None,))
with clawhdf5.File(path, "r+") as f:
d1 = f["d"]
d2 = f["d"]
attrs = d1.attrs
assert len(attrs) == 0 and "u" not in attrs
d2.resize((10,))
assert d1.shape == (10,) and d1.size == 10 and len(d1) == 10
np.testing.assert_array_equal(d1[6:], np.zeros(4))
d2.attrs["u"] = "m/s"
assert "u" in attrs and attrs["u"] == b"m/s" and len(attrs) == 1
f.attrs.create("shaped", np.arange(6), shape=(2, 3), dtype="<i2")
f.attrs.modify("shaped2", [1.5, 2.5])
with h5py.File(path, "r") as f:
assert f["d"].shape == (10,)
assert f["d"].attrs["u"] == b"m/s"
assert f.attrs["shaped"].dtype == np.dtype("<i2") and f.attrs["shaped"].shape == (2, 3)
np.testing.assert_array_equal(f.attrs["shaped2"], [1.5, 2.5])
def test_attribute_types_as_h5py_reads_them(h5py, tmp_path):
path = str(tmp_path / "attrs.h5")
with h5py.File(path, "w") as f:
f.create_group("g")
values = {
"i8": 5,
"f8": 2.5,
"i1": np.int8(-3),
"u8": np.uint64(2**64 - 1),
">f4": np.array([1.5, 2.5], dtype=">f4"),
"f2": np.float16(0.5),
"b": True,
"barr": np.array([True, False]),
"c16": np.complex128(1 + 2j),
"bytes": b"raw",
"sarr": np.array([b"a", b"bcd"]),
"str": "héllo",
"strs": ["x", "yz"],
"2d": np.arange(12, dtype="<u2").reshape(3, 4),
"empty": np.zeros((0,), dtype="<i4"),
}
with clawhdf5.File(path, "r+") as f:
for k, v in values.items():
f["g"].attrs[k] = v
# Replace one, with another type and size.
f["g"].attrs["i8"] = np.arange(100.0)
with h5py.File(path, "r") as f:
a = f["g"].attrs
np.testing.assert_array_equal(a["i8"], np.arange(100.0))
assert a["f8"] == 2.5 and a["f8"].dtype == np.float64
assert a["i1"] == -3 and a["i1"].dtype == np.int8
assert a["u8"] == 2**64 - 1 and a["u8"].dtype == np.uint64
assert a[">f4"].dtype == np.dtype(">f4")
assert a["f2"].dtype == np.float16
assert a["b"] is np.True_ or a["b"] == True # noqa: E712
assert a["barr"].dtype == np.bool_
assert a["c16"] == 1 + 2j
assert a["bytes"] == b"raw"
assert list(a["sarr"]) == [b"a", b"bcd"]
# str is stored as fixed-length UTF-8: h5py reads bytes.
assert a["str"].decode("utf-8") == "héllo"
assert [x.decode() for x in a["strs"]] == ["x", "yz"]
assert a["2d"].shape == (3, 4) and a["2d"].dtype == np.dtype("<u2")
assert a["empty"].shape == (0,)
with clawhdf5.File(path, "r") as f:
assert f["g"].attrs["c16"] == 1 + 2j
assert f["g"].attrs["str"].decode("utf-8") == "héllo"
h5dump_reads(path)
def test_many_attributes_move_to_dense_storage(h5py, tmp_path):
"""Past the compact limit (8 attributes under libver v114) the object's
attributes move to dense storage; h5py reads all of them."""
path = str(tmp_path / "dense.h5")
with h5py.File(path, "w", libver="v114") as f:
f.create_dataset("d", data=np.arange(3))
with clawhdf5.File(path, "r+") as f:
for i in range(20):
f["d"].attrs[f"a{i:02d}"] = np.full(i + 1, i, dtype="<i2")
with h5py.File(path, "r") as f:
assert sorted(f["d"].attrs.keys()) == [f"a{i:02d}" for i in range(20)]
for i in range(20):
np.testing.assert_array_equal(f["d"].attrs[f"a{i:02d}"], np.full(i + 1, i))
h5dump_reads(path)
def test_reads_never_see_a_half_written_edit(h5py, tmp_path):
"""Readers on other threads while one thread rewrites a dataset: every
read returns one whole version (all elements equal), never a mix."""
path = str(tmp_path / "race.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.zeros((64, 64)), chunks=(16, 16), compression="gzip")
errors = []
with clawhdf5.File(path, "r+") as f:
stop = threading.Event()
def read():
ds = f["d"]
while not stop.is_set():
a = ds[...]
if not (a == a.flat[0]).all():
errors.append(a)
return
readers = [threading.Thread(target=read) for _ in range(3)]
for t in readers:
t.start()
try:
for k in range(1, 25):
f["d"][...] = float(k)
finally:
stop.set()
for t in readers:
t.join()
np.testing.assert_array_equal(f["d"][...], np.full((64, 64), 24.0))
assert not errors, "a read saw a partly written dataset"
def test_close_releases_the_file(h5py, tmp_path):
path = str(tmp_path / "close.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
f = clawhdf5.File(path, "r+")
ds = f["d"]
f.close()
# The handle still reads the file as last written, but cannot edit it.
np.testing.assert_array_equal(ds[...], np.arange(4))
with pytest.raises(OSError, match="closed"):
ds[0] = 1
with h5py.File(path, "r+") as g:
g["d"][0] = 9
def _two_files(h5py, tmp_path):
"""d1/f.h5 and d2/f.h5: same name, different layouts (the review's
repro)."""
(tmp_path / "d1").mkdir()
(tmp_path / "d2").mkdir()
with h5py.File(tmp_path / "d1" / "f.h5", "w") as f:
f.create_dataset("x", data=np.arange(10, dtype="<i4"))
f.create_dataset("big", data=np.full(5000, 1.5))
with h5py.File(tmp_path / "d2" / "f.h5", "w") as f:
f.create_dataset("pad", data=np.full(3000, 2.5))
f.create_dataset("x", data=np.arange(10, dtype="<i4") + 500)
return tmp_path / "d1" / "f.h5", tmp_path / "d2" / "f.h5"
def _held_file_edited(h5py, held, other, other_bytes):
"""The edits went to `held`, planned from its own metadata; `other`
was not touched."""
assert other.read_bytes() == other_bytes, "the other file changed"
with h5py.File(held, "r") as f:
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
np.testing.assert_array_equal(f.attrs["note"], np.arange(50.0))
h5dump_reads(str(held))
def test_relative_path_and_chdir(h5py, tmp_path, monkeypatch):
"""A file opened by a relative path keeps being the file edited and read
after os.chdir (an edit was planned from the file the path named in the
new directory and written into the one open, corrupting it)."""
held, other = _two_files(h5py, tmp_path)
other_bytes = other.read_bytes()
monkeypatch.chdir(held.parent)
f = clawhdf5.File("f.h5", "r+")
ds = f["x"]
monkeypatch.chdir(other.parent)
ds[:] = np.full(10, 7, "<i4")
np.testing.assert_array_equal(ds[:5], np.full(5, 7)) # read from the held file
f.attrs["note"] = np.arange(50.0)
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
assert "pad" not in f
f.close()
_held_file_edited(h5py, held, other, other_bytes)
# 'w' writes where the path named when the file was opened.
monkeypatch.chdir(held.parent)
w = clawhdf5.File("new.h5", "w")
monkeypatch.chdir(other.parent)
w.create_dataset("d", data=np.arange(3))
w.close()
assert (held.parent / "new.h5").exists() and not (other.parent / "new.h5").exists()
def test_path_replaced_between_edits(h5py, tmp_path):
"""The path renamed away and another file put in its place between
edits: the edits go to the file held open, never mixed with the other."""
held, other = _two_files(h5py, tmp_path)
path = str(tmp_path / "f.h5")
os.replace(held, path)
moved = tmp_path / "moved.h5"
with clawhdf5.File(path, "r+") as f:
ds = f["x"]
ds[0] = 7
os.replace(path, moved)
shutil.copy(other, path)
other_bytes = open(path, "rb").read()
ds[:] = np.full(10, 7, "<i4")
f.attrs["note"] = np.arange(50.0)
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
assert "pad" not in f
from pathlib import Path
_held_file_edited(h5py, moved, Path(path), other_bytes)
def test_an_edit_releases_the_gil(h5py, tmp_path):
"""Another Python thread keeps running while a large edit is written:
had the edit held the GIL, the other thread would stall for the whole
edit (Rust code never yields it)."""
import time
path = str(tmp_path / "big.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", shape=(1024, 1024), dtype="<f8", chunks=(64, 64), compression="gzip")
value = np.random.default_rng(0).standard_normal((1024, 1024))
stamps = []
stop = threading.Event()
def spin():
while not stop.is_set():
stamps.append(time.perf_counter())
with clawhdf5.File(path, "r+") as f:
ds = f["d"]
t = threading.Thread(target=spin)
t.start()
try:
time.sleep(0.05)
t0 = time.perf_counter()
ds[...] = value
t1 = time.perf_counter()
finally:
stop.set()
t.join()
np.testing.assert_array_equal(ds[...], value)
during = [x for x in stamps if t0 <= x <= t1]
gaps = np.diff([t0] + during + [t1])
assert t1 - t0 > 0.1, "the edit is too quick to tell"
assert gaps.max() < 0.5 * (t1 - t0), (
f"the other thread stalled for {gaps.max():.3f} s of a {t1 - t0:.3f} s edit")
+29 -5
View File
@@ -178,15 +178,39 @@ def _write_fixture(h5py, path):
g.attrs["depth"] = np.int8(3) g.attrs["depth"] = np.int8(3)
@pytest.fixture(scope="module") @pytest.fixture(scope="module", params=["local", "r+", "http", "http-1k-blocks"])
def pair(h5py, tmp_path_factory): def pair(request, h5py, tmp_path_factory):
path = str(tmp_path_factory.mktemp("h5") / "fixture.h5") """The fixture file through h5py and through clawhdf5: opened locally
(read-only, and for editing: a copy, since editing locks the file), and
over HTTP range requests (a local server in this process) with the
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
through many small ranges."""
import shutil
from conftest import RangeServer
root = tmp_path_factory.mktemp("h5")
path = str(root / "fixture.h5")
_write_fixture(h5py, path) _write_fixture(h5py, path)
theirs = h5py.File(path, "r") theirs = h5py.File(path, "r")
ours = clawhdf5.File(path, "r") server = None
if request.param == "local":
ours = clawhdf5.File(path, "r")
elif request.param == "r+":
copy = str(root / "editable.h5")
shutil.copy(path, copy)
ours = clawhdf5.File(copy, "r+")
else:
server = RangeServer(root)
if request.param == "http":
ours = clawhdf5.File(server.url("fixture.h5"))
else:
ours = clawhdf5.File.open_url(server.url("fixture.h5"), block_size=1024)
yield ours, theirs, path yield ours, theirs, path
theirs.close() theirs.close()
ours.close() ours.close()
if server is not None:
server.close()
def _all_datasets(h5py, f): def _all_datasets(h5py, f):
@@ -416,7 +440,7 @@ def test_unsupported_types_are_errors_not_data(pair):
def test_boolean_masks_are_refused(pair): def test_boolean_masks_are_refused(pair):
ours, _, _ = pair ours, _, _ = pair
with pytest.raises(TypeError): with pytest.raises(NotImplementedError, match="mask"):
ours["num/le_i4_1d"][np.ones(37, dtype=bool)] ours["num/le_i4_1d"][np.ones(37, dtype=bool)]
+244
View File
@@ -0,0 +1,244 @@
"""Remote files: `clawhdf5.File(url)` / `File.open_url(url, ...)` read over
HTTP range requests (clawhdf5-remote's block cache), against a server in this
process (conftest.RangeServer). Values are compared with h5py reading the
same file locally; the rest checks what the server saw (only the blocks a
read needs are fetched), the failure modes (no range support, a missing
file, a file that changes, a server that goes away: errors, never wrong
data), and that the GIL is released while a read waits on the network."""
import os
import sys
import threading
import time
import numpy as np
import pytest
import clawhdf5
from conftest import RangeServer
def _write(h5py, path):
rng = np.random.default_rng(7)
with h5py.File(path, "w") as f:
f.create_dataset("contig", data=rng.standard_normal((400, 300)))
f.create_dataset(
"chunked",
data=rng.integers(0, 1000, size=(512, 512), dtype="<i4"),
chunks=(64, 64),
compression="gzip",
)
f.create_dataset("strings", data=["alpha", "beta", "gamma"], dtype=h5py.string_dtype())
g = f.create_group("grp")
g.create_dataset("small", data=np.arange(10, dtype="<u2"))
g.attrs["units"] = "m/s"
f.attrs["title"] = "remote test"
@pytest.fixture
def remote_file(h5py, tmp_path, range_server):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
return path, range_server.url("remote.h5"), range_server
def test_remote_reads_match_h5py(h5py, remote_file):
path, url, server = remote_file
with h5py.File(path, "r") as theirs, clawhdf5.File(url) as ours:
assert ours.filename == url
assert list(ours.keys()) == list(theirs.keys())
assert "grp/small" in ours and "nope" not in ours
for name in ["contig", "chunked", "grp/small"]:
np.testing.assert_array_equal(ours[name][...], theirs[name][...])
np.testing.assert_array_equal(ours[name][3:7], theirs[name][3:7])
np.testing.assert_array_equal(ours["chunked"][100:130, 200:300:3], theirs["chunked"][100:130, 200:300:3])
np.testing.assert_array_equal(ours["chunked"][[1, 70, 300], 5], theirs["chunked"][[1, 70, 300], 5])
assert list(ours["strings"][...]) == list(theirs["strings"][...])
assert ours["grp"].attrs["units"] == theirs["grp"].attrs["units"]
assert ours.attrs["title"] == theirs.attrs["title"]
assert ours["chunked"].maxshape == theirs["chunked"].maxshape
assert server.requests() >= 2
def test_a_small_read_fetches_only_its_blocks(h5py, remote_file):
"""With 4 KiB blocks, opening and reading one chunk of a 1 MB chunked
dataset costs a handful of requests and a few blocks, not the file."""
path, url, server = remote_file
size = os.path.getsize(path)
f = clawhdf5.File.open_url(url, block_size=4096)
opened = server.requests()
assert opened == 1, server.log
ds = f["chunked"]
got = ds[0:10, 0:10]
with h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(got, theirs["chunked"][0:10, 0:10])
stats = f.remote_stats
assert stats["bytes_fetched"] < size / 4, (stats, size)
assert server.requests() - opened <= 12, server.log
# A second read of the same region is served by the cache.
before = server.requests()
ds[0:10, 0:10]
assert server.requests() == before
assert f.remote_stats["hits"] > stats["hits"]
assert clawhdf5.File(str(path), "r").remote_stats is None
def test_server_without_range_support(h5py, tmp_path):
"""A server that ignores Range answers 200 with the whole file: that is
an OSError by default, and a whole download when allowed."""
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path, ranges=False)
try:
url = server.url("remote.h5")
with pytest.raises(OSError, match="range"):
clawhdf5.File(url)
with clawhdf5.File.open_url(url, allow_full_download=True) as ours, h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(ours["chunked"][...], theirs["chunked"][...])
np.testing.assert_array_equal(ours["contig"][5], theirs["contig"][5])
with pytest.raises(OSError):
clawhdf5.File.open_url(url, allow_full_download=True, max_full_download=1000)
finally:
server.close()
def test_errors_are_oserrors(remote_file):
_, url, server = remote_file
with pytest.raises(OSError, match="404"):
clawhdf5.File(server.url("missing.h5"))
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "r+")
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "w")
with pytest.raises(OSError, match="unsupported URL"):
clawhdf5.File("nosuchscheme://x/y.h5")
with pytest.raises(ValueError):
clawhdf5.File.open_url(url, block_size=0)
with pytest.raises(TypeError):
clawhdf5.File.open_url(url, no_such_option=1)
def test_object_store_urls_need_their_features():
"""The default wheel has no S3/GCS/Azure clients (aws-lc-rs builds C):
such a URL is an OSError naming the build feature."""
for url, feature in [("s3://bucket/k.h5", "s3"), ("gs://b/k.h5", "gcs"), ("az://c/k.h5", "azure")]:
try:
clawhdf5.File(url)
except OSError as e:
if "feature" in str(e):
assert f"`{feature}`" in str(e), str(e)
else:
pytest.fail(f"{url} opened")
def test_https_needs_the_https_feature():
"""The default wheel has no TLS stack (rustls needs ring, which builds C):
an https URL is an OSError that names the build feature."""
with pytest.raises(OSError) as e:
clawhdf5.File("https://127.0.0.1:1/x.h5")
msg = str(e.value)
# Built with `--features https` the error is the refused connection.
assert "https" in msg or "connect" in msg.lower() or "refused" in msg.lower(), msg
def test_a_changed_file_is_an_error_not_mixed_data(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=1024)
first = f["grp/small"][...]
# Rewrite the file with other values: new ETag, same name.
time.sleep(0.01)
with h5py.File(path, "w") as g:
g.create_dataset("contig", data=np.zeros((400, 300)))
with pytest.raises(OSError, match="changed"):
f["contig"][...]
np.testing.assert_array_equal(first, np.arange(10, dtype="<u2"))
def test_a_server_that_goes_away_is_an_error(h5py, tmp_path):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path)
f = clawhdf5.File.open_url(server.url("remote.h5"), block_size=1024, retries=0, timeout=2)
ds = f["contig"]
server.close()
with pytest.raises(OSError):
ds[...]
def test_threads_read_one_remote_file(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=2048)
with h5py.File(path, "r") as theirs:
expected = theirs["chunked"][...]
errors = []
def work(i):
try:
rows = slice((i * 37) % 400, (i * 37) % 400 + 64)
np.testing.assert_array_equal(f["chunked"][rows], expected[rows])
except Exception as e: # noqa: BLE001
errors.append(e)
threads = [threading.Thread(target=work, args=(i,)) for i in range(16)]
for t in threads:
t.start()
for t in threads:
t.join()
assert not errors, errors[:3]
def test_remote_reads_release_the_gil(h5py, remote_file):
"""A read waiting on a slow server lets other Python threads run: a
thread counting in a loop keeps counting (and never stalls for long)
while the main thread reads through requests that each take 0.2 s."""
_, url, server = remote_file
f = clawhdf5.File.open_url(url, block_size=1024, max_parallel=1)
ds = f["contig"]
server.delay = 0.2
server.delay_after = server.requests()
old = sys.getswitchinterval()
sys.setswitchinterval(0.001)
stop = threading.Event()
progress = {"n": 0, "worst": 0.0}
def spin():
last = time.perf_counter()
while not stop.is_set():
now = time.perf_counter()
progress["worst"] = max(progress["worst"], now - last)
last = now
progress["n"] += 1
t = threading.Thread(target=spin)
try:
t.start()
time.sleep(0.02)
t0 = time.perf_counter()
before = server.requests()
ds[0:2]
took = time.perf_counter() - t0
stop.set()
t.join()
finally:
sys.setswitchinterval(old)
assert server.requests() > before
assert took >= 0.2, took
assert progress["n"] > 1000, progress
# Held across a 0.2 s request, the spinner would stall that long.
assert progress["worst"] < 0.1, (progress, took)
def test_a_clawhdf5_written_file_reads_the_same_remotely(tmp_path, range_server):
path = tmp_path / "ours.h5"
data = np.arange(3000, dtype="<f8").reshape(100, 30)
with clawhdf5.File(str(path), "w") as f:
f.create_dataset("d", data=data, chunks=[10, 30], compression="gzip")
g = f.create_group("g")
g.create_dataset("i", data=np.arange(5, dtype="<i4"))
f.attrs["k"] = 3
with clawhdf5.File(range_server.url("ours.h5")) as f, clawhdf5.File(str(path)) as local:
np.testing.assert_array_equal(f["d"][...], data)
np.testing.assert_array_equal(f["d"][5:9, ::4], local["d"][5:9, ::4])
np.testing.assert_array_equal(f["g/i"][...], np.arange(5))
assert f.attrs["k"] == local.attrs["k"]
assert "file (read" in repr(f).lower() and "127.0.0.1" in repr(f)
@@ -1509,3 +1509,68 @@ fn unencodable_filters_are_unsupported() {
"a refused edit changed the file" "a refused edit changed the file"
); );
} }
fn fixture(dir: &Path, name: &str) -> std::path::PathBuf {
let path = dir.join(name);
std::fs::copy(
Path::new(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name),
&path,
)
.unwrap();
path
}
/// Zero extents on chunked datasets with no recorded maximum. The unfixed
/// editor left `chunk_zero_extent_no_maxshape.h5` (a 2.7.0-written file
/// resized to 1x0): a Fixed Array whose maximum, taken from the current
/// dimensions, has no chunks along one dimension, so every stride before it
/// is 0 — `h5rs check` panicked dividing by it and the next resize failed
/// with an internal error. Such a file must check clean and resize on; a
/// 2.7.0-written file taken through zero extents by the fixed editor must
/// check clean at every step and read the fill value where it grew.
#[test]
fn zero_extent_resizes_without_a_recorded_maximum() {
if !tools_ok() {
return;
}
let dir = tmpdir();
let path = fixture(dir.path(), "chunk_zero_extent_no_maxshape.h5");
check_tools(&path, true);
let mut ed = FileEditor::open(&path).unwrap();
ed.resize("d", &[0, 0]).unwrap();
ed.resize("z", &[0, 0]).unwrap();
// Their maximum is now what the index was laid out by (1 x 0).
ed.resize("d", &[1, 0]).unwrap();
assert!(matches!(
ed.resize("d", &[1, 1]),
Err(Error::InvalidArgument(_))
));
drop(ed);
check_tools(&path, true);
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
for shape in [[15, 15], [3, 2], [1, 1], [1, 0], [0, 0], [0, 20], [20, 20]] {
let mut ed = FileEditor::open(&path).unwrap();
ed.resize("d", &shape).unwrap();
ed.resize("z", &shape).unwrap();
drop(ed);
check_tools(&path, true);
}
let f = File::open(&path).unwrap();
for name in ["d", "z"] {
let d = f.dataset(name).unwrap();
assert_eq!(d.shape().unwrap(), [20, 20]);
assert!(d.read_f32().unwrap().iter().all(|&v| v == 0.0), "{name}");
}
assert_eq!(
py(&format!(
"import h5py\n\
with h5py.File({:?}) as f:\n\
\x20 print(int(abs(f['d'][()]).sum() + abs(f['z'][()]).sum()), f['d'].maxshape)",
path.to_str().unwrap()
)),
"0 (20, 20)"
);
}
+117 -9
View File
@@ -65,7 +65,8 @@ const MSG_FLAG_DONTSHARE: u8 = 0x04;
/// that may be mid-update. /// that may be mid-update.
/// ///
/// Every method is one self-contained edit: it re-reads the file's /// Every method is one self-contained edit: it re-reads the file's
/// metadata, applies the change, and syncs the file before returning. /// metadata (from the file it holds open, never by path), applies the
/// change, and syncs the file before returning.
/// ///
/// # What it can change /// # What it can change
/// ///
@@ -750,9 +751,17 @@ impl FileEditor {
/// consistent: a metadata cache image, paged or persistent free-space /// consistent: a metadata cache image, paged or persistent free-space
/// management, a multi-file driver, a file another writer has marked /// management, a multi-file driver, a file another writer has marked
/// open (superblock version 3 consistency flags). /// open (superblock version 3 consistency flags).
///
/// The path is only used to open the file: every edit is planned from
/// and written to the file opened here, even if the path is renamed,
/// replaced or (relative) resolved from another working directory
/// later. [`path`](Self::path) is the absolute path it had at open.
pub fn open<P: AsRef<Path>>(path: P) -> Result<Self, Error> { pub fn open<P: AsRef<Path>>(path: P) -> Result<Self, Error> {
let path = path.as_ref().to_path_buf(); let file = OpenOptions::new()
let file = OpenOptions::new().read(true).write(true).open(&path)?; .read(true)
.write(true)
.open(path.as_ref())?;
let path = std::fs::canonicalize(path.as_ref())?;
match file.try_lock() { match file.try_lock() {
Ok(()) => {} Ok(()) => {}
Err(TryLockError::WouldBlock) => { Err(TryLockError::WouldBlock) => {
@@ -768,16 +777,81 @@ impl FileEditor {
file, file,
free: FreeList::default(), free: FreeList::default(),
}; };
let f = File::open(&ed.path)?; let f = ed.plan_reader()?;
check_editable(&f)?; check_editable(&f)?;
Ok(ed) Ok(ed)
} }
/// The file's path. /// The file's absolute path when it was opened (it may have been
/// renamed since; the editor keeps editing the file it opened).
pub fn path(&self) -> &Path { pub fn path(&self) -> &Path {
&self.path &self.path
} }
/// A reader over the file this editor holds, as last written: the file
/// opened by [`open`](Self::open), not whatever its path names now.
///
/// It opens the file anew (read-only), so it does not share the
/// editor's lock and stays usable after the editor is dropped: on Linux
/// through `/proc/self/fd`, which reaches the held file even after its
/// path was renamed or replaced; elsewhere by the path the file had at
/// open, refused with [`Error::Io`] when that path no longer names the
/// held file (on Unix, compared by device and inode; Windows cannot
/// check). With the `mmap` feature the reader maps the file: edits
/// through the editor change the bytes it sees, so take a new reader
/// after each edit rather than reading through an old one while an edit
/// runs.
pub fn reader(&self) -> Result<File, Error> {
let dir = self.path.parent().map(Path::to_path_buf);
File::from_std_file(self.reopen()?, dir)
}
/// A new read-only open file description of the held file (see
/// [`reader`](Self::reader)).
fn reopen(&self) -> Result<std::fs::File, Error> {
#[cfg(target_os = "linux")]
{
use std::os::fd::AsRawFd;
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
if let Ok(f) = std::fs::File::open(proc) {
return Ok(f);
}
}
let f = std::fs::File::open(&self.path)?;
#[cfg(unix)]
{
use std::os::unix::fs::MetadataExt;
let (a, b) = (self.file.metadata()?, f.metadata()?);
if (a.dev(), a.ino()) != (b.dev(), b.ino()) {
return Err(Error::Io(std::io::Error::other(format!(
"{} no longer names the file being edited (renamed or replaced)",
self.path.display()
))));
}
}
Ok(f)
}
/// A reader over the held file for planning an edit (dropped before the
/// edit writes). On Linux a new open file description through
/// `/proc/self/fd`: a mapping of a clone of the held descriptor would
/// share its `flock`, and a process forked meanwhile (any
/// `std::process::Command` on another thread) would briefly keep the
/// lock alive after the editor is dropped. Elsewhere a clone of the
/// held descriptor, which follows the file wherever its path goes.
fn plan_reader(&self) -> Result<File, Error> {
let dir = self.path.parent().map(Path::to_path_buf);
#[cfg(target_os = "linux")]
{
use std::os::fd::AsRawFd;
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
if let Ok(f) = std::fs::File::open(proc) {
return File::from_std_file(f, dir);
}
}
File::from_std_file(self.file.try_clone()?, dir)
}
/// Bytes earlier edits of this editor freed that later ones can still /// Bytes earlier edits of this editor freed that later ones can still
/// reuse. /// reuse.
pub fn reusable_bytes(&self) -> u64 { pub fn reusable_bytes(&self) -> u64 {
@@ -794,7 +868,7 @@ impl FileEditor {
&mut self, &mut self,
op: impl FnOnce(&File, &mut Image<'_>) -> Result<R, Error>, op: impl FnOnce(&File, &mut Image<'_>) -> Result<R, Error>,
) -> Result<R, Error> { ) -> Result<R, Error> {
let f = File::open(&self.path)?; let f = self.plan_reader()?;
check_editable(&f)?; check_editable(&f)?;
let sb = f.superblock().clone(); let sb = f.superblock().clone();
let user_block = f.user_block_size(); let user_block = f.user_block_size();
@@ -877,7 +951,7 @@ impl FileEditor {
/// the fill value (`H5D__chunk_prune_by_extent`). /// the fill value (`H5D__chunk_prune_by_extent`).
pub fn resize(&mut self, path: &str, shape: &[u64]) -> Result<(), Error> { pub fn resize(&mut self, path: &str, shape: &[u64]) -> Result<(), Error> {
self.edit(|f, img| { self.edit(|f, img| {
let t = Target::load(f, path)?; let mut t = Target::load(f, path)?;
let dims = t.dims().to_vec(); let dims = t.dims().to_vec();
if shape.len() != dims.len() { if shape.len() != dims.len() {
return Err(Error::InvalidArgument(format!( return Err(Error::InvalidArgument(format!(
@@ -889,7 +963,12 @@ impl FileEditor {
if shape == dims.as_slice() { if shape == dims.as_slice() {
return Ok(()); return Ok(());
} }
let max = t.ds.max_dimensions.clone().unwrap_or_else(|| dims.clone()); // No maximum recorded means the current dimensions (see below).
let record_max = t.ds.max_dimensions.is_none();
let max =
t.ds.max_dimensions
.get_or_insert_with(|| dims.clone())
.clone();
for d in 0..dims.len() { for d in 0..dims.len() {
if shape[d] > max[d] { if shape[d] > max[d] {
return Err(Error::InvalidArgument(format!( return Err(Error::InvalidArgument(format!(
@@ -928,7 +1007,36 @@ impl FileEditor {
} }
put_uint(&mut dims_bytes[d * ls..], n, img.ls); put_uint(&mut dims_bytes[d * ls..], n, img.ls);
} }
hdr.patch(img, i, first, &dims_bytes)?; if !record_max {
hdr.patch(img, i, first, &dims_bytes)?;
} else {
// No maximum recorded (clawhdf5's writer, for a dataset
// created without a maxshape). libhdf5 never writes such a
// dataspace: `H5S_set_extent_simple` records the maximum,
// equal to the dimensions when none is given. Reading one,
// libhdf5 takes the maximum to be the *current* dimensions
// (`H5S_extent_get_dims`), so changing them would also
// change the maximum the chunk index was built with — the
// Fixed Array linearises chunks by it — and move every
// existing chunk. Record the maximum libhdf5 would have
// written, the dimensions before this resize, so the index
// keeps its layout and the dataset can grow back to them.
let body_len = first + dims.len() * ls;
if body.len() < body_len || body[2] & !0x01 != 0 {
return Err(Error::Unsupported("dataspace message layout".into()));
}
let mut new_body = body[..first].to_vec();
new_body[2] |= 0x01;
new_body.extend_from_slice(&dims_bytes);
let at = new_body.len();
new_body.resize(at + dims.len() * ls, 0);
for (d, &n) in dims.iter().enumerate() {
put_uint(&mut new_body[at + d * ls..], n, img.ls);
}
let (flags, corder) = (hdr.msgs[i].flags, hdr.msgs[i].corder);
hdr.delete(img, i)?;
hdr.insert(img, MSG_DATASPACE, flags, &new_body, corder)?;
}
let fill = fill_info(img, &hdr)?; let fill = fill_info(img, &hdr)?;
hdr.finish(img)?; hdr.finish(img)?;
let expand = shape.iter().zip(&dims).any(|(n, o)| n > o); let expand = shape.iter().zip(&dims).any(|(n, o)| n > o);
+32
View File
@@ -446,6 +446,38 @@ impl File {
} }
} }
/// A reader over an already open file (the file itself, not whatever
/// its path names now): mapped with the `mmap` feature, else read into
/// memory. `base_dir` resolves external Virtual Dataset sources.
pub(crate) fn from_std_file(
file: std::fs::File,
base_dir: Option<std::path::PathBuf>,
) -> Result<Self, Error> {
#[cfg(feature = "mmap")]
let mut f = {
let reader = clawhdf5_io::MmapReader::from_file(file).map_err(Error::Io)?;
let (data, superblock) = FileData::new(Backing::Mmap(reader))?;
Self {
data,
superblock,
chunk_cache: ChunkCache::new(),
base_dir: None,
vds_resolver: None,
}
};
#[cfg(not(feature = "mmap"))]
let mut f = {
use std::io::{Read, Seek, SeekFrom};
let mut file = file;
let mut bytes = Vec::new();
file.seek(SeekFrom::Start(0)).map_err(Error::Io)?;
file.read_to_end(&mut bytes).map_err(Error::Io)?;
Self::from_bytes(bytes)?
};
f.base_dir = base_dir;
Ok(f)
}
/// Open an HDF5 file by reading it entirely into memory. /// Open an HDF5 file by reading it entirely into memory.
/// ///
/// This is the pre-mmap behaviour and is useful when memory-mapping is /// This is the pre-mmap behaviour and is useful when memory-mapping is
@@ -0,0 +1,287 @@
//! `FileEditor::resize` on chunked datasets whose dataspace records no
//! maximum dimensions, as clawhdf5's writer stored a dataset created without
//! a `maxshape` up to 2.7.0 (`fixtures/chunked_no_maxshape_v2_7_0.h5`). libhdf5 never writes such a dataspace (`H5S_set_extent_simple`
//! always records the maximum, equal to the dimensions when none is given),
//! and its Fixed Array chunk index linearises chunks by the maximum
//! dimensions. The editor therefore records the maximum libhdf5 would have
//! written (the dimensions the index was built with) before it changes the
//! current ones, so existing chunks stay where the index put them and the
//! dataset can grow back to its original extent.
//!
//! The writer now records the maximum too, so h5py can resize what it writes.
//!
//! Checked against a model of the expected values, with our reader and with
//! h5py (`CLAWHDF5_PYTHON`; skipped without it unless
//! `CLAWHDF5_REQUIRE_INTEROP=1`), on files clawhdf5 (old and new) and h5py
//! wrote.
use std::path::{Path, PathBuf};
use std::process::Command;
use clawhdf5::{Error, File, FileBuilder, FileEditor};
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn h5py_ok() -> bool {
let ok = Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.is_ok_and(|o| o.status.success());
if !ok {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but h5py/numpy is not available"
);
eprintln!("SKIP (h5py part): h5py/numpy not available");
}
ok
}
fn py(script: &str) -> String {
let o = Command::new(python())
.args(["-c", script])
.output()
.expect("run python");
assert!(
o.status.success(),
"python failed:\n{script}\nSTDERR: {}",
String::from_utf8_lossy(&o.stderr)
);
String::from_utf8_lossy(&o.stdout).trim().to_string()
}
/// Row-major values of a 2-D model after resizing `data` (shape `old`) to
/// `new`: kept elements keep their values, new ones are 0 (the fill value).
fn resized(data: &[f32], old: [u64; 2], new: [u64; 2]) -> Vec<f32> {
let mut out = vec![0f32; (new[0] * new[1]) as usize];
for r in 0..old[0].min(new[0]) {
for c in 0..old[1].min(new[1]) {
out[(r * new[1] + c) as usize] = data[(r * old[1] + c) as usize];
}
}
out
}
/// Our reader and (when available) h5py read `expect` at `shape`.
fn check(path: &Path, name: &str, shape: [u64; 2], expect: &[f32], with_h5py: bool) {
let f = File::open(path).unwrap();
let d = f.dataset(name).unwrap();
assert_eq!(d.shape().unwrap(), shape);
assert_eq!(
d.read_f32().unwrap(),
expect,
"{name}: our reader at {shape:?}"
);
if with_h5py {
let got = py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'r') as f:\n\
\x20 d = f[{name:?}][()]\n\
print(d.shape, ','.join(repr(float(x)) for x in d.ravel()))",
p = path.to_str().unwrap()
));
let want = format!(
"({}, {}) {}",
shape[0],
shape[1],
expect
.iter()
.map(|x| format!("{:?}", f64::from(*x)))
.collect::<Vec<_>>()
.join(",")
);
assert_eq!(got, want.trim(), "{name}: h5py at {shape:?}");
}
}
/// Resize `name` (20 x 20, values 0..400) through a sequence of shrinks,
/// zero extents and growth back, checking every step.
fn run(path: &Path, name: &str, with_h5py: bool) {
let orig: Vec<f32> = (0..400).map(|i| i as f32).collect();
let mut shape = [20u64, 20];
let mut data = orig.clone();
check(path, name, shape, &data, with_h5py);
for next in [
[15, 15],
[3, 2],
[20, 20],
[1, 1],
[1, 0],
[0, 0],
[7, 20],
[20, 13],
[20, 20],
] {
let mut ed = FileEditor::open(path).unwrap();
ed.resize(name, &next).unwrap();
drop(ed);
data = resized(&data, shape, next);
shape = next;
check(path, name, shape, &data, with_h5py);
}
// The maximum is the extent the dataset was created with.
let mut ed = FileEditor::open(path).unwrap();
assert!(matches!(
ed.resize(name, &[21, 20]),
Err(Error::InvalidArgument(_))
));
drop(ed);
let f = File::open(path).unwrap();
assert_eq!(
f.dataset(name).unwrap().max_dimensions().unwrap(),
Some(vec![20, 20])
);
drop(f);
// A shrink keeps the values it keeps.
let mut ed = FileEditor::open(path).unwrap();
let vals: Vec<f32> = orig.iter().map(|v| v + 0.5).collect();
ed.write_values(name, &clawhdf5::Selection::All, &vals)
.unwrap();
ed.resize(name, &[15, 15]).unwrap();
drop(ed);
check(
path,
name,
[15, 15],
&resized(&vals, [20, 20], [15, 15]),
with_h5py,
);
}
fn fixture(dir: &Path, name: &str) -> PathBuf {
let path = dir.join(name);
std::fs::copy(
Path::new(env!("CARGO_MANIFEST_DIR"))
.join("tests/fixtures")
.join(name),
&path,
)
.unwrap();
path
}
/// Files clawhdf5 2.7.0 wrote: a Fixed Array index (Single Chunk for `s`)
/// and a dataspace with no maximum. Shrinking scrambled the values
/// (released in 2.7.0's `FileEditor`, PR #18).
#[test]
fn resize_without_stored_maxshape_keeps_values() {
let with_h5py = h5py_ok();
for name in ["d", "z"] {
let dir = tempfile::tempdir().unwrap();
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
run(&path, name, with_h5py);
}
// A single-chunk dataset and an empty one keep their extents as maxima.
let dir = tempfile::tempdir().unwrap();
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
let mut ed = FileEditor::open(&path).unwrap();
ed.resize("s", &[2, 4]).unwrap();
ed.resize("s", &[3, 4]).unwrap();
assert!(matches!(
ed.resize("s", &[4, 4]),
Err(Error::InvalidArgument(_))
));
assert!(matches!(
ed.resize("e", &[1, 5]),
Err(Error::InvalidArgument(_))
));
ed.resize("e", &[0, 3]).unwrap();
drop(ed);
let f = File::open(&path).unwrap();
let s = f.dataset("s").unwrap();
assert_eq!(s.max_dimensions().unwrap(), Some(vec![3, 4]));
let mut want: Vec<i32> = (0..12).collect();
want[8..].fill(0);
assert_eq!(s.read_i32().unwrap(), want);
assert_eq!(
f.dataset("e").unwrap().max_dimensions().unwrap(),
Some(vec![0, 5])
);
}
fn written(path: &Path, deflate: bool) {
let data: Vec<f32> = (0..400).map(|i| i as f32).collect();
let mut b = FileBuilder::new();
let d = b
.create_dataset("d")
.with_f32_data(&data)
.with_shape(&[20, 20])
.with_chunks(&[6, 6]);
if deflate {
d.with_deflate(4);
}
b.write(path).unwrap();
}
/// clawhdf5's writer now records the maximum, as libhdf5 does.
#[test]
fn resize_file_written_without_maxshape_keeps_values() {
let with_h5py = h5py_ok();
for deflate in [false, true] {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("cw.h5");
written(&path, deflate);
let f = File::open(&path).unwrap();
assert_eq!(
f.dataset("d").unwrap().max_dimensions().unwrap(),
Some(vec![20, 20])
);
drop(f);
run(&path, "d", with_h5py);
}
}
/// h5py resizing a file clawhdf5 wrote without a maxshape keeps its values
/// (it scrambled them while the writer recorded no maximum).
#[test]
fn h5py_resizes_what_clawhdf5_writes() {
if !h5py_ok() {
return;
}
for deflate in [false, true] {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("cw.h5");
written(&path, deflate);
let out = py(&format!(
"import h5py, numpy as np\n\
exp = np.arange(400, dtype='f4').reshape(20, 20)\n\
with h5py.File({p:?}, 'r+') as f:\n\
\x20 f['d'].resize((15, 15))\n\
\x20 ok = np.array_equal(f['d'][()], exp[:15, :15])\n\
\x20 f['d'].resize((20, 20))\n\
\x20 back = f['d'][()]\n\
want = np.zeros((20, 20), 'f4'); want[:15, :15] = exp[:15, :15]\n\
print(ok and np.array_equal(back, want))",
p = path.to_str().unwrap()
));
assert_eq!(out, "True");
let mut want = vec![0f32; 400];
for r in 0..15 {
for c in 0..15 {
want[r * 20 + c] = (r * 20 + c) as f32;
}
}
check(&path, "d", [20, 20], &want, false);
}
}
/// h5py's files record the maximum; the same sequence must hold.
#[test]
fn resize_h5py_file_without_maxshape_keeps_values() {
if !h5py_ok() {
return;
}
for libver in ["earliest", "v110", "latest"] {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hp.h5");
py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'w', libver=({libver:?}, 'latest')) as f:\n\
\x20 f.create_dataset('d', data=np.arange(400, dtype='f4').reshape(20, 20), chunks=(6, 6))",
p = path.to_str().unwrap()
));
run(&path, "d", true);
}
}
+56
View File
@@ -162,3 +162,59 @@ fn shrink_then_grow_reads_fill() {
raw[..4].fill(0.5); raw[..4].fill(0.5);
assert_eq!(f.dataset("raw").unwrap().read_f64().unwrap(), raw); assert_eq!(f.dataset("raw").unwrap().read_f64().unwrap(), raw);
} }
/// The editor plans every edit from the file it holds open, never by
/// re-opening its path: with the path renamed away and another file put in
/// its place, edits go to the held file, planned from its own metadata,
/// and the file now at the path is untouched (planning from it and writing
/// into the held file corrupted the held one).
#[test]
fn edits_go_to_the_file_held_not_the_path() {
let dir = tempfile::tempdir().unwrap();
let a = dir.path().join("a.h5");
let b = dir.path().join("b.h5");
let mut fb = FileBuilder::new();
fb.create_dataset("x")
.with_i32_data(&[0; 10])
.with_shape(&[10]);
fb.create_dataset("big")
.with_f64_data(&[1.5; 5000])
.with_shape(&[5000]);
fb.set_attr("title", AttrValue::String("a".into()));
fb.write(&a).unwrap();
let mut fb = FileBuilder::new();
fb.create_dataset("pad")
.with_f64_data(&[2.5; 3000])
.with_shape(&[3000]);
fb.create_dataset("x")
.with_i32_data(&[500; 10])
.with_shape(&[10]);
fb.write(&b).unwrap();
let mut ed = FileEditor::open(&a).unwrap();
assert!(ed.path().is_absolute());
let moved = dir.path().join("moved.h5");
std::fs::rename(&a, &moved).unwrap();
std::fs::rename(&b, &a).unwrap();
let other = std::fs::read(&a).unwrap();
ed.write_values("x", &Selection::All, &[7i32; 10]).unwrap();
let vals: Vec<f64> = (0..50).map(f64::from).collect();
ed.set_attr("/", "note", &AttrValue::F64Array(vals.clone()))
.unwrap();
// The editor's own reader sees the held file.
let r = ed.reader().unwrap();
assert_eq!(r.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
assert!(r.dataset("pad").is_err());
drop(r);
drop(ed);
assert!(
std::fs::read(&a).unwrap() == other,
"the file at the path changed"
);
let f = File::open(&moved).unwrap();
assert_eq!(f.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
assert_eq!(f.dataset("big").unwrap().read_f64().unwrap(), [1.5; 5000]);
assert!(matches!(f.root().attr("note").unwrap(), Some(AttrValue::F64Array(v)) if v == vals));
}
Binary file not shown.
Binary file not shown.
@@ -966,7 +966,10 @@ fn skipped_optional_filters_are_masked_as_libhdf5_masks_them() {
/// Files whose chunks all compress are written exactly as before optional /// Files whose chunks all compress are written exactly as before optional
/// filters could be skipped: every mask is 0 and nothing else changed. The /// filters could be skipped: every mask is 0 and nothing else changed. The
/// hashes are of the files the writer produced before that change. /// hashes are of the files the writer produced before that change, except
/// that a chunked dataset without a maxshape now records its maximum
/// dimensions (8 bytes per dimension; `lzf_fixed`, `lzf_single`, and
/// `blosc_fixed`).
#[cfg(feature = "lzf")] #[cfg(feature = "lzf")]
#[test] #[test]
fn files_whose_chunks_all_compress_are_unchanged() { fn files_whose_chunks_all_compress_are_unchanged() {
@@ -985,7 +988,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
.with_chunks(&[500]) .with_chunks(&[500])
.with_lzf(); .with_lzf();
}, },
(3965, 449169442), (3973, 452644487),
), ),
( (
"lzf_ea_noshuffle", "lzf_ea_noshuffle",
@@ -1017,7 +1020,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
.with_chunks(&[3000]) .with_chunks(&[3000])
.with_lzf(); .with_lzf();
}, },
(546, 690805477), (554, 1394027497),
), ),
]; ];
#[cfg(feature = "blosc")] #[cfg(feature = "blosc")]
@@ -1029,7 +1032,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
.with_chunks(&[1024]) .with_chunks(&[1024])
.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte); .with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
}, },
(2776, 4278611376), (2784, 1180133244),
)); ));
for (name, build, want) in &cases { for (name, build, want) in &cases {
let mut fb = clawhdf5::FileBuilder::new(); let mut fb = clawhdf5::FileBuilder::new();
+16 -1
View File
@@ -12,6 +12,10 @@ done on branch `feat/p3-m5-swmr-reader`, with its own design in
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is [`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
next. Every count in §1–§2 was next. Every count in §1–§2 was
object stores) and URLs in `h5rs` (see the M3 status below); the Python
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
change. Progress: M1, first part (the `Storage` trait and the metadata change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done; parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet. group B-tree v2 lookups, dense groups and the facade are not converted yet.
@@ -477,7 +481,7 @@ fast path within benchmark noise.
a request counter exposed for tests and users. a request counter exposed for tests and users.
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it. - Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the - *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
Python bindings, with these choices: Python bindings (done 2026-09-27, below), with these choices:
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of - A new crate, `clawhdf5-remote`, instead of a `remote` feature of
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io` `clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
sits below the facade. sits below the facade.
@@ -514,6 +518,17 @@ fast path within benchmark noise.
same work without a cache: 141 936 requests. Per file: A lists in 2 same work without a cache: 141 936 requests. Per file: A lists in 2
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
6.4 MB: 35 001 object headers spread over the file). 6.4 MB: 35 001 object headers spread over the file).
- *Status 2026-09-27, Python bindings:* done on branch
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
`File.open_url(url, **options)` (cache and HTTP options) go through
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
parsing (path lookups, headers, attributes, listings, the global heap)
moved from `File::as_bytes` to `File::storage()` and the `*_in`
functions, and every file access, metadata included, runs with the GIL
released. Checked by running the whole read-vs-h5py suite over an
in-process range server (1 MiB and 1 KiB blocks) and by request counts
in `crates/clawhdf5-py/tests/test_remote.py`.
**M4 — wasm lazy loading (1–2 weeks).** **M4 — wasm lazy loading (1–2 weeks).**
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a - `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
+87 -6
View File
@@ -30,6 +30,33 @@ h5py's SWMR reader). A file still being written is read with
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file `File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
at its length at open and is not meant for files that change while open. at its length at open and is not meant for files that change while open.
## Shrinking a chunked dataset with no recorded maximum scrambled it
**Status:** fixed 2026-09-27, before any release (`FileEditor::resize`
shipped on main in PR #18, a4c2ace). Files the writer produced before the
fix still lack the maximum; see *Existing files*.
clawhdf5's writer stored no maximum dimensions for a chunked dataset
created without a `maxshape` (libhdf5 always stores one, equal to the
dimensions when none is given). With none recorded, the maximum is the
current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index
places chunks by the maximum. `FileEditor::resize` changed only the
current dimensions, so a shrink moved every existing chunk and every
reader returned wrong values; after a shrink the dataset could not grow
back. Found by the review of the Python editing work.
**Fix:** before the first resize of such a dataset the editor records the
maximum libhdf5 would have written (the dimensions the index was built
with); the writer now records it for every chunked dataset. **Test:**
`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:**
datasets written before the fix have no recorded maximum. The fixed
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
does not** — it scrambles them the same way, and lets them grow past
their Fixed Array. Resize them once with a fixed `FileEditor` (a resize
to the same shape changes nothing; shrink and grow back) before letting
libhdf5 resize them. A dataset already shrunk by the unfixed
editor holds misplaced chunks; rewrite it from a good copy.
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768 ## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to **Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to
@@ -139,6 +166,57 @@ space it leaves is too small for its next, larger version.
**No journal.** A crash while an edit patches existing structures can leave **No journal.** A crash while an edit patches existing structures can leave
the file inconsistent; see the `FileEditor` documentation. the file inconsistent; see the `FileEditor` documentation.
**Renamed files outside Linux.** Edits always go to the file the editor
opened. `FileEditor::reader()` (which the Python `'r+'` handle reads
through) reopens it through `/proc/self/fd` on Linux; elsewhere it reopens
the path, and after the path was renamed or replaced it fails (on Unix;
Windows cannot tell and would read whatever the path names).
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
**Status:** open (added 2026-09-27). The Python bindings edit through
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
before anything is written. On top of them:
- **No new or deleted objects:** `create_dataset`/`create_group` in an
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
`NotImplementedError` (the editor changes values, shapes and
attributes only). Mode `'a'` works on an existing file only.
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
whole elements of the same structured dtype are), HDF5 array-type
elements, variable-length data, strings padded with spaces or
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
ones, h5py's, are writable), compounds containing such strings, null
dataspaces, index-list writes of more than 2²² elements (write them
in slices), and boolean-mask keys (`ds[mask] = v`, and mask reads;
h5py supports both).
- **`str` attributes are fixed-length UTF-8**, where h5py writes
variable-length strings: h5py reads them back as `bytes`
(`numpy.bytes_`), not `str`.
- **Numeric conversion follows libhdf5's native-order results, not its
bugs.** Arrays are converted as libhdf5 converts them (integers
saturate, floats are truncated toward zero and clipped), checked value by
value against h5py 3.16 / HDF5 2.0 in
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
differs from h5py on purpose:
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
minimum or 2⁶³ depending on the type);
- when the dataset or the array is not in native byte order, libhdf5's
"soft" conversions store a float in (-1, 0) as the integer minimum and
wrap an unsigned value too large for the signed type of the same size
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
native order;
- libhdf5's native casts that are undefined in C: half floats into
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
integers (→ minimum), a float equal to the integer maximum rounded up
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
- **Each edit reopens the file** (a new memory map) so that reads see it;
reads from other threads wait while an edit is written.
## Selection reads that decode more than the selection ## Selection reads that decode more than the selection
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so **Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
@@ -867,9 +945,9 @@ cache, but:
`read_*_zerocopy`) need the file in memory and answer `read_*_zerocopy`) need the file in memory and answer
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()` `FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
panics for such a file (`File::contiguous_bytes()` is the fallible form). panics for such a file (`File::contiguous_bytes()` is the fallible form).
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a `LazyFile`, `MmapFile` and the wasm bindings still read a whole file
whole file (`h5rs` reads through `File::storage`, and takes URLs with its (`h5rs` and the Python bindings read through `File::storage`, and take
`remote` feature). URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
- The file's length is read once, at open: a growing file (SWMR) is not - The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`. grows is `RemoteError::FileChanged`.
@@ -885,9 +963,12 @@ cache, but:
**Status:** open (added 2026-09-26, milestone M3 of **Status:** open (added 2026-09-26, milestone M3 of
`docs/design/range-reads.md`). `docs/design/range-reads.md`).
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3) - **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
parses through `File::as_bytes`, which a remote file does not have; the milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
wasm reader's `openUrl` is milestone M4. the default wheel reads plain `http://` only: `https://` needs a wheel
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise). - **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead. is not implemented, and only the first block is read ahead.
+9 -3
View File
@@ -131,14 +131,17 @@ run_step "cargo clippy (h5rs remote)" cargo clippy \
# js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C. # js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C.
# clawhdf5-remote is checked by default (plain HTTP) and with its # clawhdf5-remote is checked by default (plain HTTP) and with its
# object-store feature, and h5rs with URL support (remote); the https # object-store feature, and h5rs with URL support (remote); the https
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. # (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. The
# Python bindings (clawhdf5-py, remote reads over plain HTTP) are checked too:
# their https/s3/gcs/azure features are opt-in for the same reason.
no_c_in_default_build() { no_c_in_default_build() {
local entry crate features found=0 local entry crate features found=0
for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \ for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \
clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \ clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \
clawhdf5-tools \ clawhdf5-tools \
clawhdf5-wasm \ clawhdf5-wasm \
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote; do clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote \
clawhdf5-py; do
crate=${entry%%:*} crate=${entry%%:*}
features=() features=()
[ "$entry" != "$crate" ] && features=(--features "${entry#*:}") [ "$entry" != "$crate" ] && features=(--features "${entry#*:}")
@@ -269,7 +272,10 @@ python_package() {
-i "$PYTHON" \ -i "$PYTHON" \
--out "$out/wheel" || return 1 --out "$out/wheel" || return 1
"$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1 "$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \ # The editing tests run `h5rs check` on every file they edit.
cargo build -q -p clawhdf5-tools || return 1
CLAWHDF5_H5RS="${CARGO_TARGET_DIR:-$root/target}/debug/h5rs" \
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
"$root/crates/clawhdf5-py/tests" "$root/crates/clawhdf5-py/tests"
} }
if "$PYTHON" -m maturin --version >/dev/null 2>&1 && "$PYTHON" -c "import pytest" >/dev/null 2>&1; then if "$PYTHON" -m maturin --version >/dev/null 2>&1 && "$PYTHON" -c "import pytest" >/dev/null 2>&1; then