Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python
# Conflicts: # CHANGELOG.md # docs/design/range-reads.md # docs/known-issues.md
This commit is contained in:
+142
@@ -72,6 +72,148 @@ Design: `docs/design/swmr.md`.
|
|||||||
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
|
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
|
||||||
groups or attributes (a SWMR writer cannot add objects or attributes).
|
groups or attributes (a SWMR writer cannot add objects or attributes).
|
||||||
|
|
||||||
|
### Correctness: edits planned from another file after a rename or `chdir` (2026-09-27)
|
||||||
|
- **`FileEditor` planned each edit by re-opening its path but wrote
|
||||||
|
through the file it held open** (fixed 2026-09-27; on main since PR #18,
|
||||||
|
no release). When the path came to name another file between edits — a
|
||||||
|
rename or replacement, or, for a relative path, a change of working
|
||||||
|
directory — an edit was laid out from the other file's metadata and
|
||||||
|
written into the held one, corrupting it (h5py: "invalid dataset size,
|
||||||
|
likely file corruption"). The Python `'r+'` handle had the same flaw in
|
||||||
|
its reads: it reopened the path after every edit, so reads came from the
|
||||||
|
other file. The editor now plans every edit from the file it holds, and
|
||||||
|
its path is canonicalised at open. New `FileEditor::reader()` opens the
|
||||||
|
held file anew for reading (on Linux through `/proc/self/fd`, so it
|
||||||
|
follows a renamed file; elsewhere by the path, refused when the path no
|
||||||
|
longer names the held file), without sharing the editor's lock; the
|
||||||
|
Python handle reads through it, and a `'w'` file is written at the
|
||||||
|
absolute path it was opened with. Tests: `edit_tests.rs`'s
|
||||||
|
`edits_go_to_the_file_held_not_the_path`; `test_edit.py`'s
|
||||||
|
`test_relative_path_and_chdir` and `test_path_replaced_between_edits`
|
||||||
|
(the review's repro).
|
||||||
|
|
||||||
|
### Correctness: zero extents in Fixed/Extensible Array chunk indexes (2026-09-27)
|
||||||
|
- **A chunked dataset whose maximum (or, with none recorded, current)
|
||||||
|
extent is 0 along a dimension made the reader divide by zero** (fixed
|
||||||
|
2026-09-27): `h5rs check` panicked ("attempt to divide by zero",
|
||||||
|
`chunk_grid.rs`) and the next `FileEditor::resize` failed with an
|
||||||
|
internal error. The unfixed editor produced such files by resizing a
|
||||||
|
clawhdf5-written dataset to a zero extent (12 of 30 extra random-edit
|
||||||
|
seeds on clawhdf5-written files). Such an index has no slot for any
|
||||||
|
chunk of the dataset, and `ChunkGrid::offsets` now says so instead of
|
||||||
|
dividing by the zero stride. Tests: `chunk_grid`'s
|
||||||
|
`zero_extent_has_no_chunks`, `edit_interop.rs`'s
|
||||||
|
`zero_extent_resizes_without_a_recorded_maximum` (a file the unfixed
|
||||||
|
editor left checks clean and resizes on; a 2.7.0-written file through
|
||||||
|
zero extents checks clean at each step), and `test_edit.py`'s random
|
||||||
|
edits on seeds 10 to 39 of a clawhdf5-written file.
|
||||||
|
|
||||||
|
### Correctness: resizing chunked datasets with no recorded maximum (2026-09-27)
|
||||||
|
- **`FileEditor::resize` scrambled the values of a chunked dataset whose
|
||||||
|
dataspace records no maximum dimensions when it shrank it** (fixed
|
||||||
|
2026-09-27). The editor shipped on main in PR #18 (a4c2ace) and reached
|
||||||
|
Python as `Dataset.resize` in `'r+'` files; no release has it. clawhdf5's
|
||||||
|
writer stores such a dataspace for every chunked dataset created without
|
||||||
|
a `maxshape`, with a Fixed Array (or Single Chunk) chunk index. The
|
||||||
|
editor patched only the current dimensions, and with no maximum
|
||||||
|
recorded the maximum is the current dimensions — which is also what the
|
||||||
|
Fixed Array linearises chunks by — so a shrink moved every chunk after
|
||||||
|
the first row: h5py, h5dump and our reader all read wrong values
|
||||||
|
without complaint (20x20, chunks 6x6, resized to 15x15: row 6 read
|
||||||
|
`0 0 0 0 0 0 120 ...`). After a shrink the dataset could not grow back
|
||||||
|
either (`3 exceeds the maximum 0`). libhdf5 itself never writes such a
|
||||||
|
dataspace (`H5S_set_extent_simple` records the maximum, equal to the
|
||||||
|
dimensions when none is given); reading one, `H5S_extent_get_dims`
|
||||||
|
reports the current dimensions as the maximum and `H5S_set_extent`
|
||||||
|
checks against no maximum at all, so libhdf5's own `H5Dset_extent`
|
||||||
|
scrambles such a file the same way (and lets it grow past its Fixed
|
||||||
|
Array). The editor now records the maximum libhdf5 would have written
|
||||||
|
— the dimensions before the first resize, the ones the index was built
|
||||||
|
with — then changes the current ones (the dataspace message grows by one
|
||||||
|
length per dimension and moves in the header when it has to). The
|
||||||
|
dataset then shrinks, grows back to that extent and refuses more, as
|
||||||
|
one libhdf5 wrote would. The writer (`FileBuilder`) now records the
|
||||||
|
maximum of every chunked dataset too, as libhdf5 does, so h5py can
|
||||||
|
resize what it writes (8 more bytes per dimension). Tests:
|
||||||
|
`crates/clawhdf5/tests/edit_resize_interop.rs` (a 2.7.0-written fixture,
|
||||||
|
new `FileBuilder` files and h5py files through shrinks, zero extents
|
||||||
|
and growth, checked against a model with our reader and h5py; h5py
|
||||||
|
resizing a `FileBuilder` file) and `test_edit.py`'s numpy-model checks.
|
||||||
|
|
||||||
|
### Python bindings: in-place editing (2026-09-27)
|
||||||
|
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
|
||||||
|
file for editing through `clawhdf5::FileEditor`, holding its exclusive
|
||||||
|
lock until `close()`:
|
||||||
|
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
|
||||||
|
one increasing index list) and broadcasting (numpy's rules for slices
|
||||||
|
and integers, allowing extra leading length-1 axes; the exact shape for
|
||||||
|
an index list, a scalar only where h5py expands it). A numpy array is
|
||||||
|
converted to the dataset's dtype as libhdf5 converts it in native byte
|
||||||
|
order (integers saturate; floats are truncated toward zero and clipped;
|
||||||
|
integers go into h5py's bool enum by value, as libhdf5 stores them);
|
||||||
|
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
|
||||||
|
does. NaN into an integer dataset is a `ValueError`.
|
||||||
|
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
|
||||||
|
and errors (`TypeError` for a dataset that is not chunked).
|
||||||
|
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
|
||||||
|
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
|
||||||
|
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
|
||||||
|
enum, complex as the `r`/`i` compound), except that `str` becomes
|
||||||
|
fixed-length UTF-8.
|
||||||
|
- Every edit is written and synced before it returns, then the file is
|
||||||
|
reopened: datasets and attrs objects taken earlier see new shapes and
|
||||||
|
attributes, and reads on other threads wait while an edit is written.
|
||||||
|
- What the editor cannot do raises `NotImplementedError` and writes
|
||||||
|
nothing (deleting attributes or objects, creating datasets or groups,
|
||||||
|
compound fields by name, variable-length data, ...;
|
||||||
|
`docs/known-issues.md`).
|
||||||
|
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
|
||||||
|
- Boolean-mask keys (`ds[mask]`, `ds[mask] = v`), which h5py supports,
|
||||||
|
raise `NotImplementedError` (they raised `TypeError`).
|
||||||
|
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
|
||||||
|
by h5py and by clawhdf5, and both read back through h5py after every
|
||||||
|
edit, on files h5py writes with `libver` earliest, v114 and latest and on
|
||||||
|
a clawhdf5-written one: a fixed sequence over every chunk index kind,
|
||||||
|
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
|
||||||
|
and compound types, and 16 random sequences of 40 edits (writes,
|
||||||
|
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
|
||||||
|
and leave its file unchanged. A matrix of every numeric source dtype into
|
||||||
|
every numeric dataset dtype at the edge values, dense attribute storage,
|
||||||
|
locking, objects seeing each other's edits, readers racing a writer
|
||||||
|
(never a partly written dataset). Every edited file must pass `h5dump`
|
||||||
|
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
|
||||||
|
a file opened `'r+'`.
|
||||||
|
|
||||||
|
### Python bindings: remote files (2026-09-27)
|
||||||
|
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
|
||||||
|
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
|
||||||
|
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
|
||||||
|
through the block cache, the whole read API (groups, attributes, every
|
||||||
|
dataset type and index the local reader handles). A URL is any
|
||||||
|
`scheme://…`; a remote file is read-only (another mode is a
|
||||||
|
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
|
||||||
|
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
|
||||||
|
`allow_full_download`, `max_full_download`, `require_validator`,
|
||||||
|
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
|
||||||
|
cache's counters. The default wheel builds plain HTTP only (no C: rustls
|
||||||
|
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
|
||||||
|
check now covers `clawhdf5-py`.
|
||||||
|
- **Every read parses through `File::storage()`** instead of
|
||||||
|
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
|
||||||
|
group listings, variable-length data through the global heap), inside one
|
||||||
|
shared file handle that releases the GIL for all file access, not only
|
||||||
|
dataset reads: a read waiting on the network lets other Python threads
|
||||||
|
run. A failed read of the storage (a network error, a file changed on
|
||||||
|
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
|
||||||
|
data; `key in group` raises it instead of answering `False`.
|
||||||
|
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
|
||||||
|
blocks and 1 KiB blocks) against a range-capable `http.server` in the
|
||||||
|
test process, plus `tests/test_remote.py`: request counts of a small
|
||||||
|
read, cache hits, a server without `Range` support (refused, or a
|
||||||
|
whole download when allowed), a file changed on the server, a server
|
||||||
|
that hangs up, 16 threads on one remote file, and a thread that keeps
|
||||||
|
running while a read waits on 0.2 s requests.
|
||||||
|
|
||||||
### Range reads, milestone M3: remote files (2026-09-26)
|
### Range reads, milestone M3: remote files (2026-09-26)
|
||||||
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
||||||
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
|
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
|
||||||
|
|||||||
@@ -574,8 +574,36 @@ with clawhdf5.File("data.h5", "r") as f:
|
|||||||
|
|
||||||
records = f["table"] # compound -> numpy structured array
|
records = f["table"] # compound -> numpy structured array
|
||||||
ids = records["id"] # one field
|
ids = records["id"] # one field
|
||||||
|
|
||||||
|
# A file on a web server: range requests through a block cache, nothing
|
||||||
|
# downloaded up front; the same read API. The GIL is released while waiting.
|
||||||
|
with clawhdf5.File("http://data.example.org/run42.h5") as f:
|
||||||
|
first = f["group/temperatures"][0]
|
||||||
|
f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024,
|
||||||
|
headers={"Authorization": "Bearer ..."})
|
||||||
```
|
```
|
||||||
|
|
||||||
|
An existing file opened with `"r+"` is edited in place (through
|
||||||
|
`clawhdf5::FileEditor`), with h5py's indexing, broadcasting and numeric
|
||||||
|
conversion; each edit is on disk when the statement returns:
|
||||||
|
|
||||||
|
```python
|
||||||
|
with clawhdf5.File("data.h5", "r+") as f:
|
||||||
|
f["group/temperatures"][100:200, ::4] = 0.0
|
||||||
|
f["series"].resize(5000, axis=0) # chunked datasets, within maxshape
|
||||||
|
f["series"][4000:] = new_values
|
||||||
|
f["group"].attrs["calibrated"] = True
|
||||||
|
```
|
||||||
|
|
||||||
|
Creating or deleting datasets, groups and attributes in an existing file is
|
||||||
|
not supported (`NotImplementedError`); limits are in
|
||||||
|
[known issues](docs/known-issues.md).
|
||||||
|
|
||||||
|
The default build reads `http://` URLs only; build with
|
||||||
|
`maturin develop --release --features https` (rustls with ring, which
|
||||||
|
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
|
||||||
|
object-store URLs.
|
||||||
|
|
||||||
Reads cover integers and IEEE floats of every width in either byte order,
|
Reads cover integers and IEEE floats of every width in either byte order,
|
||||||
`bool`, enums, complex, fixed and variable-length strings, variable-length
|
`bool`, enums, complex, fixed and variable-length strings, variable-length
|
||||||
sequences, opaque, HDF5 array types and compounds; other types (references,
|
sequences, opaque, HDF5 array types and compounds; other types (references,
|
||||||
@@ -590,7 +618,8 @@ non-default fill value (`docs/known-issues.md`). An index list is read one
|
|||||||
group of neighbouring chunks at a time.
|
group of neighbouring chunks at a time.
|
||||||
Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`)
|
Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`)
|
||||||
covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests
|
covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests
|
||||||
in `crates/clawhdf5-py/tests` compare every read with h5py; run them with
|
in `crates/clawhdf5-py/tests` compare every read and every in-place edit
|
||||||
|
with h5py; run them with
|
||||||
`pip install pytest h5py && pytest crates/clawhdf5-py/tests`.
|
`pip install pytest h5py && pytest crates/clawhdf5-py/tests`.
|
||||||
|
|
||||||
### Agent Memory
|
### Agent Memory
|
||||||
|
|||||||
@@ -135,6 +135,12 @@ impl ChunkGrid {
|
|||||||
let mut rem = index;
|
let mut rem = index;
|
||||||
for p in 0..rank {
|
for p in 0..rank {
|
||||||
let d = self.order[p];
|
let d = self.order[p];
|
||||||
|
// A zero stride: a later dimension has no chunks (its maximum,
|
||||||
|
// or with none recorded its current extent, is 0), so no slot of
|
||||||
|
// the index is a chunk of the dataset.
|
||||||
|
if self.down[p] == 0 {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
let scaled = rem / self.down[p];
|
let scaled = rem / self.down[p];
|
||||||
rem %= self.down[p];
|
rem %= self.down[p];
|
||||||
if scaled >= self.cur_chunks[d] {
|
if scaled >= self.cur_chunks[d] {
|
||||||
@@ -193,6 +199,28 @@ mod tests {
|
|||||||
assert_eq!(g.offsets(11), Some(vec![2, 3]));
|
assert_eq!(g.offsets(11), Some(vec![2, 3]));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn zero_extent_has_no_chunks() {
|
||||||
|
// No maximum recorded and a zero current dimension: every stride
|
||||||
|
// before it is 0 (this divided by zero).
|
||||||
|
let g = ChunkGrid::fixed_array(&[1, 0], None, &[6, 6]).unwrap();
|
||||||
|
for i in 0..16 {
|
||||||
|
assert_eq!(g.offsets(i), None);
|
||||||
|
}
|
||||||
|
let g = ChunkGrid::fixed_array(&[0, 0, 3], Some(&[4, 0, 3]), &[2, 2, 3]).unwrap();
|
||||||
|
for i in 0..16 {
|
||||||
|
assert_eq!(g.offsets(i), None);
|
||||||
|
}
|
||||||
|
let g = ChunkGrid::extensible_array(&[0, 5], Some(&[u64::MAX, 0]), &[2, 2]).unwrap();
|
||||||
|
for i in 0..16 {
|
||||||
|
assert_eq!(g.offsets(i), None);
|
||||||
|
}
|
||||||
|
// A zero last dimension leaves the other strides alone.
|
||||||
|
let g = ChunkGrid::fixed_array(&[4, 0], Some(&[4, 6]), &[2, 3]).unwrap();
|
||||||
|
assert_eq!(g.offsets(0), None);
|
||||||
|
assert_eq!(g.linear_index(&[1, 1]), 3);
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn rejects_two_unlimited_dims_after_the_first() {
|
fn rejects_two_unlimited_dims_after_the_first() {
|
||||||
assert!(ChunkGrid::fixed_array(&[4, 6], Some(&[u64::MAX, u64::MAX]), &[2, 3]).is_err());
|
assert!(ChunkGrid::fixed_array(&[4, 6], Some(&[u64::MAX, u64::MAX]), &[2, 3]).is_err());
|
||||||
|
|||||||
@@ -130,6 +130,22 @@ pub(crate) fn build_chunked_dataset_oh(
|
|||||||
fill_message: &[u8],
|
fill_message: &[u8],
|
||||||
refcount: u32,
|
refcount: u32,
|
||||||
) -> Result<Vec<u8>, FormatError> {
|
) -> Result<Vec<u8>, FormatError> {
|
||||||
|
// libhdf5 records every simple dataspace's maximum dimensions (the
|
||||||
|
// dimensions themselves when none are given, `H5S_set_extent_simple`).
|
||||||
|
// Without them libhdf5 takes the maximum to be the current dimensions,
|
||||||
|
// so a resize by libhdf5 (h5py's `Dataset.resize`) would also change
|
||||||
|
// the maximum a Fixed Array chunk index is laid out by, and move every
|
||||||
|
// chunk already written.
|
||||||
|
let recorded;
|
||||||
|
let ds = if ds.space_type == DataspaceType::Simple && ds.max_dimensions.is_none() {
|
||||||
|
recorded = Dataspace {
|
||||||
|
max_dimensions: Some(ds.dimensions.clone()),
|
||||||
|
..ds.clone()
|
||||||
|
};
|
||||||
|
&recorded
|
||||||
|
} else {
|
||||||
|
ds
|
||||||
|
};
|
||||||
let mut w = ObjectHeaderWriter::new();
|
let mut w = ObjectHeaderWriter::new();
|
||||||
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
|
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
|
||||||
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
|
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
|
||||||
|
|||||||
@@ -39,6 +39,23 @@ impl MmapReader {
|
|||||||
Ok(Self { _file: file, mmap })
|
Ok(Self { _file: file, mmap })
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Memory-map a file that is already open (for reading).
|
||||||
|
///
|
||||||
|
/// The mapping references `file`'s open file description for as long as
|
||||||
|
/// it lives, so a `flock` taken through that description (or a
|
||||||
|
/// `try_clone` of it) is held until the reader is dropped.
|
||||||
|
///
|
||||||
|
/// # Safety
|
||||||
|
///
|
||||||
|
/// The same contract as [`open`](Self::open): the file must not be
|
||||||
|
/// modified while the mapping is active.
|
||||||
|
pub fn from_file(file: fs::File) -> io::Result<Self> {
|
||||||
|
// SAFETY: a read-only mapping; the caller keeps the file unmodified
|
||||||
|
// while it is alive.
|
||||||
|
let mmap = unsafe { Mmap::map(&file)? };
|
||||||
|
Ok(Self { _file: file, mmap })
|
||||||
|
}
|
||||||
|
|
||||||
/// Zero-copy access to the entire file contents.
|
/// Zero-copy access to the entire file contents.
|
||||||
pub fn as_bytes(&self) -> &[u8] {
|
pub fn as_bytes(&self) -> &[u8] {
|
||||||
&self.mmap
|
&self.mmap
|
||||||
|
|||||||
@@ -17,11 +17,20 @@ crate-type = ["cdylib", "rlib"]
|
|||||||
[dependencies]
|
[dependencies]
|
||||||
clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" }
|
clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" }
|
||||||
clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
|
clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
|
||||||
|
# Remote files (`clawhdf5.File(url)`): plain HTTP by default, which builds
|
||||||
|
# no C. HTTPS and the object stores are opt-in features below.
|
||||||
|
clawhdf5-remote = { path = "../clawhdf5-remote", version = "2.7.0" }
|
||||||
pyo3 = "0.29"
|
pyo3 = "0.29"
|
||||||
numpy = "0.29"
|
numpy = "0.29"
|
||||||
|
|
||||||
[features]
|
[features]
|
||||||
extension-module = ["pyo3/extension-module"]
|
extension-module = ["pyo3/extension-module"]
|
||||||
|
# https:// URLs (rustls with ring, which compiles C and assembly).
|
||||||
|
https = ["clawhdf5-remote/https"]
|
||||||
|
# s3://, gs://, az:// URLs (object_store; its cloud clients build aws-lc-rs, C).
|
||||||
|
s3 = ["clawhdf5-remote/s3"]
|
||||||
|
gcs = ["clawhdf5-remote/gcs"]
|
||||||
|
azure = ["clawhdf5-remote/azure"]
|
||||||
|
|
||||||
[package.metadata.docs.rs]
|
[package.metadata.docs.rs]
|
||||||
features = []
|
features = []
|
||||||
|
|||||||
@@ -43,8 +43,9 @@ with clawhdf5.File("data.h5", "r") as f:
|
|||||||
Other types raise `TypeError`.
|
Other types raise `TypeError`.
|
||||||
- Keys are h5py's: integers, slices with a positive step, `...`, one
|
- Keys are h5py's: integers, slices with a positive step, `...`, one
|
||||||
increasing list of integers, compound field names. Each maps onto a
|
increasing list of integers, compound field names. Each maps onto a
|
||||||
hyperslab selection. `None`, negative steps and boolean masks are refused
|
hyperslab selection. `None` and negative steps are refused
|
||||||
with h5py's errors.
|
with h5py's errors; boolean masks (which h5py supports) raise
|
||||||
|
`NotImplementedError`, for reads and writes.
|
||||||
- What is read from the file: a selection whose bounding box covers at
|
- What is read from the file: a selection whose bounding box covers at
|
||||||
most half the dataset decodes only the chunks (or contiguous rows) the box
|
most half the dataset decodes only the chunks (or contiguous rows) the box
|
||||||
overlaps. The library decodes the whole dataset for a larger box
|
overlaps. The library decodes the whole dataset for a larger box
|
||||||
@@ -61,6 +62,36 @@ with clawhdf5.File("data.h5", "r") as f:
|
|||||||
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
|
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
|
||||||
dataspace (h5py's `Empty`).
|
dataspace (h5py's `Empty`).
|
||||||
|
|
||||||
|
## Remote files
|
||||||
|
|
||||||
|
A URL instead of a path reads the file where it is, through
|
||||||
|
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
|
||||||
|
64 MiB budget by default), fetching only the blocks a read needs. The whole
|
||||||
|
read API works the same, and the GIL is released while waiting on the
|
||||||
|
network.
|
||||||
|
|
||||||
|
```python
|
||||||
|
f = clawhdf5.File("http://host/data.h5") # default options
|
||||||
|
f = clawhdf5.File.open_url(
|
||||||
|
"http://host/data.h5",
|
||||||
|
block_size=256 * 1024, cache_size=128 << 20, # the block cache
|
||||||
|
headers={"Authorization": "Bearer ..."}, # sent to this origin only
|
||||||
|
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
|
||||||
|
allow_full_download=False, # a server without Range support: refuse
|
||||||
|
require_validator=False, # refuse servers without ETag/Last-Modified
|
||||||
|
)
|
||||||
|
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
|
||||||
|
```
|
||||||
|
|
||||||
|
- The file is pinned when opened (ETag or Last-Modified, and length): if it
|
||||||
|
changes on the server, reads raise `OSError` instead of mixing versions.
|
||||||
|
Network failures are `OSError` too.
|
||||||
|
- Remote files are read-only.
|
||||||
|
- Schemes: the default build (no C) reads `http://`. `https://` needs
|
||||||
|
`maturin develop --release --features https` (rustls with ring, which
|
||||||
|
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
|
||||||
|
`azure` features (credentials from the environment; aws-lc-rs, C).
|
||||||
|
|
||||||
## Writing
|
## Writing
|
||||||
|
|
||||||
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
|
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
|
||||||
@@ -68,6 +99,41 @@ chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
|
|||||||
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
|
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
|
||||||
written on `close()`.
|
written on `close()`.
|
||||||
|
|
||||||
|
## Editing a file in place
|
||||||
|
|
||||||
|
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
|
||||||
|
where it is, through clawhdf5's `FileEditor`; the file is locked until
|
||||||
|
`close()`, and every edit is written and synced before the statement
|
||||||
|
returns.
|
||||||
|
|
||||||
|
```python
|
||||||
|
with clawhdf5.File("data.h5", "r+") as f:
|
||||||
|
ds = f["grid"]
|
||||||
|
ds[10:20, ::2] = 0 # h5py keys and broadcasting
|
||||||
|
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
|
||||||
|
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
|
||||||
|
f["series"].attrs["units"] = "K"
|
||||||
|
f.attrs.create("version", 2, dtype="u1")
|
||||||
|
```
|
||||||
|
|
||||||
|
- Values: a numpy array is converted to the dataset's dtype as libhdf5
|
||||||
|
converts it (integers saturate at the target's limits; floats are
|
||||||
|
truncated toward zero and clipped); anything else goes through
|
||||||
|
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
|
||||||
|
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
|
||||||
|
value). A few libhdf5 edge cases differ on purpose; see
|
||||||
|
`docs/known-issues.md`.
|
||||||
|
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
|
||||||
|
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
|
||||||
|
see its result.
|
||||||
|
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
|
||||||
|
`str` is stored as a fixed-length UTF-8 string (h5py stores a
|
||||||
|
variable-length one), so h5py reads it back as `bytes`.
|
||||||
|
- Not supported (`NotImplementedError`, nothing written): creating or
|
||||||
|
deleting datasets, groups and attributes, writing compound fields by
|
||||||
|
name, variable-length data, HDF5 array types, and whatever
|
||||||
|
`FileEditor` refuses (listed in `docs/known-issues.md`).
|
||||||
|
|
||||||
## Tests
|
## Tests
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
@@ -76,7 +142,12 @@ pytest crates/clawhdf5-py/tests
|
|||||||
```
|
```
|
||||||
|
|
||||||
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
|
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
|
||||||
writes. `scripts/ci-test.sh` builds the wheel and runs these in CI.
|
writes, opened locally and over HTTP (an in-process range server,
|
||||||
|
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
|
||||||
|
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
|
||||||
|
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
|
||||||
|
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
|
||||||
|
the wheel and runs these in CI.
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
+166
-65
@@ -1,22 +1,62 @@
|
|||||||
//! PyAttrs — dict-like access to HDF5 attributes.
|
//! PyAttrs — dict-like access to HDF5 attributes.
|
||||||
|
|
||||||
use std::sync::{Arc, Mutex};
|
use std::sync::{Arc, Mutex, PoisonError};
|
||||||
|
|
||||||
use clawhdf5_format::attribute::AttributeMessage;
|
use clawhdf5_format::attribute::AttributeMessage;
|
||||||
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
|
use pyo3::exceptions::{PyKeyError, PyNotImplementedError, PyTypeError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::{PyList, PyTuple};
|
use pyo3::types::{PyList, PyTuple};
|
||||||
|
|
||||||
use crate::convert::{Converter, Elements, resolve_vl};
|
use crate::convert::{Converter, Elements, resolve_vl};
|
||||||
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value};
|
use crate::handle::Handle;
|
||||||
|
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, edit, node, py_to_attr_value};
|
||||||
|
|
||||||
|
/// The attributes of an object in a file opened for reading (or editing).
|
||||||
|
struct ReadAttrs {
|
||||||
|
handle: Arc<Handle>,
|
||||||
|
addr: u64,
|
||||||
|
path: String,
|
||||||
|
/// Sorted by name, with the file generation they were read at: an edit
|
||||||
|
/// (`attrs[name] = value`, here or through another handle on the same
|
||||||
|
/// object) makes them re-read.
|
||||||
|
cache: Mutex<(u64, Arc<Vec<AttributeMessage>>)>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl ReadAttrs {
|
||||||
|
fn current(&self, py: Python<'_>) -> PyResult<Arc<Vec<AttributeMessage>>> {
|
||||||
|
let generation = self.handle.generation();
|
||||||
|
{
|
||||||
|
let cached = self.cache.lock().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
if cached.0 == generation {
|
||||||
|
return Ok(Arc::clone(&cached.1));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let (addr, path) = (self.addr, &self.path);
|
||||||
|
let attrs = Arc::new(self.handle.with(py, |f| node::attributes(f, addr, path))?);
|
||||||
|
*self.cache.lock().unwrap_or_else(PoisonError::into_inner) =
|
||||||
|
(generation, Arc::clone(&attrs));
|
||||||
|
Ok(attrs)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn check_writable(&self) -> PyResult<()> {
|
||||||
|
if self.handle.is_writable() {
|
||||||
|
return Ok(());
|
||||||
|
}
|
||||||
|
Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||||
|
"cannot set attributes on a read-only file (open it with mode 'r+')",
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn set(&self, py: Python<'_>, name: &str, value: clawhdf5_rs::AttrValue) -> PyResult<()> {
|
||||||
|
self.check_writable()?;
|
||||||
|
let path = node::name(&self.path);
|
||||||
|
self.handle.edit(py, |ed| ed.set_attr(&path, name, &value))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Backing storage for attributes.
|
/// Backing storage for attributes.
|
||||||
enum AttrsInner {
|
enum AttrsInner {
|
||||||
/// Attributes of an object in a file opened for reading, sorted by name.
|
Read(ReadAttrs),
|
||||||
Read {
|
|
||||||
file: Arc<clawhdf5_rs::File>,
|
|
||||||
attrs: Vec<AttributeMessage>,
|
|
||||||
},
|
|
||||||
/// Writable attribute list shared with a parent (PyFile or PyGroup).
|
/// Writable attribute list shared with a parent (PyFile or PyGroup).
|
||||||
Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>),
|
Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>),
|
||||||
}
|
}
|
||||||
@@ -26,8 +66,11 @@ enum AttrsInner {
|
|||||||
/// In read mode, values are what h5py returns: numpy scalars for scalar
|
/// In read mode, values are what h5py returns: numpy scalars for scalar
|
||||||
/// attributes, numpy arrays otherwise, `str` for variable-length strings,
|
/// attributes, numpy arrays otherwise, `str` for variable-length strings,
|
||||||
/// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace.
|
/// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace.
|
||||||
/// In write mode, attributes set here are accumulated and written when
|
/// In a file opened with `'r+'`, `attrs[name] = value` adds or replaces an
|
||||||
/// the parent file is closed.
|
/// attribute in the file at once (as h5py stores it, except that `str`
|
||||||
|
/// values become fixed-length UTF-8 strings). In write mode (`'w'`),
|
||||||
|
/// attributes set here are accumulated and written when the parent file is
|
||||||
|
/// closed.
|
||||||
#[pyclass(name = "Attrs")]
|
#[pyclass(name = "Attrs")]
|
||||||
pub struct PyAttrs {
|
pub struct PyAttrs {
|
||||||
inner: AttrsInner,
|
inner: AttrsInner,
|
||||||
@@ -36,10 +79,21 @@ pub struct PyAttrs {
|
|||||||
impl PyAttrs {
|
impl PyAttrs {
|
||||||
/// The attributes of the object at `addr` (whose path is `path`) in a
|
/// The attributes of the object at `addr` (whose path is `path`) in a
|
||||||
/// file opened for reading.
|
/// file opened for reading.
|
||||||
pub(crate) fn read(file: Arc<clawhdf5_rs::File>, addr: u64, path: &str) -> PyResult<Self> {
|
pub(crate) fn read(
|
||||||
let attrs = node::attributes(&file, addr, path)?;
|
py: Python<'_>,
|
||||||
|
handle: Arc<Handle>,
|
||||||
|
addr: u64,
|
||||||
|
path: &str,
|
||||||
|
) -> PyResult<Self> {
|
||||||
|
let generation = handle.generation();
|
||||||
|
let attrs = Arc::new(handle.with(py, |f| node::attributes(f, addr, path))?);
|
||||||
Ok(Self {
|
Ok(Self {
|
||||||
inner: AttrsInner::Read { file, attrs },
|
inner: AttrsInner::Read(ReadAttrs {
|
||||||
|
handle,
|
||||||
|
addr,
|
||||||
|
path: path.to_string(),
|
||||||
|
cache: Mutex::new((generation, attrs)),
|
||||||
|
}),
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -49,14 +103,48 @@ impl PyAttrs {
|
|||||||
inner: AttrsInner::Write(store),
|
inner: AttrsInner::Write(store),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn set_value(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
key: &str,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
dtype: Option<&Bound<'_, PyAny>>,
|
||||||
|
shape: Option<&Bound<'_, PyAny>>,
|
||||||
|
) -> PyResult<()> {
|
||||||
|
match &self.inner {
|
||||||
|
AttrsInner::Read(r) => {
|
||||||
|
r.check_writable()?;
|
||||||
|
let value = edit::attr_value(py, value, dtype, shape)?;
|
||||||
|
r.set(py, key, value)
|
||||||
|
}
|
||||||
|
AttrsInner::Write(store) => {
|
||||||
|
if dtype.is_some() || shape.is_some() {
|
||||||
|
return Err(PyNotImplementedError::new_err(
|
||||||
|
"attrs.create with a dtype or shape is only supported in a file opened \
|
||||||
|
with 'r+'",
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let owned = py_to_attr_value(value)?;
|
||||||
|
let mut guard = store.lock().unwrap();
|
||||||
|
// Replace existing key if present.
|
||||||
|
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
|
||||||
|
entry.1 = owned;
|
||||||
|
} else {
|
||||||
|
guard.push((key.to_string(), owned));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
#[pymethods]
|
#[pymethods]
|
||||||
impl PyAttrs {
|
impl PyAttrs {
|
||||||
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { file, attrs } => match attrs.iter().find(|a| a.name == key) {
|
AttrsInner::Read(r) => match r.current(py)?.iter().find(|a| a.name == key) {
|
||||||
Some(attr) => Ok(attr_to_py(py, file, attr)?.unbind()),
|
Some(attr) => Ok(attr_to_py(py, &r.handle, attr)?.unbind()),
|
||||||
None => Err(PyKeyError::new_err(format!(
|
None => Err(PyKeyError::new_err(format!(
|
||||||
"Can't open attribute (can't locate attribute: '{key}')"
|
"Can't open attribute (can't locate attribute: '{key}')"
|
||||||
))),
|
))),
|
||||||
@@ -74,36 +162,63 @@ impl PyAttrs {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __setitem__(&self, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
/// `attrs[name] = value`. In a file opened with `'r+'` this writes the
|
||||||
|
/// attribute (numeric, bool, complex, bytes and str data, any shape)
|
||||||
|
/// into the file before returning; see the class docs.
|
||||||
|
fn __setitem__(&self, py: Python<'_>, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
||||||
|
self.set_value(py, key, value, None, None)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Deleting attributes is not supported in a file (the in-place editor
|
||||||
|
/// cannot remove them); in write mode it removes a pending attribute.
|
||||||
|
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { .. } => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
AttrsInner::Read(_) => Err(PyNotImplementedError::new_err(format!(
|
||||||
"cannot set attributes on a read-only file",
|
"cannot delete attribute '{key}': deleting attributes is not supported by \
|
||||||
)),
|
clawhdf5's in-place editor"
|
||||||
|
))),
|
||||||
AttrsInner::Write(store) => {
|
AttrsInner::Write(store) => {
|
||||||
let owned = py_to_attr_value(value)?;
|
|
||||||
let mut guard = store.lock().unwrap();
|
let mut guard = store.lock().unwrap();
|
||||||
// Replace existing key if present.
|
let before = guard.len();
|
||||||
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
|
guard.retain(|(k, _)| k != key);
|
||||||
entry.1 = owned;
|
if guard.len() == before {
|
||||||
} else {
|
return Err(PyKeyError::new_err(key.to_string()));
|
||||||
guard.push((key.to_string(), owned));
|
|
||||||
}
|
}
|
||||||
Ok(())
|
Ok(())
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __len__(&self) -> usize {
|
/// h5py's `attrs.create(name, data, shape=None, dtype=None)`: `data`
|
||||||
|
/// converted to `dtype` and reshaped to `shape` first.
|
||||||
|
#[pyo3(signature = (name, data, shape=None, dtype=None))]
|
||||||
|
fn create(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
name: &str,
|
||||||
|
data: &Bound<'_, PyAny>,
|
||||||
|
shape: Option<&Bound<'_, PyAny>>,
|
||||||
|
dtype: Option<&Bound<'_, PyAny>>,
|
||||||
|
) -> PyResult<()> {
|
||||||
|
self.set_value(py, name, data, dtype, shape)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// h5py's `attrs.modify(name, value)`: same as `attrs[name] = value`.
|
||||||
|
fn modify(&self, py: Python<'_>, name: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
||||||
|
self.set_value(py, name, value, None, None)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { attrs, .. } => attrs.len(),
|
AttrsInner::Read(r) => Ok(r.current(py)?.len()),
|
||||||
AttrsInner::Write(store) => store.lock().unwrap().len(),
|
AttrsInner::Write(store) => Ok(store.lock().unwrap().len()),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __contains__(&self, key: &str) -> bool {
|
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
AttrsInner::Read { attrs, .. } => attrs.iter().any(|a| a.name == key),
|
AttrsInner::Read(r) => Ok(r.current(py)?.iter().any(|a| a.name == key)),
|
||||||
AttrsInner::Write(store) => store.lock().unwrap().iter().any(|(k, _)| k == key),
|
AttrsInner::Write(store) => Ok(store.lock().unwrap().iter().any(|(k, _)| k == key)),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -113,15 +228,17 @@ impl PyAttrs {
|
|||||||
Ok(iter)
|
Ok(iter)
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __repr__(&self) -> String {
|
fn __repr__(&self, py: Python<'_>) -> String {
|
||||||
let n = self.__len__();
|
match self.__len__(py) {
|
||||||
format!("<HDF5 Attrs ({n} members)>")
|
Ok(n) => format!("<HDF5 Attrs ({n} members)>"),
|
||||||
|
Err(_) => "<HDF5 Attrs>".to_string(),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The value of `key`, or `default` if there is no such attribute.
|
/// The value of `key`, or `default` if there is no such attribute.
|
||||||
#[pyo3(signature = (key, default=None))]
|
#[pyo3(signature = (key, default=None))]
|
||||||
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
|
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
|
||||||
if self.__contains__(key) {
|
if self.__contains__(py, key)? {
|
||||||
self.__getitem__(py, key)
|
self.__getitem__(py, key)
|
||||||
} else {
|
} else {
|
||||||
Ok(default.unwrap_or_else(|| py.None()))
|
Ok(default.unwrap_or_else(|| py.None()))
|
||||||
@@ -131,7 +248,7 @@ impl PyAttrs {
|
|||||||
/// Return attribute names as a list.
|
/// Return attribute names as a list.
|
||||||
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
let names: Vec<String> = match &self.inner {
|
let names: Vec<String> = match &self.inner {
|
||||||
AttrsInner::Read { attrs, .. } => attrs.iter().map(|a| a.name.clone()).collect(),
|
AttrsInner::Read(r) => r.current(py)?.iter().map(|a| a.name.clone()).collect(),
|
||||||
AttrsInner::Write(store) => store
|
AttrsInner::Write(store) => store
|
||||||
.lock()
|
.lock()
|
||||||
.unwrap()
|
.unwrap()
|
||||||
@@ -146,9 +263,10 @@ impl PyAttrs {
|
|||||||
/// Return attribute values as a list.
|
/// Return attribute values as a list.
|
||||||
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
let vals: Vec<Py<PyAny>> = match &self.inner {
|
let vals: Vec<Py<PyAny>> = match &self.inner {
|
||||||
AttrsInner::Read { file, attrs } => attrs
|
AttrsInner::Read(r) => r
|
||||||
|
.current(py)?
|
||||||
.iter()
|
.iter()
|
||||||
.map(|a| attr_to_py(py, file, a).map(Bound::unbind))
|
.map(|a| attr_to_py(py, &r.handle, a).map(Bound::unbind))
|
||||||
.collect::<PyResult<_>>()?,
|
.collect::<PyResult<_>>()?,
|
||||||
AttrsInner::Write(store) => store
|
AttrsInner::Write(store) => store
|
||||||
.lock()
|
.lock()
|
||||||
@@ -167,9 +285,10 @@ impl PyAttrs {
|
|||||||
/// Return attribute (key, value) pairs as a list of tuples.
|
/// Return attribute (key, value) pairs as a list of tuples.
|
||||||
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
|
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
|
||||||
AttrsInner::Read { file, attrs } => attrs
|
AttrsInner::Read(r) => r
|
||||||
|
.current(py)?
|
||||||
.iter()
|
.iter()
|
||||||
.map(|a| Ok((a.name.clone(), attr_to_py(py, file, a)?.unbind())))
|
.map(|a| Ok((a.name.clone(), attr_to_py(py, &r.handle, a)?.unbind())))
|
||||||
.collect::<PyResult<_>>()?,
|
.collect::<PyResult<_>>()?,
|
||||||
AttrsInner::Write(store) => store
|
AttrsInner::Write(store) => store
|
||||||
.lock()
|
.lock()
|
||||||
@@ -189,12 +308,11 @@ impl PyAttrs {
|
|||||||
/// An attribute's value as h5py returns it.
|
/// An attribute's value as h5py returns it.
|
||||||
fn attr_to_py<'py>(
|
fn attr_to_py<'py>(
|
||||||
py: Python<'py>,
|
py: Python<'py>,
|
||||||
file: &clawhdf5_rs::File,
|
handle: &Handle,
|
||||||
attr: &AttributeMessage,
|
attr: &AttributeMessage,
|
||||||
) -> PyResult<Bound<'py, PyAny>> {
|
) -> PyResult<Bound<'py, PyAny>> {
|
||||||
crate::no_panic(|| {
|
crate::no_panic(|| {
|
||||||
let sb = file.superblock();
|
let conv = Converter::new(py, &attr.datatype, handle.offset_size)
|
||||||
let conv = Converter::new(py, &attr.datatype, sb.offset_size)
|
|
||||||
.map_err(|e| prefix_err(py, &attr.name, e))?;
|
.map_err(|e| prefix_err(py, &attr.name, e))?;
|
||||||
if node::is_null(&attr.dataspace) {
|
if node::is_null(&attr.dataspace) {
|
||||||
return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any());
|
return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any());
|
||||||
@@ -216,12 +334,11 @@ fn attr_to_py<'py>(
|
|||||||
)));
|
)));
|
||||||
}
|
}
|
||||||
let raw = &attr.raw_data[..want];
|
let raw = &attr.raw_data[..want];
|
||||||
let file_data = file.as_bytes();
|
let (osz, lsz, unit) = (handle.offset_size, handle.length_size, conv.vl_unit);
|
||||||
let (osz, lsz, unit) = (sb.offset_size, sb.length_size, conv.vl_unit);
|
let what = format!("attribute {}", attr.name);
|
||||||
Elements::Vl(
|
Elements::Vl(handle.with(py, |f| {
|
||||||
py.detach(|| resolve_vl(file_data, raw, n, osz, lsz, unit))
|
resolve_vl(f.storage(), raw, n, osz, lsz, unit).map_err(|e| e.into_py(&what))
|
||||||
.map_err(|e| PyValueError::new_err(format!("attribute {}: {e}", attr.name)))?,
|
})?)
|
||||||
)
|
|
||||||
} else {
|
} else {
|
||||||
Elements::Bytes(attr.raw_data.clone())
|
Elements::Bytes(attr.raw_data.clone())
|
||||||
};
|
};
|
||||||
@@ -244,19 +361,3 @@ fn prefix_err(py: Python<'_>, name: &str, e: PyErr) -> PyErr {
|
|||||||
PyValueError::new_err(msg)
|
PyValueError::new_err(msg)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
#[cfg(test)]
|
|
||||||
mod tests {
|
|
||||||
use super::*;
|
|
||||||
|
|
||||||
#[test]
|
|
||||||
fn write_attrs_len() {
|
|
||||||
let store = Arc::new(Mutex::new(Vec::new()));
|
|
||||||
store
|
|
||||||
.lock()
|
|
||||||
.unwrap()
|
|
||||||
.push(("key".into(), OwnedAttrValue::I64(99)));
|
|
||||||
let attrs = PyAttrs::from_write(store);
|
|
||||||
assert_eq!(attrs.__len__(), 1);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|||||||
@@ -17,6 +17,7 @@ use std::collections::HashMap;
|
|||||||
|
|
||||||
use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder};
|
use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder};
|
||||||
use clawhdf5_format::global_heap::GlobalHeapCollection;
|
use clawhdf5_format::global_heap::GlobalHeapCollection;
|
||||||
|
use clawhdf5_format::storage::Storage;
|
||||||
use numpy::PyArray1;
|
use numpy::PyArray1;
|
||||||
use pyo3::exceptions::{PyTypeError, PyValueError};
|
use pyo3::exceptions::{PyTypeError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
@@ -538,16 +539,16 @@ fn object_array<'py>(
|
|||||||
/// their bytes: each element's stored length times `unit` (1 for strings,
|
/// their bytes: each element's stored length times `unit` (1 for strings,
|
||||||
/// the base type's size for sequences). Pure Rust, so it runs without the
|
/// the base type's size for sequences). Pure Rust, so it runs without the
|
||||||
/// GIL.
|
/// GIL.
|
||||||
pub(crate) fn resolve_vl(
|
pub(crate) fn resolve_vl<S: Storage + ?Sized>(
|
||||||
file_data: &[u8],
|
file: &S,
|
||||||
raw: &[u8],
|
raw: &[u8],
|
||||||
count: usize,
|
count: usize,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
length_size: u8,
|
length_size: u8,
|
||||||
unit: usize,
|
unit: usize,
|
||||||
) -> Result<Vec<Vec<u8>>, String> {
|
) -> Result<Vec<Vec<u8>>, VlError> {
|
||||||
let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size)
|
let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size)
|
||||||
.map_err(|e| e.to_string())?;
|
.map_err(|e| VlError::Invalid(e.to_string()))?;
|
||||||
let undefined = match offset_size {
|
let undefined = match offset_size {
|
||||||
2 => 0xFFFF,
|
2 => 0xFFFF,
|
||||||
4 => 0xFFFF_FFFF,
|
4 => 0xFFFF_FFFF,
|
||||||
@@ -558,43 +559,70 @@ pub(crate) fn resolve_vl(
|
|||||||
for vl in &refs {
|
for vl in &refs {
|
||||||
if vl.collection_address == 0 || vl.collection_address == undefined {
|
if vl.collection_address == 0 || vl.collection_address == undefined {
|
||||||
if vl.length != 0 {
|
if vl.length != 0 {
|
||||||
return Err(format!(
|
return Err(VlError::Invalid(format!(
|
||||||
"variable-length element of length {} has no heap address",
|
"variable-length element of length {} has no heap address",
|
||||||
vl.length
|
vl.length
|
||||||
));
|
)));
|
||||||
}
|
}
|
||||||
out.push(Vec::new());
|
out.push(Vec::new());
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
let coll = match collections.entry(vl.collection_address) {
|
let coll = match collections.entry(vl.collection_address) {
|
||||||
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
|
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
|
||||||
std::collections::hash_map::Entry::Vacant(e) => {
|
std::collections::hash_map::Entry::Vacant(e) => e.insert(
|
||||||
let addr = usize::try_from(vl.collection_address)
|
GlobalHeapCollection::parse_in(file, vl.collection_address, length_size)
|
||||||
.map_err(|_| "global heap address out of range".to_string())?;
|
.map_err(VlError::from_format)?,
|
||||||
e.insert(
|
),
|
||||||
GlobalHeapCollection::parse(file_data, addr, length_size)
|
|
||||||
.map_err(|e| e.to_string())?,
|
|
||||||
)
|
|
||||||
}
|
|
||||||
};
|
};
|
||||||
let index = u16::try_from(vl.object_index)
|
let index = u16::try_from(vl.object_index).map_err(|_| {
|
||||||
.map_err(|_| format!("global heap object index {} out of range", vl.object_index))?;
|
VlError::Invalid(format!(
|
||||||
|
"global heap object index {} out of range",
|
||||||
|
vl.object_index
|
||||||
|
))
|
||||||
|
})?;
|
||||||
let obj = coll.get_object(index).ok_or_else(|| {
|
let obj = coll.get_object(index).ok_or_else(|| {
|
||||||
format!(
|
VlError::Invalid(format!(
|
||||||
"global heap object {index} not found in the collection at {}",
|
"global heap object {index} not found in the collection at {}",
|
||||||
vl.collection_address
|
vl.collection_address
|
||||||
)
|
))
|
||||||
})?;
|
})?;
|
||||||
let need = (vl.length as usize)
|
let need = (vl.length as usize)
|
||||||
.checked_mul(unit)
|
.checked_mul(unit)
|
||||||
.ok_or("variable-length element too long")?;
|
.ok_or_else(|| VlError::Invalid("variable-length element too long".into()))?;
|
||||||
if need > obj.data.len() {
|
if need > obj.data.len() {
|
||||||
return Err(format!(
|
return Err(VlError::Invalid(format!(
|
||||||
"variable-length element of {need} bytes in a {}-byte heap object",
|
"variable-length element of {need} bytes in a {}-byte heap object",
|
||||||
obj.data.len()
|
obj.data.len()
|
||||||
));
|
)));
|
||||||
}
|
}
|
||||||
out.push(obj.data[..need].to_vec());
|
out.push(obj.data[..need].to_vec());
|
||||||
}
|
}
|
||||||
Ok(out)
|
Ok(out)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Why variable-length elements could not be resolved.
|
||||||
|
#[derive(Debug)]
|
||||||
|
pub(crate) enum VlError {
|
||||||
|
/// Reading the file failed (a network error on a remote file).
|
||||||
|
Storage(String),
|
||||||
|
/// The references or the heap are not valid.
|
||||||
|
Invalid(String),
|
||||||
|
}
|
||||||
|
|
||||||
|
impl VlError {
|
||||||
|
fn from_format(e: clawhdf5_format::error::FormatError) -> Self {
|
||||||
|
match e {
|
||||||
|
clawhdf5_format::error::FormatError::Storage(_) => VlError::Storage(e.to_string()),
|
||||||
|
e => VlError::Invalid(e.to_string()),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// As a Python exception, the message prefixed with `what`: a storage
|
||||||
|
/// failure is an `OSError`, anything else a `ValueError`.
|
||||||
|
pub(crate) fn into_py(self, what: &str) -> PyErr {
|
||||||
|
match self {
|
||||||
|
VlError::Storage(m) => pyo3::exceptions::PyOSError::new_err(format!("{what}: {m}")),
|
||||||
|
VlError::Invalid(m) => PyValueError::new_err(format!("{what}: {m}")),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@@ -6,20 +6,52 @@
|
|||||||
//! whole dataset instead); the
|
//! whole dataset instead); the
|
||||||
//! bytes it returns become the numpy array's buffer without a copy (see
|
//! bytes it returns become the numpy array's buffer without a copy (see
|
||||||
//! `convert`). All file access and decoding runs with the GIL released, so
|
//! `convert`). All file access and decoding runs with the GIL released, so
|
||||||
//! Python threads reading the same or different datasets run in parallel.
|
//! Python threads reading the same or different datasets run in parallel,
|
||||||
|
//! and a remote file's network reads never hold the GIL.
|
||||||
|
|
||||||
use std::sync::Arc;
|
use std::sync::{Arc, Mutex, PoisonError};
|
||||||
|
|
||||||
use clawhdf5_format::datatype::Datatype;
|
use clawhdf5_format::datatype::Datatype;
|
||||||
use clawhdf5_format::object_header::ObjectHeader;
|
use clawhdf5_format::object_header::ObjectHeader;
|
||||||
use pyo3::exceptions::{PyTypeError, PyValueError};
|
use clawhdf5_rs::File;
|
||||||
|
use pyo3::exceptions::{PyNotImplementedError, PyOSError, PyTypeError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::{PyList, PyTuple};
|
use pyo3::types::{PyList, PyTuple};
|
||||||
|
|
||||||
use crate::attrs::PyAttrs;
|
use crate::attrs::PyAttrs;
|
||||||
use crate::convert::{Converter, Elements, resolve_vl};
|
use crate::convert::{Converter, Elements, VlError, resolve_vl};
|
||||||
|
use crate::handle::Handle;
|
||||||
use crate::select::{self, Plan};
|
use crate::select::{self, Plan};
|
||||||
use crate::{PyEmpty, node, to_py_err};
|
use crate::{PyEmpty, edit, node, to_py_err};
|
||||||
|
|
||||||
|
/// What opening a dataset reads from the file (without the GIL).
|
||||||
|
pub(crate) struct DatasetMeta {
|
||||||
|
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
|
||||||
|
shape: Option<Vec<u64>>,
|
||||||
|
chunks: Option<Vec<u64>>,
|
||||||
|
datatype: Datatype,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl DatasetMeta {
|
||||||
|
pub(crate) fn load(f: &File, addr: u64, hdr: &ObjectHeader, path: &str) -> PyResult<Self> {
|
||||||
|
let null = node::is_null(&node::dataspace(f, hdr, path)?);
|
||||||
|
let ds = f.dataset_at(addr).map_err(to_py_err)?;
|
||||||
|
let shape = if null {
|
||||||
|
None
|
||||||
|
} else {
|
||||||
|
Some(ds.shape().map_err(to_py_err)?)
|
||||||
|
};
|
||||||
|
let datatype = ds.raw_datatype().map_err(to_py_err)?;
|
||||||
|
let chunks = shape
|
||||||
|
.as_ref()
|
||||||
|
.and_then(|s| node::chunk_shape(f, hdr, s.len()));
|
||||||
|
Ok(Self {
|
||||||
|
shape,
|
||||||
|
chunks,
|
||||||
|
datatype,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// A dataset in a file opened for reading.
|
/// A dataset in a file opened for reading.
|
||||||
///
|
///
|
||||||
@@ -30,13 +62,14 @@ use crate::{PyEmpty, node, to_py_err};
|
|||||||
/// ```
|
/// ```
|
||||||
#[pyclass(name = "Dataset")]
|
#[pyclass(name = "Dataset")]
|
||||||
pub struct PyDataset {
|
pub struct PyDataset {
|
||||||
file: Arc<clawhdf5_rs::File>,
|
handle: Arc<Handle>,
|
||||||
path: String,
|
path: String,
|
||||||
/// Where the dataset's object header is: reads open it from here rather
|
/// Where the dataset's object header is: reads open it from here rather
|
||||||
/// than resolve `path` again.
|
/// than resolve `path` again.
|
||||||
addr: u64,
|
addr: u64,
|
||||||
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
|
/// The shape (`None` for a null dataspace, h5py's `Empty`), with the
|
||||||
shape: Option<Vec<u64>>,
|
/// file generation it was read at: an edit (a resize) may change it.
|
||||||
|
shape: Mutex<(u64, Option<Vec<u64>>)>,
|
||||||
/// The chunk shape, for a chunked dataset.
|
/// The chunk shape, for a chunked dataset.
|
||||||
chunks: Option<Vec<u64>>,
|
chunks: Option<Vec<u64>>,
|
||||||
datatype: Datatype,
|
datatype: Datatype,
|
||||||
@@ -45,39 +78,25 @@ pub struct PyDataset {
|
|||||||
}
|
}
|
||||||
|
|
||||||
impl PyDataset {
|
impl PyDataset {
|
||||||
pub(crate) fn open(
|
pub(crate) fn new(
|
||||||
py: Python<'_>,
|
py: Python<'_>,
|
||||||
file: Arc<clawhdf5_rs::File>,
|
handle: Arc<Handle>,
|
||||||
path: String,
|
path: String,
|
||||||
addr: u64,
|
addr: u64,
|
||||||
hdr: &ObjectHeader,
|
meta: DatasetMeta,
|
||||||
) -> PyResult<Self> {
|
) -> Self {
|
||||||
crate::no_panic(|| {
|
let generation = handle.generation();
|
||||||
let null = node::is_null(&node::dataspace(&file, hdr)?);
|
let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
|
||||||
let (shape, datatype) = {
|
.map_err(|e| e.value(py).to_string());
|
||||||
let ds = file.dataset_at(addr).map_err(to_py_err)?;
|
Self {
|
||||||
let shape = if null {
|
handle,
|
||||||
None
|
path,
|
||||||
} else {
|
addr,
|
||||||
Some(ds.shape().map_err(to_py_err)?)
|
shape: Mutex::new((generation, meta.shape)),
|
||||||
};
|
chunks: meta.chunks,
|
||||||
(shape, ds.raw_datatype().map_err(to_py_err)?)
|
datatype: meta.datatype,
|
||||||
};
|
conv,
|
||||||
let conv = Converter::new(py, &datatype, file.superblock().offset_size)
|
}
|
||||||
.map_err(|e| e.value(py).to_string());
|
|
||||||
let chunks = shape
|
|
||||||
.as_ref()
|
|
||||||
.and_then(|s| node::chunk_shape(&file, hdr, s.len()));
|
|
||||||
Ok(Self {
|
|
||||||
file,
|
|
||||||
path,
|
|
||||||
addr,
|
|
||||||
shape,
|
|
||||||
chunks,
|
|
||||||
datatype,
|
|
||||||
conv,
|
|
||||||
})
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
fn converter(&self) -> PyResult<&Converter> {
|
fn converter(&self) -> PyResult<&Converter> {
|
||||||
@@ -86,10 +105,54 @@ impl PyDataset {
|
|||||||
.map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path))))
|
.map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path))))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The current shape: the one read at open, or re-read after an edit.
|
||||||
|
fn dims(&self, py: Python<'_>) -> PyResult<Option<Vec<u64>>> {
|
||||||
|
let generation = self.handle.generation();
|
||||||
|
{
|
||||||
|
let cached = self.shape.lock().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
if cached.0 == generation {
|
||||||
|
return Ok(cached.1.clone());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let addr = self.addr;
|
||||||
|
let null = self
|
||||||
|
.shape
|
||||||
|
.lock()
|
||||||
|
.unwrap_or_else(PoisonError::into_inner)
|
||||||
|
.1
|
||||||
|
.is_none();
|
||||||
|
let shape = if null {
|
||||||
|
None
|
||||||
|
} else {
|
||||||
|
Some(self.handle.with(py, |f| {
|
||||||
|
f.dataset_at(addr)
|
||||||
|
.and_then(|ds| ds.shape())
|
||||||
|
.map_err(to_py_err)
|
||||||
|
})?)
|
||||||
|
};
|
||||||
|
*self.shape.lock().unwrap_or_else(PoisonError::into_inner) = (generation, shape.clone());
|
||||||
|
Ok(shape)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn check_writable(&self) -> PyResult<()> {
|
||||||
|
if self.handle.is_writable() {
|
||||||
|
Ok(())
|
||||||
|
} else {
|
||||||
|
Err(PyOSError::new_err(format!(
|
||||||
|
"{}: the file is open read-only; open it with mode 'r+' to change it",
|
||||||
|
node::name(&self.path)
|
||||||
|
)))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Read the selection described by `plan` into a numpy array.
|
/// Read the selection described by `plan` into a numpy array.
|
||||||
fn read_plan<'py>(&self, py: Python<'py>, plan: &Plan) -> PyResult<Bound<'py, PyAny>> {
|
fn read_plan<'py>(
|
||||||
|
&self,
|
||||||
|
py: Python<'py>,
|
||||||
|
plan: &Plan,
|
||||||
|
dims: &[u64],
|
||||||
|
) -> PyResult<Bound<'py, PyAny>> {
|
||||||
let conv = self.converter()?;
|
let conv = self.converter()?;
|
||||||
let dims = self.shape.as_deref().unwrap_or(&[]);
|
|
||||||
let out_shape = plan.out_shape();
|
let out_shape = plan.out_shape();
|
||||||
|
|
||||||
let arr = if plan.is_empty() {
|
let arr = if plan.is_empty() {
|
||||||
@@ -103,10 +166,10 @@ impl PyDataset {
|
|||||||
};
|
};
|
||||||
let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size);
|
let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size);
|
||||||
let read_shape = plan.read_shape();
|
let read_shape = plan.read_shape();
|
||||||
let file = &*self.file;
|
let handle = &*self.handle;
|
||||||
let addr = self.addr;
|
let addr = self.addr;
|
||||||
// Everything below touches only Rust data: release the GIL.
|
// Everything below touches only Rust data: release the GIL.
|
||||||
let read = || -> Result<Elements, ReadError> {
|
let read = |file: &File| -> Result<Elements, ReadError> {
|
||||||
let ds = file.dataset_at(addr)?;
|
let ds = file.dataset_at(addr)?;
|
||||||
let mut blocks = Vec::with_capacity(reads.len());
|
let mut blocks = Vec::with_capacity(reads.len());
|
||||||
for read in reads {
|
for read in reads {
|
||||||
@@ -147,7 +210,7 @@ impl PyDataset {
|
|||||||
let sb = file.superblock();
|
let sb = file.superblock();
|
||||||
let n = read_shape.iter().product();
|
let n = read_shape.iter().product();
|
||||||
resolve_vl(
|
resolve_vl(
|
||||||
file.as_bytes(),
|
file.storage(),
|
||||||
&raw,
|
&raw,
|
||||||
n,
|
n,
|
||||||
sb.offset_size,
|
sb.offset_size,
|
||||||
@@ -155,12 +218,20 @@ impl PyDataset {
|
|||||||
unit,
|
unit,
|
||||||
)
|
)
|
||||||
.map(Elements::Vl)
|
.map(Elements::Vl)
|
||||||
.map_err(ReadError::Other)
|
.map_err(ReadError::Vl)
|
||||||
};
|
};
|
||||||
let data = py
|
let data = py
|
||||||
.detach(|| {
|
.detach(|| {
|
||||||
std::panic::catch_unwind(std::panic::AssertUnwindSafe(read))
|
handle
|
||||||
.unwrap_or_else(|p| Err(ReadError::Panic(crate::panic_text(&*p))))
|
.with_detached(|f| {
|
||||||
|
Ok(
|
||||||
|
std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| read(f)))
|
||||||
|
.unwrap_or_else(|p| {
|
||||||
|
Err(ReadError::Panic(crate::panic_text(&*p)))
|
||||||
|
}),
|
||||||
|
)
|
||||||
|
})
|
||||||
|
.unwrap_or_else(|e| Err(ReadError::Py(e)))
|
||||||
})
|
})
|
||||||
.map_err(|e| e.into_py(&self.path))?;
|
.map_err(|e| e.into_py(&self.path))?;
|
||||||
let joined = conv.to_array(py, data, &read_shape, false)?;
|
let joined = conv.to_array(py, data, &read_shape, false)?;
|
||||||
@@ -183,6 +254,8 @@ impl PyDataset {
|
|||||||
/// An error from the read closure, turned into a Python error with the GIL.
|
/// An error from the read closure, turned into a Python error with the GIL.
|
||||||
enum ReadError {
|
enum ReadError {
|
||||||
Lib(clawhdf5_rs::Error),
|
Lib(clawhdf5_rs::Error),
|
||||||
|
Vl(VlError),
|
||||||
|
Py(PyErr),
|
||||||
Other(String),
|
Other(String),
|
||||||
Panic(String),
|
Panic(String),
|
||||||
}
|
}
|
||||||
@@ -197,6 +270,8 @@ impl ReadError {
|
|||||||
fn into_py(self, path: &str) -> PyErr {
|
fn into_py(self, path: &str) -> PyErr {
|
||||||
match self {
|
match self {
|
||||||
ReadError::Lib(e) => to_py_err(e),
|
ReadError::Lib(e) => to_py_err(e),
|
||||||
|
ReadError::Vl(e) => e.into_py(&node::name(path)),
|
||||||
|
ReadError::Py(e) => e,
|
||||||
ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))),
|
ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))),
|
||||||
ReadError::Panic(msg) => crate::InternalError::new_err(format!(
|
ReadError::Panic(msg) => crate::InternalError::new_err(format!(
|
||||||
"{}: clawhdf5 internal error (please report it): {msg}",
|
"{}: clawhdf5 internal error (please report it): {msg}",
|
||||||
@@ -243,7 +318,7 @@ impl PyDataset {
|
|||||||
/// The shape of the dataset (`None` for an empty/null dataspace).
|
/// The shape of the dataset (`None` for an empty/null dataspace).
|
||||||
#[getter]
|
#[getter]
|
||||||
fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||||
match &self.shape {
|
match self.dims(py)? {
|
||||||
Some(s) => Ok(PyTuple::new(py, s)?.into_any()),
|
Some(s) => Ok(PyTuple::new(py, s)?.into_any()),
|
||||||
None => Ok(py.None().into_bound(py)),
|
None => Ok(py.None().into_bound(py)),
|
||||||
}
|
}
|
||||||
@@ -252,22 +327,32 @@ impl PyDataset {
|
|||||||
/// The maximum shape (`None` per unlimited dimension), like h5py.
|
/// The maximum shape (`None` per unlimited dimension), like h5py.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||||
crate::no_panic(|| {
|
let Some(shape) = self.dims(py)? else {
|
||||||
let Some(shape) = &self.shape else {
|
return Ok(py.None().into_bound(py));
|
||||||
return Ok(py.None().into_bound(py));
|
};
|
||||||
};
|
let addr = self.addr;
|
||||||
let max = self
|
let max = self
|
||||||
.file
|
.handle
|
||||||
.dataset_at(self.addr)
|
.with(py, |f| {
|
||||||
.and_then(|ds| ds.max_dimensions())
|
f.dataset_at(addr)
|
||||||
.map_err(to_py_err)?
|
.and_then(|ds| ds.max_dimensions())
|
||||||
.unwrap_or_else(|| shape.clone());
|
.map_err(to_py_err)
|
||||||
let items: Vec<Option<u64>> = max
|
})?
|
||||||
.into_iter()
|
.unwrap_or(shape);
|
||||||
.map(|d| (d != u64::MAX).then_some(d))
|
let items: Vec<Option<u64>> = max
|
||||||
.collect();
|
.into_iter()
|
||||||
Ok(PyTuple::new(py, items)?.into_any())
|
.map(|d| (d != u64::MAX).then_some(d))
|
||||||
})
|
.collect();
|
||||||
|
Ok(PyTuple::new(py, items)?.into_any())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The chunk shape, or `None` for a dataset that is not chunked.
|
||||||
|
#[getter]
|
||||||
|
fn chunks<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||||
|
match &self.chunks {
|
||||||
|
Some(c) => Ok(PyTuple::new(py, c)?.into_any()),
|
||||||
|
None => Ok(py.None().into_bound(py)),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The dataset's numpy dtype, as h5py reports it.
|
/// The dataset's numpy dtype, as h5py reports it.
|
||||||
@@ -277,14 +362,14 @@ impl PyDataset {
|
|||||||
}
|
}
|
||||||
|
|
||||||
#[getter]
|
#[getter]
|
||||||
fn ndim(&self) -> usize {
|
fn ndim(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
self.shape.as_ref().map_or(0, Vec::len)
|
Ok(self.dims(py)?.map_or(0, |s| s.len()))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Number of elements (`None` for an empty/null dataspace, as h5py).
|
/// Number of elements (`None` for an empty/null dataspace, as h5py).
|
||||||
#[getter]
|
#[getter]
|
||||||
fn size(&self) -> Option<u64> {
|
fn size(&self, py: Python<'_>) -> PyResult<Option<u64>> {
|
||||||
self.shape.as_ref().map(|s| s.iter().product())
|
Ok(self.dims(py)?.map(|s| s.iter().product()))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The dataset's full name, e.g. `/group/data`.
|
/// The dataset's full name, e.g. `/group/data`.
|
||||||
@@ -293,10 +378,11 @@ impl PyDataset {
|
|||||||
node::name(&self.path)
|
node::name(&self.path)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The dataset's attributes (read-only, dict-like).
|
/// The dataset's attributes (dict-like; writable in a file opened with
|
||||||
|
/// `'r+'`).
|
||||||
#[getter]
|
#[getter]
|
||||||
fn attrs(&self) -> PyResult<PyAttrs> {
|
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||||
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
|
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Read with h5py indexing: integers, slices with positive steps,
|
/// Read with h5py indexing: integers, slices with positive steps,
|
||||||
@@ -308,7 +394,7 @@ impl PyDataset {
|
|||||||
py: Python<'py>,
|
py: Python<'py>,
|
||||||
key: &Bound<'py, PyAny>,
|
key: &Bound<'py, PyAny>,
|
||||||
) -> PyResult<Bound<'py, PyAny>> {
|
) -> PyResult<Bound<'py, PyAny>> {
|
||||||
let Some(dims) = &self.shape else {
|
let Some(dims) = self.dims(py)? else {
|
||||||
let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty());
|
let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty());
|
||||||
let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>();
|
let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>();
|
||||||
if is_empty_tuple || is_ellipsis {
|
if is_empty_tuple || is_ellipsis {
|
||||||
@@ -317,8 +403,109 @@ impl PyDataset {
|
|||||||
}
|
}
|
||||||
return Err(PyValueError::new_err("Empty datasets cannot be sliced"));
|
return Err(PyValueError::new_err("Empty datasets cannot be sliced"));
|
||||||
};
|
};
|
||||||
let plan = select::parse(key, dims)?;
|
let plan = select::parse(key, &dims)?;
|
||||||
self.read_plan(py, &plan)
|
self.read_plan(py, &plan, &dims)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Write with h5py indexing (file opened with `'r+'`): `ds[key] = value`.
|
||||||
|
///
|
||||||
|
/// The key is what `ds[key]` reads (without compound field names). The
|
||||||
|
/// value is converted to the dataset's dtype as h5py converts it (a
|
||||||
|
/// numpy array as libhdf5 does, clipping out-of-range numbers; anything
|
||||||
|
/// else through `numpy.asarray(value, dtype=ds.dtype)`), and broadcast
|
||||||
|
/// to the selection as h5py broadcasts. The edit is written and synced
|
||||||
|
/// before this returns; what the in-place editor cannot write raises
|
||||||
|
/// `NotImplementedError` and leaves the file as it was.
|
||||||
|
fn __setitem__(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
key: &Bound<'_, PyAny>,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
) -> PyResult<()> {
|
||||||
|
self.check_writable()?;
|
||||||
|
let Some(dims) = self.dims(py)? else {
|
||||||
|
return Err(PyNotImplementedError::new_err(
|
||||||
|
"writing to an empty (null dataspace) dataset is not supported",
|
||||||
|
));
|
||||||
|
};
|
||||||
|
let plan = select::parse(key, &dims)?;
|
||||||
|
if !plan.fields.is_empty() {
|
||||||
|
return Err(PyNotImplementedError::new_err(
|
||||||
|
"writing compound fields by name is not supported by clawhdf5's in-place editor; \
|
||||||
|
write whole elements",
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let category = edit::category(&self.datatype)?;
|
||||||
|
let conv = self.converter()?;
|
||||||
|
let bytes = edit::dataset_bytes(
|
||||||
|
py,
|
||||||
|
value,
|
||||||
|
conv.dtype.bind(py),
|
||||||
|
category,
|
||||||
|
&plan,
|
||||||
|
self.chunks.as_deref(),
|
||||||
|
)?;
|
||||||
|
if plan.is_empty() {
|
||||||
|
return Ok(());
|
||||||
|
}
|
||||||
|
let sel = edit::selection(&plan, &dims)?;
|
||||||
|
let path = node::name(&self.path);
|
||||||
|
self.handle
|
||||||
|
.edit(py, |ed| ed.write_selection(&path, &sel, &bytes))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Change the dataset's shape (file opened with `'r+'`), as h5py's
|
||||||
|
/// `Dataset.resize`: `ds.resize((100, 20))`, or `ds.resize(100, axis=0)`.
|
||||||
|
/// Only chunked datasets, within their maximum shape; new elements read
|
||||||
|
/// as the fill value.
|
||||||
|
#[pyo3(signature = (size, axis=None))]
|
||||||
|
fn resize(&self, py: Python<'_>, size: &Bound<'_, PyAny>, axis: Option<isize>) -> PyResult<()> {
|
||||||
|
self.check_writable()?;
|
||||||
|
let Some(dims) = self.dims(py)? else {
|
||||||
|
return Err(PyTypeError::new_err("Empty datasets cannot be resized"));
|
||||||
|
};
|
||||||
|
if self.chunks.is_none() {
|
||||||
|
return Err(PyTypeError::new_err("Only chunked datasets can be resized"));
|
||||||
|
}
|
||||||
|
let shape: Vec<u64> = match axis {
|
||||||
|
Some(axis) => {
|
||||||
|
let rank = dims.len();
|
||||||
|
let a = usize::try_from(axis)
|
||||||
|
.ok()
|
||||||
|
.filter(|&a| a < rank)
|
||||||
|
.ok_or_else(|| {
|
||||||
|
PyValueError::new_err(format!(
|
||||||
|
"Invalid axis (0 to {} allowed)",
|
||||||
|
rank.saturating_sub(1)
|
||||||
|
))
|
||||||
|
})?;
|
||||||
|
let n: u64 = size.extract().map_err(|_| {
|
||||||
|
PyTypeError::new_err("Argument must be a single int if axis is specified")
|
||||||
|
})?;
|
||||||
|
let mut s = dims.clone();
|
||||||
|
s[a] = n;
|
||||||
|
s
|
||||||
|
}
|
||||||
|
// As h5py: without `axis` the size is a sequence (`tuple(size)`).
|
||||||
|
None => size.extract().map_err(|_| {
|
||||||
|
PyTypeError::new_err(format!(
|
||||||
|
"'{}' object is not iterable",
|
||||||
|
size.get_type()
|
||||||
|
.name()
|
||||||
|
.map(|n| n.to_string())
|
||||||
|
.unwrap_or_default()
|
||||||
|
))
|
||||||
|
})?,
|
||||||
|
};
|
||||||
|
if shape.len() != dims.len() {
|
||||||
|
return Err(PyValueError::new_err(format!(
|
||||||
|
"new shape {shape:?} has {} dimensions, the dataset {}",
|
||||||
|
shape.len(),
|
||||||
|
dims.len()
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
let path = node::name(&self.path);
|
||||||
|
self.handle.edit(py, |ed| ed.resize(&path, &shape))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// `numpy.asarray(ds)` reads the whole dataset.
|
/// `numpy.asarray(ds)` reads the whole dataset.
|
||||||
@@ -330,20 +517,20 @@ impl PyDataset {
|
|||||||
copy: Option<bool>,
|
copy: Option<bool>,
|
||||||
) -> PyResult<Bound<'py, PyAny>> {
|
) -> PyResult<Bound<'py, PyAny>> {
|
||||||
let _ = copy; // every read is a fresh array
|
let _ = copy; // every read is a fresh array
|
||||||
let Some(dims) = &self.shape else {
|
let Some(dims) = self.dims(py)? else {
|
||||||
return Err(PyValueError::new_err("an empty dataset has no array value"));
|
return Err(PyValueError::new_err("an empty dataset has no array value"));
|
||||||
};
|
};
|
||||||
let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any();
|
let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any();
|
||||||
let plan = select::parse(&ellipsis, dims)?;
|
let plan = select::parse(&ellipsis, &dims)?;
|
||||||
let arr = self.read_plan(py, &plan)?;
|
let arr = self.read_plan(py, &plan, &dims)?;
|
||||||
match dtype {
|
match dtype {
|
||||||
Some(dt) => arr.call_method1("astype", (dt,)),
|
Some(dt) => arr.call_method1("astype", (dt,)),
|
||||||
None => Ok(arr),
|
None => Ok(arr),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __len__(&self) -> PyResult<usize> {
|
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
match self.shape.as_deref() {
|
match self.dims(py)?.as_deref() {
|
||||||
Some([first, ..]) => Ok(*first as usize),
|
Some([first, ..]) => Ok(*first as usize),
|
||||||
_ => Err(PyTypeError::new_err(
|
_ => Err(PyTypeError::new_err(
|
||||||
"Attempt to take len() of scalar dataset",
|
"Attempt to take len() of scalar dataset",
|
||||||
@@ -361,9 +548,10 @@ impl PyDataset {
|
|||||||
.unwrap_or_default(),
|
.unwrap_or_default(),
|
||||||
Err(_) => format!("{:?}", self.datatype),
|
Err(_) => format!("{:?}", self.datatype),
|
||||||
};
|
};
|
||||||
let shape = match &self.shape {
|
let shape = match self.dims(py) {
|
||||||
Some(s) => format!("{s:?}"),
|
Ok(Some(s)) => format!("{s:?}"),
|
||||||
None => "None".to_string(),
|
Ok(None) => "None".to_string(),
|
||||||
|
Err(_) => "?".to_string(),
|
||||||
};
|
};
|
||||||
format!(
|
format!(
|
||||||
"<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">",
|
"<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">",
|
||||||
|
|||||||
@@ -0,0 +1,387 @@
|
|||||||
|
//! In-place editing (`clawhdf5.File(path, 'r+')`) through `FileEditor`:
|
||||||
|
//! turning what Python assigns into the bytes, selections and attribute
|
||||||
|
//! values the editor takes.
|
||||||
|
//!
|
||||||
|
//! Value conversion follows h5py (see `edit_helpers.py`, run inside the
|
||||||
|
//! extension module); what `FileEditor` cannot do is `NotImplementedError`
|
||||||
|
//! before anything is written.
|
||||||
|
|
||||||
|
use std::ffi::CString;
|
||||||
|
|
||||||
|
use clawhdf5_format::datatype::{
|
||||||
|
CharacterSet, CompoundMember, Datatype, DatatypeByteOrder, EnumMember, StringPadding,
|
||||||
|
};
|
||||||
|
use clawhdf5_format::selection::Selection;
|
||||||
|
use clawhdf5_rs::AttrValue;
|
||||||
|
use pyo3::exceptions::{PyNotImplementedError, PyTypeError};
|
||||||
|
use pyo3::prelude::*;
|
||||||
|
use pyo3::sync::PyOnceLock;
|
||||||
|
use pyo3::types::{PyBytes, PyModule, PyTuple};
|
||||||
|
|
||||||
|
use crate::select::{Axis, Plan};
|
||||||
|
|
||||||
|
/// The helper module, compiled once.
|
||||||
|
pub(crate) fn helpers(py: Python<'_>) -> PyResult<&Bound<'_, PyModule>> {
|
||||||
|
static HELPERS: PyOnceLock<Py<PyModule>> = PyOnceLock::new();
|
||||||
|
let module = HELPERS.get_or_try_init(py, || -> PyResult<Py<PyModule>> {
|
||||||
|
let code = CString::new(include_str!("edit_helpers.py"))
|
||||||
|
.map_err(|e| PyTypeError::new_err(e.to_string()))?;
|
||||||
|
Ok(PyModule::from_code(
|
||||||
|
py,
|
||||||
|
&code,
|
||||||
|
c"clawhdf5/edit_helpers.py",
|
||||||
|
c"clawhdf5._edit_helpers",
|
||||||
|
)?
|
||||||
|
.unbind())
|
||||||
|
})?;
|
||||||
|
Ok(module.bind(py))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn not_implemented(what: impl std::fmt::Display) -> PyErr {
|
||||||
|
PyNotImplementedError::new_err(format!(
|
||||||
|
"{what} is not supported by clawhdf5's in-place editor"
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// How values for a dataset of type `dt` are converted (a category of
|
||||||
|
/// `edit_helpers._convert_array`), or why they cannot be written.
|
||||||
|
pub(crate) fn category(dt: &Datatype) -> PyResult<&'static str> {
|
||||||
|
match dt {
|
||||||
|
Datatype::FixedPoint { .. } => Ok("int"),
|
||||||
|
Datatype::FloatingPoint { .. } => Ok("float"),
|
||||||
|
Datatype::Enumeration {
|
||||||
|
base_type, members, ..
|
||||||
|
} => {
|
||||||
|
let is_bool = base_type.type_size() == 1
|
||||||
|
&& members.len() == 2
|
||||||
|
&& members
|
||||||
|
.iter()
|
||||||
|
.any(|m| m.name == "FALSE" && m.value.first() == Some(&0))
|
||||||
|
&& members
|
||||||
|
.iter()
|
||||||
|
.any(|m| m.name == "TRUE" && m.value.first() == Some(&1));
|
||||||
|
Ok(if is_bool { "bool" } else { "enum" })
|
||||||
|
}
|
||||||
|
Datatype::String {
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
..
|
||||||
|
} => Ok("string"),
|
||||||
|
Datatype::String { padding, .. } => Err(not_implemented(format!(
|
||||||
|
"writing fixed-length strings padded {padding:?} (libhdf5 converts them \
|
||||||
|
differently from numpy)"
|
||||||
|
))),
|
||||||
|
Datatype::Compound { size, members } => {
|
||||||
|
if is_complex(*size, members) {
|
||||||
|
return Ok("complex");
|
||||||
|
}
|
||||||
|
check_exact(dt)?;
|
||||||
|
Ok("exact")
|
||||||
|
}
|
||||||
|
Datatype::Opaque { .. } => Ok("exact"),
|
||||||
|
Datatype::Array { .. } => Err(not_implemented("writing HDF5 array-type elements")),
|
||||||
|
Datatype::VariableLength { .. } => Err(not_implemented("writing variable-length data")),
|
||||||
|
Datatype::Reference { .. } => Err(not_implemented("writing references")),
|
||||||
|
Datatype::BitField { .. } => Err(not_implemented("writing bitfields")),
|
||||||
|
Datatype::Time { .. } => Err(not_implemented("writing time values")),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// h5py's complex numbers: a compound of two identical floats `r`, `i`.
|
||||||
|
fn is_complex(size: u32, members: &[CompoundMember]) -> bool {
|
||||||
|
matches!(members, [r, i] if r.name == "r" && i.name == "i"
|
||||||
|
&& r.datatype == i.datatype
|
||||||
|
&& matches!(r.datatype, Datatype::FloatingPoint { size: fs, .. }
|
||||||
|
if r.byte_offset == 0 && i.byte_offset == u64::from(fs) && size == 2 * fs))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Compound members written byte for byte from the same numpy dtype: fine
|
||||||
|
/// unless libhdf5 would convert them on the way (strings padded other than
|
||||||
|
/// with NULs), or the editor cannot write them at all.
|
||||||
|
fn check_exact(dt: &Datatype) -> PyResult<()> {
|
||||||
|
match dt {
|
||||||
|
Datatype::Compound { members, .. } => {
|
||||||
|
members.iter().try_for_each(|m| check_exact(&m.datatype))
|
||||||
|
}
|
||||||
|
Datatype::Array { base_type, .. } => check_exact(base_type),
|
||||||
|
Datatype::String {
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
..
|
||||||
|
}
|
||||||
|
| Datatype::FixedPoint { .. }
|
||||||
|
| Datatype::FloatingPoint { .. }
|
||||||
|
| Datatype::Enumeration { .. }
|
||||||
|
| Datatype::Opaque { .. }
|
||||||
|
| Datatype::BitField { .. } => Ok(()),
|
||||||
|
Datatype::String { .. } => Err(not_implemented(
|
||||||
|
"writing compounds with strings not padded with NULs",
|
||||||
|
)),
|
||||||
|
Datatype::VariableLength { .. } => Err(not_implemented(
|
||||||
|
"writing compounds with variable-length members",
|
||||||
|
)),
|
||||||
|
Datatype::Reference { .. } => Err(not_implemented("writing references")),
|
||||||
|
Datatype::Time { .. } => Err(not_implemented("writing time values")),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Largest point selection an index-list write builds (one coordinate
|
||||||
|
/// vector per element).
|
||||||
|
const MAX_POINTS: usize = 1 << 22;
|
||||||
|
|
||||||
|
/// The selection `plan` writes, whose elements are numbered as the value's
|
||||||
|
/// (row-major over the selection's shape).
|
||||||
|
pub(crate) fn selection(plan: &Plan, dims: &[u64]) -> PyResult<Selection> {
|
||||||
|
if plan.axes.is_empty() {
|
||||||
|
return Ok(Selection::All);
|
||||||
|
}
|
||||||
|
if plan.list_axis().is_none() {
|
||||||
|
let (reads, _) = plan.reads(dims, None, 1);
|
||||||
|
return match <[_; 1]>::try_from(reads) {
|
||||||
|
Ok([read]) => Ok(read.sel),
|
||||||
|
Err(_) => Err(PyTypeError::new_err("internal error: several hyperslabs")),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
// An index list: the points, in the value's order.
|
||||||
|
let per_axis: Vec<Vec<u64>> = plan
|
||||||
|
.axes
|
||||||
|
.iter()
|
||||||
|
.map(|a| match a {
|
||||||
|
Axis::Index(i) => vec![*i],
|
||||||
|
Axis::Slice { start, step, count } => (0..*count).map(|k| start + k * step).collect(),
|
||||||
|
Axis::List(v) => v.clone(),
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
let n = per_axis
|
||||||
|
.iter()
|
||||||
|
.try_fold(1usize, |acc, v| acc.checked_mul(v.len()))
|
||||||
|
.filter(|&n| n <= MAX_POINTS)
|
||||||
|
.ok_or_else(|| {
|
||||||
|
not_implemented(format!(
|
||||||
|
"an index-list write of more than {MAX_POINTS} elements (write it in slices)"
|
||||||
|
))
|
||||||
|
})?;
|
||||||
|
let mut points = Vec::with_capacity(n);
|
||||||
|
let mut at = vec![0usize; per_axis.len()];
|
||||||
|
for _ in 0..n {
|
||||||
|
points.push(at.iter().zip(&per_axis).map(|(&i, v)| v[i]).collect());
|
||||||
|
for d in (0..at.len()).rev() {
|
||||||
|
at[d] += 1;
|
||||||
|
if at[d] < per_axis[d].len() {
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
at[d] = 0;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Ok(Selection::Points(points))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The bytes to write for `value` under `plan`, in the dataset's dtype.
|
||||||
|
pub(crate) fn dataset_bytes(
|
||||||
|
py: Python<'_>,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
dtype: &Bound<'_, PyAny>,
|
||||||
|
category: &str,
|
||||||
|
plan: &Plan,
|
||||||
|
chunks: Option<&[u64]>,
|
||||||
|
) -> PyResult<Vec<u8>> {
|
||||||
|
let shape = PyTuple::new(py, plan.out_shape())?;
|
||||||
|
let fancy = plan.list_axis().is_some();
|
||||||
|
let chunk_elems = chunks.map_or(0, |c| c.iter().fold(1u64, |a, &d| a.saturating_mul(d)));
|
||||||
|
let bytes = helpers(py)?.call_method1(
|
||||||
|
"dataset_values",
|
||||||
|
(value, dtype, category, shape, fancy, chunk_elems),
|
||||||
|
)?;
|
||||||
|
Ok(bytes.cast::<PyBytes>()?.as_bytes().to_vec())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn ieee_float(size: u32, byte_order: DatatypeByteOrder) -> Option<Datatype> {
|
||||||
|
let (exponent_location, exponent_size, mantissa_size, exponent_bias) = match size {
|
||||||
|
2 => (10, 5, 10, 15),
|
||||||
|
4 => (23, 8, 23, 127),
|
||||||
|
8 => (52, 11, 52, 1023),
|
||||||
|
_ => return None,
|
||||||
|
};
|
||||||
|
Some(Datatype::FloatingPoint {
|
||||||
|
size,
|
||||||
|
byte_order,
|
||||||
|
bit_offset: 0,
|
||||||
|
bit_precision: (size * 8) as u16,
|
||||||
|
exponent_location,
|
||||||
|
exponent_size,
|
||||||
|
mantissa_location: 0,
|
||||||
|
mantissa_size,
|
||||||
|
exponent_bias,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The HDF5 datatype h5py writes for a numpy dtype string (`'<i4'`,
|
||||||
|
/// `'|b1'`, `'>f8'`, `'<c16'`, `'|S5'`).
|
||||||
|
fn datatype_of(dtype: &str) -> Option<Datatype> {
|
||||||
|
let order = match dtype.as_bytes().first()? {
|
||||||
|
b'<' | b'|' | b'=' => DatatypeByteOrder::LittleEndian,
|
||||||
|
b'>' => DatatypeByteOrder::BigEndian,
|
||||||
|
_ => return None,
|
||||||
|
};
|
||||||
|
let kind = dtype.as_bytes().get(1)?;
|
||||||
|
let size: u32 = dtype.get(2..)?.parse().ok()?;
|
||||||
|
match kind {
|
||||||
|
b'b' if size == 1 => Some(Datatype::Enumeration {
|
||||||
|
size: 1,
|
||||||
|
base_type: Box::new(Datatype::FixedPoint {
|
||||||
|
size: 1,
|
||||||
|
byte_order: DatatypeByteOrder::LittleEndian,
|
||||||
|
signed: true,
|
||||||
|
bit_offset: 0,
|
||||||
|
bit_precision: 8,
|
||||||
|
}),
|
||||||
|
members: vec![
|
||||||
|
EnumMember {
|
||||||
|
name: "FALSE".into(),
|
||||||
|
value: vec![0],
|
||||||
|
},
|
||||||
|
EnumMember {
|
||||||
|
name: "TRUE".into(),
|
||||||
|
value: vec![1],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
}),
|
||||||
|
b'i' | b'u' if matches!(size, 1 | 2 | 4 | 8) => Some(Datatype::FixedPoint {
|
||||||
|
size,
|
||||||
|
byte_order: order,
|
||||||
|
signed: *kind == b'i',
|
||||||
|
bit_offset: 0,
|
||||||
|
bit_precision: (size * 8) as u16,
|
||||||
|
}),
|
||||||
|
b'f' => ieee_float(size, order),
|
||||||
|
b'c' => {
|
||||||
|
let part = ieee_float(size / 2, order)?;
|
||||||
|
Some(Datatype::Compound {
|
||||||
|
size,
|
||||||
|
members: vec![
|
||||||
|
CompoundMember {
|
||||||
|
name: "r".into(),
|
||||||
|
byte_offset: 0,
|
||||||
|
datatype: part.clone(),
|
||||||
|
},
|
||||||
|
CompoundMember {
|
||||||
|
name: "i".into(),
|
||||||
|
byte_offset: u64::from(size / 2),
|
||||||
|
datatype: part,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
})
|
||||||
|
}
|
||||||
|
b'S' if size > 0 => Some(Datatype::String {
|
||||||
|
size,
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
charset: CharacterSet::Ascii,
|
||||||
|
}),
|
||||||
|
_ => None,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// An attribute value as h5py would store it (`attrs[name] = value`, or
|
||||||
|
/// `attrs.create(name, data, shape, dtype)`), except that `str` data is
|
||||||
|
/// stored as fixed-length UTF-8 strings (h5py stores variable-length ones,
|
||||||
|
/// which the editor cannot write).
|
||||||
|
pub(crate) fn attr_value(
|
||||||
|
py: Python<'_>,
|
||||||
|
value: &Bound<'_, PyAny>,
|
||||||
|
dtype: Option<&Bound<'_, PyAny>>,
|
||||||
|
shape: Option<&Bound<'_, PyAny>>,
|
||||||
|
) -> PyResult<AttrValue> {
|
||||||
|
if value.is_instance_of::<crate::PyEmpty>() {
|
||||||
|
return Err(not_implemented(
|
||||||
|
"writing an empty (null dataspace) attribute",
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let (kind, dt, dims, data): (String, String, Vec<u64>, Vec<u8>) = helpers(py)?
|
||||||
|
.call_method1("attr_value", (value, dtype, shape))?
|
||||||
|
.extract()?;
|
||||||
|
let datatype = if kind == "str" {
|
||||||
|
let size: u32 = dt
|
||||||
|
.parse()
|
||||||
|
.map_err(|_| PyTypeError::new_err("bad string size"))?;
|
||||||
|
Datatype::String {
|
||||||
|
size,
|
||||||
|
padding: StringPadding::NullPad,
|
||||||
|
charset: CharacterSet::Utf8,
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
datatype_of(&dt).ok_or_else(|| not_implemented(format!("an attribute of dtype {dt}")))?
|
||||||
|
};
|
||||||
|
Ok(AttrValue::Raw {
|
||||||
|
datatype,
|
||||||
|
shape: dims,
|
||||||
|
data,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod tests {
|
||||||
|
use super::*;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn numpy_dtypes_map_to_h5py_types() {
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of("<i4"),
|
||||||
|
Some(Datatype::FixedPoint {
|
||||||
|
size: 4,
|
||||||
|
signed: true,
|
||||||
|
byte_order: DatatypeByteOrder::LittleEndian,
|
||||||
|
..
|
||||||
|
})
|
||||||
|
));
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of(">u2"),
|
||||||
|
Some(Datatype::FixedPoint {
|
||||||
|
size: 2,
|
||||||
|
signed: false,
|
||||||
|
byte_order: DatatypeByteOrder::BigEndian,
|
||||||
|
..
|
||||||
|
})
|
||||||
|
));
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of("<f2"),
|
||||||
|
Some(Datatype::FloatingPoint { size: 2, .. })
|
||||||
|
));
|
||||||
|
let c = datatype_of("<c16").unwrap();
|
||||||
|
assert!(matches!(&c, Datatype::Compound { size: 16, members } if is_complex(16, members)));
|
||||||
|
assert_eq!(category(&c).unwrap(), "complex");
|
||||||
|
let b = datatype_of("|b1").unwrap();
|
||||||
|
assert_eq!(category(&b).unwrap(), "bool");
|
||||||
|
assert!(matches!(
|
||||||
|
datatype_of("|S5"),
|
||||||
|
Some(Datatype::String { size: 5, .. })
|
||||||
|
));
|
||||||
|
assert!(datatype_of("<f16").is_none());
|
||||||
|
assert!(datatype_of("<M8").is_none());
|
||||||
|
assert!(datatype_of("|S0").is_none());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn list_writes_become_points_in_value_order() {
|
||||||
|
let plan = Plan {
|
||||||
|
axes: vec![
|
||||||
|
Axis::List(vec![1, 4]),
|
||||||
|
Axis::Slice {
|
||||||
|
start: 0,
|
||||||
|
step: 2,
|
||||||
|
count: 2,
|
||||||
|
},
|
||||||
|
Axis::Index(3),
|
||||||
|
],
|
||||||
|
fields: vec![],
|
||||||
|
scalar: false,
|
||||||
|
};
|
||||||
|
let sel = selection(&plan, &[5, 4, 4]).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
sel,
|
||||||
|
Selection::Points(vec![
|
||||||
|
vec![1, 0, 3],
|
||||||
|
vec![1, 2, 3],
|
||||||
|
vec![4, 0, 3],
|
||||||
|
vec![4, 2, 3]
|
||||||
|
])
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,175 @@
|
|||||||
|
"""Values for in-place writes (clawhdf5.File(path, 'r+')), prepared the way
|
||||||
|
h5py prepares them, so `ds[key] = value` stores what h5py would store.
|
||||||
|
|
||||||
|
Loaded by the extension module (src/edit.rs); not a public API.
|
||||||
|
|
||||||
|
h5py converts in two ways, and so does this module:
|
||||||
|
|
||||||
|
- a value that is not a numpy array (a list, a Python or numpy scalar) is
|
||||||
|
converted by numpy straight to the dataset's dtype
|
||||||
|
(`numpy.asarray(value, dtype=ds.dtype)`), with numpy's rules and errors;
|
||||||
|
- a numpy array is converted by libhdf5, whose numeric conversions clip to
|
||||||
|
the target's range instead of wrapping: integers saturate, floats are
|
||||||
|
truncated toward zero and clipped, a double too large for a float becomes
|
||||||
|
infinity. That is what `_convert_array` reproduces. Where libhdf5 has no
|
||||||
|
meaningful answer — NaN into an integer, for which it writes a different
|
||||||
|
arbitrary value per type — this raises ValueError instead of guessing.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
|
||||||
|
def _no_path(src, dst):
|
||||||
|
return TypeError(f"No conversion path for dtype: {src!r} -> {dst!r}")
|
||||||
|
|
||||||
|
|
||||||
|
def _to_int(arr, dtype):
|
||||||
|
"""Integer target: libhdf5's saturating conversion."""
|
||||||
|
info = np.iinfo(dtype)
|
||||||
|
kind = arr.dtype.kind
|
||||||
|
if kind == "b":
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if kind in "iu":
|
||||||
|
src = np.iinfo(arr.dtype)
|
||||||
|
lo = max(info.min, src.min)
|
||||||
|
hi = min(info.max, src.max)
|
||||||
|
clipped = np.clip(arr, np.array(lo, arr.dtype), np.array(hi, arr.dtype))
|
||||||
|
return clipped.astype(dtype)
|
||||||
|
if kind == "f":
|
||||||
|
if np.isnan(arr).any():
|
||||||
|
raise ValueError(
|
||||||
|
"cannot write NaN to an integer dataset (libhdf5 would store an arbitrary value)"
|
||||||
|
)
|
||||||
|
t = np.trunc(arr.astype(np.float64))
|
||||||
|
# info.max + 1 and info.min are powers of two: exact as floats.
|
||||||
|
over = t >= float(info.max + 1)
|
||||||
|
under = t < float(info.min)
|
||||||
|
out = np.where(over | under, 0.0, t).astype(dtype)
|
||||||
|
out[over] = info.max
|
||||||
|
out[under] = info.min
|
||||||
|
return out
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
|
||||||
|
|
||||||
|
def _convert_array(arr, dtype, category):
|
||||||
|
kind = arr.dtype.kind
|
||||||
|
if category in ("int", "enum"):
|
||||||
|
if arr.dtype == dtype and kind in "iu":
|
||||||
|
return arr
|
||||||
|
if category == "enum" and kind not in "iu":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
return _to_int(arr, dtype)
|
||||||
|
if category == "bool":
|
||||||
|
if kind == "b":
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if kind in "iu":
|
||||||
|
# h5py's bool is an enum over int8; libhdf5 converts integers
|
||||||
|
# into it by value (saturating), not to FALSE/TRUE, so 3 is
|
||||||
|
# stored as 3. Keep those bytes: a view, not a cast.
|
||||||
|
return _to_int(arr, np.dtype("i1")).view(dtype)
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
if category == "float":
|
||||||
|
if kind not in "biuf":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
with np.errstate(over="ignore", invalid="ignore"):
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if category == "complex":
|
||||||
|
if kind != "c":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
with np.errstate(over="ignore", invalid="ignore"):
|
||||||
|
return arr.astype(dtype)
|
||||||
|
if category == "string":
|
||||||
|
if kind != "S":
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
return arr.astype(dtype)
|
||||||
|
# "exact": compound and opaque types, written only from the same dtype.
|
||||||
|
if arr.dtype == dtype:
|
||||||
|
return arr
|
||||||
|
raise _no_path(arr.dtype, dtype)
|
||||||
|
|
||||||
|
|
||||||
|
def _convert_other(value, dtype, category):
|
||||||
|
if category == "string":
|
||||||
|
items = np.asarray(value, dtype=object)
|
||||||
|
if any(isinstance(x, str) for x in items.flat):
|
||||||
|
meta = dtype.metadata or {}
|
||||||
|
if meta.get("h5py_encoding") == "utf-8":
|
||||||
|
enc = [x.encode("utf-8") if isinstance(x, str) else x for x in items.flat]
|
||||||
|
return np.array(enc, dtype=dtype).reshape(items.shape)
|
||||||
|
return np.asarray(value, dtype=dtype)
|
||||||
|
|
||||||
|
|
||||||
|
def _broadcast(arr, shape, fancy, chunk_elems):
|
||||||
|
"""h5py's broadcasting: numpy's rules against the selection's shape
|
||||||
|
(extra leading length-1 axes allowed) for slices and integers. For an
|
||||||
|
index list, the exact shape; a scalar only where h5py expands it to the
|
||||||
|
whole selection (a chunked dataset whose chunk holds at least as many
|
||||||
|
elements as the selection)."""
|
||||||
|
if arr.shape == shape:
|
||||||
|
return arr
|
||||||
|
if fancy:
|
||||||
|
size = int(np.prod(shape))
|
||||||
|
if arr.ndim == 0 and ((chunk_elems > 0 and size <= chunk_elems) or len(shape) == 1):
|
||||||
|
return np.broadcast_to(arr, shape)
|
||||||
|
raise TypeError("Broadcasting is not supported for complex selections")
|
||||||
|
if arr.ndim == 0:
|
||||||
|
return np.broadcast_to(arr, shape)
|
||||||
|
err = TypeError(f"Can't broadcast {arr.shape} -> {shape}")
|
||||||
|
src = arr.shape
|
||||||
|
while len(src) > len(shape) and src[0] == 1:
|
||||||
|
src = src[1:]
|
||||||
|
if len(src) > len(shape):
|
||||||
|
raise err
|
||||||
|
try:
|
||||||
|
return np.broadcast_to(arr.reshape(src), shape)
|
||||||
|
except ValueError:
|
||||||
|
raise err from None
|
||||||
|
|
||||||
|
|
||||||
|
def dataset_values(value, dtype, category, shape, fancy, chunk_elems):
|
||||||
|
"""The bytes to write for `value` under a selection of `shape`, as a
|
||||||
|
C-ordered array of the dataset's dtype."""
|
||||||
|
if isinstance(value, np.ndarray):
|
||||||
|
arr = _convert_array(value, dtype, category)
|
||||||
|
else:
|
||||||
|
arr = _convert_other(value, dtype, category)
|
||||||
|
arr = _broadcast(arr, tuple(shape), fancy, chunk_elems)
|
||||||
|
return np.ascontiguousarray(arr, dtype=dtype).tobytes()
|
||||||
|
|
||||||
|
|
||||||
|
def attr_value(value, dtype=None, shape=None):
|
||||||
|
"""(kind, dtype string, shape, bytes) for an attribute value, h5py's
|
||||||
|
`attrs[name] = value` / `attrs.create(name, data, shape, dtype)`:
|
||||||
|
|
||||||
|
- "str": `str` data (h5py would store a variable-length string; this
|
||||||
|
stores a fixed-length UTF-8 string, which clawhdf5 can write);
|
||||||
|
the dtype string is the byte length of the longest element;
|
||||||
|
- "raw": a numeric, bool or bytes array, as numpy lays it out.
|
||||||
|
"""
|
||||||
|
if dtype is not None:
|
||||||
|
arr = np.asarray(value, dtype=dtype, order="C")
|
||||||
|
else:
|
||||||
|
arr = np.asarray(value, order="C")
|
||||||
|
if shape is not None:
|
||||||
|
arr = arr.reshape(shape)
|
||||||
|
kind = arr.dtype.kind
|
||||||
|
if kind == "O":
|
||||||
|
if arr.size and all(isinstance(x, str) for x in arr.flat):
|
||||||
|
kind = "U"
|
||||||
|
elif arr.size and all(isinstance(x, bytes) for x in arr.flat):
|
||||||
|
arr = arr.astype(bytes)
|
||||||
|
kind = "S"
|
||||||
|
else:
|
||||||
|
raise TypeError(
|
||||||
|
f"clawhdf5 cannot write an attribute of Python objects ({value!r:.60})"
|
||||||
|
)
|
||||||
|
if kind == "U":
|
||||||
|
enc = [str(x).encode("utf-8") for x in arr.flat]
|
||||||
|
size = max([len(b) for b in enc] + [1])
|
||||||
|
data = np.array(enc, dtype=f"S{size}").reshape(arr.shape)
|
||||||
|
return ("str", str(size), arr.shape, data.tobytes())
|
||||||
|
if kind in "biufcS":
|
||||||
|
return ("raw", arr.dtype.str, arr.shape, np.ascontiguousarray(arr).tobytes())
|
||||||
|
raise NotImplementedError(
|
||||||
|
f"clawhdf5 cannot write an attribute of dtype {arr.dtype} in place"
|
||||||
|
)
|
||||||
+228
-32
@@ -1,13 +1,17 @@
|
|||||||
//! PyFile — the main entry point for opening and creating HDF5 files.
|
//! PyFile — the main entry point for opening and creating HDF5 files.
|
||||||
|
|
||||||
|
use std::collections::HashMap;
|
||||||
use std::path::PathBuf;
|
use std::path::PathBuf;
|
||||||
use std::sync::{Arc, Mutex};
|
use std::sync::{Arc, Mutex};
|
||||||
|
use std::time::Duration;
|
||||||
|
|
||||||
|
use pyo3::exceptions::{PyNotImplementedError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::PyList;
|
use pyo3::types::{PyDict, PyList};
|
||||||
|
|
||||||
use crate::attrs::PyAttrs;
|
use crate::attrs::PyAttrs;
|
||||||
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
|
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
|
||||||
|
use crate::handle::Handle;
|
||||||
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
|
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
|
||||||
|
|
||||||
/// Internal state for write mode.
|
/// Internal state for write mode.
|
||||||
@@ -23,8 +27,10 @@ struct WriteState {
|
|||||||
/// Mirrors the h5py.File interface:
|
/// Mirrors the h5py.File interface:
|
||||||
///
|
///
|
||||||
/// ```python
|
/// ```python
|
||||||
/// # Reading
|
/// # Reading, a local file or a URL (range requests, nothing downloaded
|
||||||
|
/// # up front)
|
||||||
/// f = clawhdf5.File('data.h5', 'r')
|
/// f = clawhdf5.File('data.h5', 'r')
|
||||||
|
/// f = clawhdf5.File('https://example.org/data.h5')
|
||||||
/// ds = f['dataset']
|
/// ds = f['dataset']
|
||||||
/// f.close()
|
/// f.close()
|
||||||
///
|
///
|
||||||
@@ -44,49 +50,200 @@ enum FileInner {
|
|||||||
Write(WriteState),
|
Write(WriteState),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Whether `s` is a URL (`scheme://…`) rather than a path: the scheme is a
|
||||||
|
/// letter followed by letters, digits, `+`, `-` or `.` (RFC 3986).
|
||||||
|
fn is_url(s: &str) -> bool {
|
||||||
|
let Some((scheme, _)) = s.split_once("://") else {
|
||||||
|
return false;
|
||||||
|
};
|
||||||
|
let mut chars = scheme.chars();
|
||||||
|
chars.next().is_some_and(|c| c.is_ascii_alphabetic())
|
||||||
|
&& chars.all(|c| c.is_ascii_alphanumeric() || matches!(c, '+' | '-' | '.'))
|
||||||
|
}
|
||||||
|
|
||||||
|
impl PyFile {
|
||||||
|
fn from_handle(handle: Arc<Handle>, filename: String) -> Self {
|
||||||
|
let root = handle.root;
|
||||||
|
Self {
|
||||||
|
inner: Some(FileInner::Read(ReadGroup::new(handle, String::new(), root))),
|
||||||
|
filename,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
#[pymethods]
|
#[pymethods]
|
||||||
impl PyFile {
|
impl PyFile {
|
||||||
/// Open or create an HDF5 file.
|
/// Open or create an HDF5 file.
|
||||||
///
|
///
|
||||||
/// Parameters:
|
/// Parameters:
|
||||||
/// path: file path
|
/// path: file path, or a URL (`http://`, `https://`, `s3://`, `gs://`,
|
||||||
|
/// `az://`; which schemes work depends on how the wheel was built)
|
||||||
|
/// to read the file remotely with default options (see `open_url`)
|
||||||
/// mode: 'r' for read (default), 'w' for write
|
/// mode: 'r' for read (default), 'w' for write
|
||||||
#[new]
|
#[new]
|
||||||
#[pyo3(signature = (path, mode="r"))]
|
#[pyo3(signature = (path, mode="r"))]
|
||||||
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
|
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
|
||||||
let filename = path.to_string();
|
let filename = path.to_string();
|
||||||
match mode {
|
if is_url(path) {
|
||||||
"r" => {
|
if mode != "r" {
|
||||||
let file = py.detach(|| {
|
return Err(PyValueError::new_err(format!(
|
||||||
crate::no_panic(|| clawhdf5_rs::File::open(path).map_err(to_py_err))
|
"remote files are read-only: mode '{mode}' is not supported for a URL"
|
||||||
})?;
|
)));
|
||||||
Ok(Self {
|
|
||||||
inner: Some(FileInner::Read(root_group(Arc::new(file)))),
|
|
||||||
filename,
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
let handle = Handle::open_url(py, path, &clawhdf5_remote::Options::default())?;
|
||||||
|
return Ok(Self::from_handle(handle, filename));
|
||||||
|
}
|
||||||
|
match mode {
|
||||||
|
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
|
||||||
|
"r+" => Ok(Self::from_handle(
|
||||||
|
Handle::open_editable(py, path)?,
|
||||||
|
filename,
|
||||||
|
)),
|
||||||
|
"a" if std::path::Path::new(path).exists() => Ok(Self::from_handle(
|
||||||
|
Handle::open_editable(py, path)?,
|
||||||
|
filename,
|
||||||
|
)),
|
||||||
|
"a" => Err(PyNotImplementedError::new_err(format!(
|
||||||
|
"mode 'a' on {path}, which does not exist: clawhdf5 can only edit an existing \
|
||||||
|
file in place; create a new one with mode 'w'"
|
||||||
|
))),
|
||||||
"w" => Ok(Self {
|
"w" => Ok(Self {
|
||||||
filename,
|
filename,
|
||||||
inner: Some(FileInner::Write(WriteState {
|
inner: Some(FileInner::Write(WriteState {
|
||||||
path: PathBuf::from(path),
|
// Absolute now: the file is written at close, possibly
|
||||||
|
// after the working directory changed.
|
||||||
|
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
|
||||||
root_datasets: Vec::new(),
|
root_datasets: Vec::new(),
|
||||||
root_attrs: Arc::new(Mutex::new(Vec::new())),
|
root_attrs: Arc::new(Mutex::new(Vec::new())),
|
||||||
groups: Vec::new(),
|
groups: Vec::new(),
|
||||||
})),
|
})),
|
||||||
}),
|
}),
|
||||||
other => Err(PyErr::new::<pyo3::exceptions::PyValueError, _>(format!(
|
other => Err(PyValueError::new_err(format!(
|
||||||
"unsupported mode '{other}'; expected 'r' or 'w'"
|
"unsupported mode '{other}'; expected 'r', 'r+', 'a' or 'w'"
|
||||||
))),
|
))),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Open a remote file for reading, with options.
|
||||||
|
///
|
||||||
|
/// The file is read through a block cache with range requests: opening
|
||||||
|
/// costs one request (it also fetches the first block), and a read
|
||||||
|
/// fetches only the blocks it needs. The GIL is released while waiting
|
||||||
|
/// on the network.
|
||||||
|
///
|
||||||
|
/// Parameters (all optional):
|
||||||
|
/// block_size: bytes per cached block (default 1 MiB)
|
||||||
|
/// cache_size: byte budget of the block cache (default 64 MiB)
|
||||||
|
/// headers: dict of extra HTTP headers (e.g. Authorization), sent only
|
||||||
|
/// to the URL's own origin
|
||||||
|
/// retries: retries of a request that failed transiently (default 3)
|
||||||
|
/// timeout: seconds to connect and receive response headers (default 30)
|
||||||
|
/// allow_full_download: when the server ignores Range requests,
|
||||||
|
/// download the whole file once instead of failing (default False)
|
||||||
|
/// max_full_download: largest file such a download may fetch
|
||||||
|
/// (default 1 GiB)
|
||||||
|
/// require_validator: refuse a server that sends neither ETag nor
|
||||||
|
/// Last-Modified (default False)
|
||||||
|
/// max_redirects: redirects followed per request (default 5)
|
||||||
|
/// max_parallel: requests of one read in flight at once (default 8)
|
||||||
|
#[staticmethod]
|
||||||
|
#[allow(clippy::too_many_arguments)]
|
||||||
|
#[pyo3(signature = (url, *, block_size=None, cache_size=None, headers=None, retries=None,
|
||||||
|
timeout=None, allow_full_download=None, max_full_download=None,
|
||||||
|
require_validator=None, max_redirects=None, max_parallel=None))]
|
||||||
|
fn open_url(
|
||||||
|
py: Python<'_>,
|
||||||
|
url: &str,
|
||||||
|
block_size: Option<u64>,
|
||||||
|
cache_size: Option<u64>,
|
||||||
|
headers: Option<HashMap<String, String>>,
|
||||||
|
retries: Option<u32>,
|
||||||
|
timeout: Option<f64>,
|
||||||
|
allow_full_download: Option<bool>,
|
||||||
|
max_full_download: Option<u64>,
|
||||||
|
require_validator: Option<bool>,
|
||||||
|
max_redirects: Option<u32>,
|
||||||
|
max_parallel: Option<usize>,
|
||||||
|
) -> PyResult<Self> {
|
||||||
|
let mut options = clawhdf5_remote::Options::default();
|
||||||
|
if let Some(b) = block_size {
|
||||||
|
if b == 0 {
|
||||||
|
return Err(PyValueError::new_err("block_size must be positive"));
|
||||||
|
}
|
||||||
|
options.cache.block_size = b;
|
||||||
|
options.cache.coalesce_gap = b;
|
||||||
|
// The opening request fetches the first block, not 1 MiB.
|
||||||
|
options.http.first_request = b;
|
||||||
|
}
|
||||||
|
if let Some(c) = cache_size {
|
||||||
|
options.cache.capacity = c;
|
||||||
|
}
|
||||||
|
let http = &mut options.http;
|
||||||
|
if let Some(h) = headers {
|
||||||
|
http.headers = h.into_iter().collect();
|
||||||
|
}
|
||||||
|
if let Some(r) = retries {
|
||||||
|
http.retries = r;
|
||||||
|
}
|
||||||
|
if let Some(t) = timeout {
|
||||||
|
if !(t.is_finite() && t > 0.0) {
|
||||||
|
return Err(PyValueError::new_err("timeout must be a positive number"));
|
||||||
|
}
|
||||||
|
http.timeout = Duration::from_secs_f64(t);
|
||||||
|
}
|
||||||
|
if let Some(a) = allow_full_download {
|
||||||
|
http.allow_full_download = a;
|
||||||
|
}
|
||||||
|
if let Some(m) = max_full_download {
|
||||||
|
http.max_full_download = m;
|
||||||
|
}
|
||||||
|
if let Some(v) = require_validator {
|
||||||
|
http.require_validator = v;
|
||||||
|
}
|
||||||
|
if let Some(r) = max_redirects {
|
||||||
|
http.max_redirects = r;
|
||||||
|
}
|
||||||
|
if let Some(p) = max_parallel {
|
||||||
|
if p == 0 {
|
||||||
|
return Err(PyValueError::new_err("max_parallel must be positive"));
|
||||||
|
}
|
||||||
|
http.max_parallel = p;
|
||||||
|
}
|
||||||
|
let handle = Handle::open_url(py, url, &options)?;
|
||||||
|
Ok(Self::from_handle(handle, url.to_string()))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// For a remote file, what its block cache has done so far (reads,
|
||||||
|
/// hits, misses, requests, bytes fetched, ...); `None` for a local file.
|
||||||
|
#[getter]
|
||||||
|
fn remote_stats<'py>(&self, py: Python<'py>) -> PyResult<Option<Bound<'py, PyDict>>> {
|
||||||
|
let Some(storage) = self.read_file()?.handle.remote_storage() else {
|
||||||
|
return Ok(None);
|
||||||
|
};
|
||||||
|
let s = storage.stats();
|
||||||
|
let d = PyDict::new(py);
|
||||||
|
d.set_item("reads", s.reads)?;
|
||||||
|
d.set_item("hits", s.hits)?;
|
||||||
|
d.set_item("misses", s.misses)?;
|
||||||
|
d.set_item("waits", s.waits)?;
|
||||||
|
d.set_item("requests", s.requests)?;
|
||||||
|
d.set_item("fetch_calls", s.fetch_calls)?;
|
||||||
|
d.set_item("bytes_fetched", s.bytes_fetched)?;
|
||||||
|
d.set_item("evictions", s.evictions)?;
|
||||||
|
d.set_item("cached_bytes", s.cached_bytes)?;
|
||||||
|
Ok(Some(d))
|
||||||
|
}
|
||||||
|
|
||||||
/// Close the file. In write mode, this finalizes and writes the file.
|
/// Close the file. In write mode, this finalizes and writes the file.
|
||||||
fn close(&mut self) -> PyResult<()> {
|
fn close(&mut self) -> PyResult<()> {
|
||||||
let inner = self.inner.take().ok_or_else(|| {
|
let inner = self.inner.take().ok_or_else(|| {
|
||||||
PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed")
|
PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed")
|
||||||
})?;
|
})?;
|
||||||
match inner {
|
match inner {
|
||||||
FileInner::Read(_) => Ok(()),
|
FileInner::Read(root) => {
|
||||||
|
root.handle.close();
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
FileInner::Write(state) => finalize_write(state),
|
FileInner::Write(state) => finalize_write(state),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -121,7 +278,7 @@ impl PyFile {
|
|||||||
|
|
||||||
/// List the names of all children in the root group.
|
/// List the names of all children in the root group.
|
||||||
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
let names = self.read_file()?.member_names()?;
|
let names = self.read_file()?.member_names(py)?;
|
||||||
Ok(PyList::new(py, names)?.into_any().unbind())
|
Ok(PyList::new(py, names)?.into_any().unbind())
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -139,8 +296,8 @@ impl PyFile {
|
|||||||
self.keys(py)?.call_method0(py, "__iter__")
|
self.keys(py)?.call_method0(py, "__iter__")
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __len__(&self) -> PyResult<usize> {
|
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
Ok(self.read_file()?.member_names()?.len())
|
Ok(self.read_file()?.member_names(py)?.len())
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The root group's name, `/`.
|
/// The root group's name, `/`.
|
||||||
@@ -149,7 +306,31 @@ impl PyFile {
|
|||||||
"/"
|
"/"
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The path the file was opened with.
|
/// `'r'` for a file opened read-only (a local file or a URL), `'r+'`
|
||||||
|
/// for one open for editing or writing, as h5py reports it.
|
||||||
|
#[getter]
|
||||||
|
fn mode(&self) -> PyResult<&'static str> {
|
||||||
|
match &self.inner {
|
||||||
|
Some(FileInner::Read(root)) if !root.handle.is_writable() => Ok("r"),
|
||||||
|
Some(_) => Ok("r+"),
|
||||||
|
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||||
|
"file is closed",
|
||||||
|
)),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Nothing to do: every edit is written and synced when it is made, and
|
||||||
|
/// a file opened with 'w' is written on `close()`.
|
||||||
|
fn flush(&self) {}
|
||||||
|
|
||||||
|
/// Deleting objects is not supported (h5py's `del f[name]`).
|
||||||
|
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||||
|
Err(PyNotImplementedError::new_err(format!(
|
||||||
|
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
|
||||||
|
)))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The path (or URL) the file was opened with.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn filename(&self) -> &str {
|
fn filename(&self) -> &str {
|
||||||
&self.filename
|
&self.filename
|
||||||
@@ -204,9 +385,9 @@ impl PyFile {
|
|||||||
/// Attribute access. In read mode, returns attributes of the root group.
|
/// Attribute access. In read mode, returns attributes of the root group.
|
||||||
/// In write mode, returns a writable attrs handle.
|
/// In write mode, returns a writable attrs handle.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn attrs(&self) -> PyResult<PyAttrs> {
|
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||||
match self.inner.as_ref() {
|
match self.inner.as_ref() {
|
||||||
Some(FileInner::Read(root)) => root.attrs(),
|
Some(FileInner::Read(root)) => root.attrs(py),
|
||||||
Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))),
|
Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))),
|
||||||
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||||
"file is closed",
|
"file is closed",
|
||||||
@@ -216,9 +397,10 @@ impl PyFile {
|
|||||||
|
|
||||||
fn __repr__(&self) -> String {
|
fn __repr__(&self) -> String {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
Some(FileInner::Read(root)) => {
|
Some(FileInner::Read(root)) => match root.handle.redacted_url() {
|
||||||
format!("<HDF5 File (read, {} bytes)>", root.file.as_bytes().len())
|
Some(url) => format!("<HDF5 File (read, \"{url}\")>"),
|
||||||
}
|
None => format!("<HDF5 File (read, \"{}\")>", self.filename),
|
||||||
|
},
|
||||||
Some(FileInner::Write(s)) => {
|
Some(FileInner::Write(s)) => {
|
||||||
format!("<HDF5 File (write, \"{}\")>", s.path.display())
|
format!("<HDF5 File (write, \"{}\")>", s.path.display())
|
||||||
}
|
}
|
||||||
@@ -226,8 +408,8 @@ impl PyFile {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __contains__(&self, key: &str) -> PyResult<bool> {
|
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||||
Ok(self.read_file()?.contains(key))
|
self.read_file()?.contains(py, key)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -248,6 +430,13 @@ impl PyFile {
|
|||||||
fn write_state_mut(&mut self) -> PyResult<&mut WriteState> {
|
fn write_state_mut(&mut self) -> PyResult<&mut WriteState> {
|
||||||
match &mut self.inner {
|
match &mut self.inner {
|
||||||
Some(FileInner::Write(s)) => Ok(s),
|
Some(FileInner::Write(s)) => Ok(s),
|
||||||
|
Some(FileInner::Read(root)) if root.handle.is_writable() => {
|
||||||
|
Err(PyNotImplementedError::new_err(
|
||||||
|
"creating datasets or groups in an existing file is not supported by \
|
||||||
|
clawhdf5's in-place editor (mode 'r+' changes values, shapes and \
|
||||||
|
attributes)",
|
||||||
|
))
|
||||||
|
}
|
||||||
Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||||
"cannot write to a file opened for reading",
|
"cannot write to a file opened for reading",
|
||||||
)),
|
)),
|
||||||
@@ -271,11 +460,6 @@ fn parse_compression(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn root_group(file: Arc<clawhdf5_rs::File>) -> ReadGroup {
|
|
||||||
let root = file.superblock().root_group_address;
|
|
||||||
ReadGroup::new(file, String::new(), root)
|
|
||||||
}
|
|
||||||
|
|
||||||
/// Build and write the HDF5 file from accumulated write state.
|
/// Build and write the HDF5 file from accumulated write state.
|
||||||
fn finalize_write(state: WriteState) -> PyResult<()> {
|
fn finalize_write(state: WriteState) -> PyResult<()> {
|
||||||
crate::no_panic(|| {
|
crate::no_panic(|| {
|
||||||
@@ -309,6 +493,18 @@ fn finalize_write(state: WriteState) -> PyResult<()> {
|
|||||||
mod tests {
|
mod tests {
|
||||||
use super::*;
|
use super::*;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn urls_and_paths() {
|
||||||
|
assert!(is_url("http://h/f.h5"));
|
||||||
|
assert!(is_url("s3://bucket/key.h5"));
|
||||||
|
assert!(is_url("git+https://x"));
|
||||||
|
assert!(!is_url("data.h5"));
|
||||||
|
assert!(!is_url("/tmp/a://b.h5"));
|
||||||
|
assert!(!is_url("dir/x://y"));
|
||||||
|
assert!(!is_url("1http://x"));
|
||||||
|
assert!(!is_url("://x"));
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn parse_gzip_compression() {
|
fn parse_gzip_compression() {
|
||||||
assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6));
|
assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6));
|
||||||
|
|||||||
@@ -3,11 +3,12 @@
|
|||||||
use std::collections::HashMap;
|
use std::collections::HashMap;
|
||||||
use std::sync::{Arc, Mutex, OnceLock};
|
use std::sync::{Arc, Mutex, OnceLock};
|
||||||
|
|
||||||
use pyo3::exceptions::{PyIOError, PyKeyError, PyValueError};
|
use pyo3::exceptions::{PyIOError, PyKeyError, PyNotImplementedError, PyOSError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::PyList;
|
use pyo3::types::PyList;
|
||||||
|
|
||||||
use crate::attrs::PyAttrs;
|
use crate::attrs::PyAttrs;
|
||||||
|
use crate::handle::Handle;
|
||||||
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
|
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
|
||||||
|
|
||||||
/// Shared state for a group being written.
|
/// Shared state for a group being written.
|
||||||
@@ -34,9 +35,9 @@ enum GroupInner {
|
|||||||
}
|
}
|
||||||
|
|
||||||
impl PyGroup {
|
impl PyGroup {
|
||||||
pub(crate) fn from_read(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
|
pub(crate) fn from_read(handle: Arc<Handle>, path: String, addr: u64) -> Self {
|
||||||
Self {
|
Self {
|
||||||
inner: GroupInner::Read(ReadGroup::new(file, path, addr)),
|
inner: GroupInner::Read(ReadGroup::new(handle, path, addr)),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -60,9 +61,10 @@ impl PyGroup {
|
|||||||
/// h5py). It keeps its own address and, once listed, its links, so looking
|
/// h5py). It keeps its own address and, once listed, its links, so looking
|
||||||
/// up a child neither resolves the path from the root nor scans the group's
|
/// up a child neither resolves the path from the root nor scans the group's
|
||||||
/// links again: visiting every member of a large group is linear, not
|
/// links again: visiting every member of a large group is linear, not
|
||||||
/// quadratic.
|
/// quadratic. (Edits never add or remove links, so these stay valid in a
|
||||||
|
/// file open for editing.)
|
||||||
pub(crate) struct ReadGroup {
|
pub(crate) struct ReadGroup {
|
||||||
pub file: Arc<clawhdf5_rs::File>,
|
pub handle: Arc<Handle>,
|
||||||
pub path: String,
|
pub path: String,
|
||||||
pub addr: u64,
|
pub addr: u64,
|
||||||
/// Link name -> object address (soft links resolved), filled on first use.
|
/// Link name -> object address (soft links resolved), filled on first use.
|
||||||
@@ -72,9 +74,9 @@ pub(crate) struct ReadGroup {
|
|||||||
}
|
}
|
||||||
|
|
||||||
impl ReadGroup {
|
impl ReadGroup {
|
||||||
pub(crate) fn new(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
|
pub(crate) fn new(handle: Arc<Handle>, path: String, addr: u64) -> Self {
|
||||||
Self {
|
Self {
|
||||||
file,
|
handle,
|
||||||
path,
|
path,
|
||||||
addr,
|
addr,
|
||||||
links: OnceLock::new(),
|
links: OnceLock::new(),
|
||||||
@@ -82,17 +84,14 @@ impl ReadGroup {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn links(&self) -> PyResult<&HashMap<String, u64>> {
|
fn links(&self, py: Python<'_>) -> PyResult<&HashMap<String, u64>> {
|
||||||
if let Some(links) = self.links.get() {
|
if let Some(links) = self.links.get() {
|
||||||
return Ok(links);
|
return Ok(links);
|
||||||
}
|
}
|
||||||
let entries = crate::no_panic(|| {
|
let (addr, path) = (self.addr, &self.path);
|
||||||
clawhdf5_format::group_v2::resolve_group_children(
|
let entries = self.handle.with(py, |f| {
|
||||||
self.file.as_bytes(),
|
clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr)
|
||||||
self.file.superblock(),
|
.map_err(|e| node::format_err(path, e, PyValueError::new_err))
|
||||||
self.addr,
|
|
||||||
)
|
|
||||||
.map_err(|e| PyValueError::new_err(format!("{}: {e}", node::name(&self.path))))
|
|
||||||
})?;
|
})?;
|
||||||
let map = entries
|
let map = entries
|
||||||
.into_iter()
|
.into_iter()
|
||||||
@@ -102,7 +101,7 @@ impl ReadGroup {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// The path and address of `key` (a name, a relative or an absolute path).
|
/// The path and address of `key` (a name, a relative or an absolute path).
|
||||||
fn locate(&self, key: &str) -> PyResult<(String, u64)> {
|
fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> {
|
||||||
let path = node::join(&self.path, key);
|
let path = node::join(&self.path, key);
|
||||||
let rel = if self.path.is_empty() {
|
let rel = if self.path.is_empty() {
|
||||||
Some(path.as_str())
|
Some(path.as_str())
|
||||||
@@ -112,24 +111,24 @@ impl ReadGroup {
|
|||||||
path.strip_prefix(self.path.as_str())
|
path.strip_prefix(self.path.as_str())
|
||||||
.and_then(|r| r.strip_prefix('/'))
|
.and_then(|r| r.strip_prefix('/'))
|
||||||
};
|
};
|
||||||
let addr = match rel {
|
// A direct child: the link table, when it has the name.
|
||||||
// A direct child: the link table, when it has the name.
|
if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/'))
|
||||||
Some(name) if !name.is_empty() && !name.contains('/') => {
|
&& let Some(&a) = self.links(py)?.get(name)
|
||||||
match self.links()?.get(name) {
|
{
|
||||||
Some(&a) => a,
|
return Ok((path, a));
|
||||||
None => node::resolve_from(&self.file, self.addr, name, &path)?,
|
}
|
||||||
}
|
let addr = self.addr;
|
||||||
}
|
let found = self.handle.with(py, |f| match rel {
|
||||||
Some(rel) => node::resolve_from(&self.file, self.addr, rel, &path)?,
|
Some(rel) => node::resolve_from(f, addr, rel, &path),
|
||||||
None => node::address(&self.file, &path)?,
|
None => node::address(f, &path),
|
||||||
};
|
})?;
|
||||||
Ok((path, addr))
|
Ok((path, found))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// `group[key]`.
|
/// `group[key]`.
|
||||||
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
||||||
let (path, addr) = self.locate(key)?;
|
let (path, addr) = self.locate(py, key)?;
|
||||||
node::open(py, &self.file, path, addr)
|
node::open(py, &self.handle, path, addr)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// `group.get(key, default)`.
|
/// `group.get(key, default)`.
|
||||||
@@ -148,48 +147,57 @@ impl ReadGroup {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// Names of the group's datasets and subgroups, sorted (h5py's order).
|
/// Names of the group's datasets and subgroups, sorted (h5py's order).
|
||||||
pub(crate) fn member_names(&self) -> PyResult<&[String]> {
|
pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> {
|
||||||
if let Some(m) = self.members.get() {
|
if let Some(m) = self.members.get() {
|
||||||
return Ok(m);
|
return Ok(m);
|
||||||
}
|
}
|
||||||
let mut names = Vec::new();
|
let links = self.links(py)?;
|
||||||
for (name, &addr) in self.links()? {
|
let path = &self.path;
|
||||||
let hdr = node::header_at(&self.file, addr, &node::join(&self.path, name))?;
|
let mut names = self.handle.with(py, |f| {
|
||||||
if matches!(
|
let mut names = Vec::new();
|
||||||
node::kind(&hdr),
|
for (name, &addr) in links {
|
||||||
Some(node::Kind::Dataset | node::Kind::Group)
|
if matches!(
|
||||||
) {
|
node::kind_at(f, addr, &node::join(path, name))?,
|
||||||
names.push(name.clone());
|
Some(node::Kind::Dataset | node::Kind::Group)
|
||||||
|
) {
|
||||||
|
names.push(name.clone());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
Ok(names)
|
||||||
|
})?;
|
||||||
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
|
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
|
||||||
Ok(self.members.get_or_init(|| names))
|
Ok(self.members.get_or_init(|| names))
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(crate) fn contains(&self, key: &str) -> bool {
|
/// `key in group`: whether `key` names a dataset or group. A failed
|
||||||
self.locate(key)
|
/// read of the file (a network error) is raised, not `False`.
|
||||||
.and_then(|(path, addr)| node::header_at(&self.file, addr, &path))
|
pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||||
.ok()
|
let found = self
|
||||||
.and_then(|h| node::kind(&h))
|
.locate(py, key)
|
||||||
.is_some_and(|k| k != node::Kind::Datatype)
|
.and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path)));
|
||||||
|
match found {
|
||||||
|
Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)),
|
||||||
|
Err(e) if e.is_instance_of::<PyOSError>(py) => Err(e),
|
||||||
|
Err(_) => Ok(false),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
|
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
|
||||||
self.member_names()?
|
self.member_names(py)?
|
||||||
.iter()
|
.iter()
|
||||||
.map(|n| self.get_item(py, n))
|
.map(|n| self.get_item(py, n))
|
||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
|
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
|
||||||
self.member_names()?
|
self.member_names(py)?
|
||||||
.iter()
|
.iter()
|
||||||
.map(|n| Ok((n.clone(), self.get_item(py, n)?)))
|
.map(|n| Ok((n.clone(), self.get_item(py, n)?)))
|
||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(crate) fn attrs(&self) -> PyResult<PyAttrs> {
|
pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||||
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
|
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -210,7 +218,7 @@ impl PyGroup {
|
|||||||
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
GroupInner::Read(g) => {
|
GroupInner::Read(g) => {
|
||||||
let list = PyList::new(py, g.member_names()?)?;
|
let list = PyList::new(py, g.member_names(py)?)?;
|
||||||
Ok(list.into_any().unbind())
|
Ok(list.into_any().unbind())
|
||||||
}
|
}
|
||||||
GroupInner::Write(state) => {
|
GroupInner::Write(state) => {
|
||||||
@@ -236,9 +244,9 @@ impl PyGroup {
|
|||||||
self.keys(py)?.call_method0(py, "__iter__")
|
self.keys(py)?.call_method0(py, "__iter__")
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __len__(&self) -> PyResult<usize> {
|
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
GroupInner::Read(g) => Ok(g.member_names()?.len()),
|
GroupInner::Read(g) => Ok(g.member_names(py)?.len()),
|
||||||
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
|
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -293,17 +301,28 @@ impl PyGroup {
|
|||||||
state.lock().unwrap().datasets.push(spec);
|
state.lock().unwrap().datasets.push(spec);
|
||||||
Ok(())
|
Ok(())
|
||||||
}
|
}
|
||||||
GroupInner::Read { .. } => Err(PyIOError::new_err(
|
GroupInner::Read(g) if g.handle.is_writable() => Err(PyNotImplementedError::new_err(
|
||||||
|
"creating datasets or groups in an existing file is not supported by \
|
||||||
|
clawhdf5's in-place editor (mode 'r+' changes values, shapes and attributes)",
|
||||||
|
)),
|
||||||
|
GroupInner::Read(_) => Err(PyIOError::new_err(
|
||||||
"cannot create datasets on a read-only group",
|
"cannot create datasets on a read-only group",
|
||||||
)),
|
)),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Deleting objects is not supported (h5py's `del group[name]`).
|
||||||
|
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||||
|
Err(PyNotImplementedError::new_err(format!(
|
||||||
|
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
|
||||||
|
)))
|
||||||
|
}
|
||||||
|
|
||||||
/// Attribute access.
|
/// Attribute access.
|
||||||
#[getter]
|
#[getter]
|
||||||
fn attrs(&self) -> PyResult<PyAttrs> {
|
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
GroupInner::Read(g) => g.attrs(),
|
GroupInner::Read(g) => g.attrs(py),
|
||||||
GroupInner::Write(state) => {
|
GroupInner::Write(state) => {
|
||||||
let store = Arc::clone(&state.lock().unwrap().attrs);
|
let store = Arc::clone(&state.lock().unwrap().attrs);
|
||||||
Ok(PyAttrs::from_write(store))
|
Ok(PyAttrs::from_write(store))
|
||||||
@@ -311,10 +330,10 @@ impl PyGroup {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __repr__(&self) -> String {
|
fn __repr__(&self, py: Python<'_>) -> String {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
GroupInner::Read(g) => {
|
GroupInner::Read(g) => {
|
||||||
let n = g.member_names().map_or(0, |m| m.len());
|
let n = g.member_names(py).map_or(0, |m| m.len());
|
||||||
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
|
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
|
||||||
}
|
}
|
||||||
GroupInner::Write(state) => {
|
GroupInner::Write(state) => {
|
||||||
@@ -324,9 +343,9 @@ impl PyGroup {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn __contains__(&self, key: &str) -> PyResult<bool> {
|
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||||
match &self.inner {
|
match &self.inner {
|
||||||
GroupInner::Read(g) => Ok(g.contains(key)),
|
GroupInner::Read(g) => g.contains(py, key),
|
||||||
GroupInner::Write(state) => {
|
GroupInner::Write(state) => {
|
||||||
let guard = state.lock().unwrap();
|
let guard = state.lock().unwrap();
|
||||||
Ok(guard.datasets.iter().any(|d| d.name == key))
|
Ok(guard.datasets.iter().any(|d| d.name == key))
|
||||||
@@ -357,31 +376,6 @@ pub(crate) fn finalize_write_group(
|
|||||||
mod tests {
|
mod tests {
|
||||||
use super::*;
|
use super::*;
|
||||||
|
|
||||||
#[test]
|
|
||||||
fn member_names_are_sorted() {
|
|
||||||
let mut b = clawhdf5_rs::FileBuilder::new();
|
|
||||||
b.create_dataset("zeta").with_f64_data(&[1.0]);
|
|
||||||
b.create_dataset("alpha").with_f64_data(&[1.0]);
|
|
||||||
let mut g = b.create_group("mid");
|
|
||||||
g.create_dataset("x").with_f64_data(&[1.0]);
|
|
||||||
let finished = g.finish();
|
|
||||||
b.add_group(finished);
|
|
||||||
let bytes = b.finish().unwrap();
|
|
||||||
let file = Arc::new(clawhdf5_rs::File::from_bytes(bytes).unwrap());
|
|
||||||
let root = file.superblock().root_group_address;
|
|
||||||
let top = ReadGroup::new(Arc::clone(&file), String::new(), root);
|
|
||||||
assert_eq!(top.member_names().unwrap(), ["alpha", "mid", "zeta"]);
|
|
||||||
let (path, addr) = top.locate("mid").unwrap();
|
|
||||||
assert_eq!(path, "mid");
|
|
||||||
let mid = ReadGroup::new(Arc::clone(&file), path, addr);
|
|
||||||
assert_eq!(mid.member_names().unwrap(), ["x"]);
|
|
||||||
assert!(top.contains("mid/x"));
|
|
||||||
assert!(mid.contains("/alpha"));
|
|
||||||
assert!(mid.contains("x") && mid.contains("./x"));
|
|
||||||
assert!(!top.contains("nope"));
|
|
||||||
assert!(!mid.contains("alpha"));
|
|
||||||
}
|
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn finalize_group() {
|
fn finalize_group() {
|
||||||
let state = WriteGroupState {
|
let state = WriteGroupState {
|
||||||
|
|||||||
@@ -0,0 +1,230 @@
|
|||||||
|
//! The open file every object of a `File` shares.
|
||||||
|
//!
|
||||||
|
//! Every read goes through [`Handle::with`], which releases the GIL and
|
||||||
|
//! parses through `File::storage()`, so the same code serves a local file
|
||||||
|
//! (memory-mapped), a remote one (`clawhdf5-remote`: range requests through
|
||||||
|
//! a block cache, so a network read never holds the GIL) and a file open
|
||||||
|
//! for editing.
|
||||||
|
//!
|
||||||
|
//! A file opened with `'r+'` also holds a [`FileEditor`]. An edit takes the
|
||||||
|
//! file's write lock, so no read runs while the file changes underneath it,
|
||||||
|
//! and reopens the file afterwards, through the editor's own open file
|
||||||
|
//! rather than its path: reads after an edit see the new bytes
|
||||||
|
//! (a grown file, a new dataspace), never a stale mapping or chunk cache.
|
||||||
|
//! Objects that cache something an edit can change compare
|
||||||
|
//! [`Handle::generation`] with the value they cached it at.
|
||||||
|
//!
|
||||||
|
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
|
||||||
|
//! with the GIL released, and code that holds it never touches Python.
|
||||||
|
|
||||||
|
use std::sync::atomic::{AtomicU64, Ordering};
|
||||||
|
use std::sync::{Arc, Mutex, PoisonError, RwLock};
|
||||||
|
|
||||||
|
use clawhdf5_rs::{File, FileEditor};
|
||||||
|
use pyo3::exceptions::PyOSError;
|
||||||
|
use pyo3::prelude::*;
|
||||||
|
|
||||||
|
use crate::{panic_text, to_py_err};
|
||||||
|
|
||||||
|
/// Where the file's bytes come from.
|
||||||
|
pub(crate) enum Source {
|
||||||
|
/// A local file (memory-mapped). Its path is not kept: nothing reopens
|
||||||
|
/// it by path (see `open_editable`).
|
||||||
|
Local,
|
||||||
|
/// A URL, read through `clawhdf5-remote`'s block cache.
|
||||||
|
Remote {
|
||||||
|
url: String,
|
||||||
|
storage: Arc<clawhdf5_remote::RemoteStorage>,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
pub(crate) struct Handle {
|
||||||
|
/// The file as last opened; `None` if reopening it after an edit failed
|
||||||
|
/// (every read is then an error rather than a read of stale bytes).
|
||||||
|
file: RwLock<Option<File>>,
|
||||||
|
/// For `'r+'`: the editor, until the file is closed.
|
||||||
|
editor: Option<Mutex<Option<FileEditor>>>,
|
||||||
|
source: Source,
|
||||||
|
/// Bumped by every edit.
|
||||||
|
generation: AtomicU64,
|
||||||
|
pub offset_size: u8,
|
||||||
|
pub length_size: u8,
|
||||||
|
pub root: u64,
|
||||||
|
}
|
||||||
|
|
||||||
|
fn closed_after_failed_reopen() -> PyErr {
|
||||||
|
PyOSError::new_err("the file could not be reopened after an edit; open it again")
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Handle {
|
||||||
|
fn new(file: File, source: Source, editor: Option<FileEditor>) -> Arc<Self> {
|
||||||
|
let sb = file.superblock();
|
||||||
|
let (offset_size, length_size, root) =
|
||||||
|
(sb.offset_size, sb.length_size, sb.root_group_address);
|
||||||
|
Arc::new(Self {
|
||||||
|
file: RwLock::new(Some(file)),
|
||||||
|
editor: editor.map(|e| Mutex::new(Some(e))),
|
||||||
|
source,
|
||||||
|
generation: AtomicU64::new(0),
|
||||||
|
offset_size,
|
||||||
|
length_size,
|
||||||
|
root,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A local file, read-only.
|
||||||
|
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
|
||||||
|
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
|
||||||
|
Ok(Self::new(file, Source::Local, None))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A local file, open for in-place editing (`'r+'`): the editor takes
|
||||||
|
/// the file's exclusive lock and checks that it can edit the file, then
|
||||||
|
/// the file is read through the editor's own file — never by path
|
||||||
|
/// again, so a later `os.chdir` or a rename or replacement of the path
|
||||||
|
/// cannot make reads (or the editor's plans) come from another file.
|
||||||
|
pub(crate) fn open_editable(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
|
||||||
|
let (file, editor) = py.detach(|| {
|
||||||
|
crate::no_panic(|| {
|
||||||
|
let editor = FileEditor::open(path).map_err(to_py_err)?;
|
||||||
|
let file = editor.reader().map_err(to_py_err)?;
|
||||||
|
Ok((file, editor))
|
||||||
|
})
|
||||||
|
})?;
|
||||||
|
Ok(Self::new(file, Source::Local, Some(editor)))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A remote file (`http(s)://`, `s3://`, ...).
|
||||||
|
pub(crate) fn open_url(
|
||||||
|
py: Python<'_>,
|
||||||
|
url: &str,
|
||||||
|
options: &clawhdf5_remote::Options,
|
||||||
|
) -> PyResult<Arc<Self>> {
|
||||||
|
let (file, storage) = py.detach(|| {
|
||||||
|
crate::no_panic(|| {
|
||||||
|
let storage = clawhdf5_remote::storage_for_url(url, options).map_err(remote_err)?;
|
||||||
|
let file = File::open_storage(storage.clone()).map_err(to_py_err)?;
|
||||||
|
Ok((file, storage))
|
||||||
|
})
|
||||||
|
})?;
|
||||||
|
Ok(Self::new(
|
||||||
|
file,
|
||||||
|
Source::Remote {
|
||||||
|
url: url.to_string(),
|
||||||
|
storage,
|
||||||
|
},
|
||||||
|
None,
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Run `f` on the file with the GIL released (a remote read may wait
|
||||||
|
/// on the network; other Python threads run meanwhile). `f` must not
|
||||||
|
/// touch Python.
|
||||||
|
pub(crate) fn with<R: Send>(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
f: impl FnOnce(&File) -> PyResult<R> + Send,
|
||||||
|
) -> PyResult<R> {
|
||||||
|
py.detach(|| self.with_detached(f))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// [`with`](Self::with) for code that already runs without the GIL.
|
||||||
|
pub(crate) fn with_detached<R>(&self, f: impl FnOnce(&File) -> PyResult<R>) -> PyResult<R> {
|
||||||
|
crate::no_panic(|| {
|
||||||
|
let guard = self.file.read().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
let file = guard.as_ref().ok_or_else(closed_after_failed_reopen)?;
|
||||||
|
f(file)
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Edits so far: objects that cache something an edit can change (a
|
||||||
|
/// dataset's shape, an object's attributes) re-read it when this moved.
|
||||||
|
pub(crate) fn generation(&self) -> u64 {
|
||||||
|
self.generation.load(Ordering::Acquire)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Whether the file was opened for editing (`'r+'`), even if closed since.
|
||||||
|
pub(crate) fn is_writable(&self) -> bool {
|
||||||
|
self.editor.is_some()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The remote file's block cache.
|
||||||
|
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
|
||||||
|
match &self.source {
|
||||||
|
Source::Remote { storage, .. } => Some(storage),
|
||||||
|
Source::Local => None,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The URL of a remote file, credentials and query values redacted.
|
||||||
|
pub(crate) fn redacted_url(&self) -> Option<String> {
|
||||||
|
match &self.source {
|
||||||
|
Source::Remote { url, .. } => Some(clawhdf5_remote::redact_url(url)),
|
||||||
|
Source::Local => None,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Release the editor, and with it the file's lock. Objects still
|
||||||
|
/// open keep reading the file as it was last written; an edit through
|
||||||
|
/// them is an error.
|
||||||
|
pub(crate) fn close(&self) {
|
||||||
|
if let Some(ed) = &self.editor {
|
||||||
|
ed.lock().unwrap_or_else(PoisonError::into_inner).take();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Apply one edit with the GIL released. No read runs while it writes,
|
||||||
|
/// and the file is reopened afterwards — also after a failed edit, since
|
||||||
|
/// a commit that failed part-way may have changed the file.
|
||||||
|
pub(crate) fn edit<R: Send>(
|
||||||
|
&self,
|
||||||
|
py: Python<'_>,
|
||||||
|
f: impl FnOnce(&mut FileEditor) -> Result<R, clawhdf5_rs::Error> + Send,
|
||||||
|
) -> PyResult<R> {
|
||||||
|
let Some(editor) = &self.editor else {
|
||||||
|
return Err(PyOSError::new_err(match self.source {
|
||||||
|
Source::Remote { .. } => "remote files are read-only",
|
||||||
|
Source::Local => "the file is open read-only; open it with mode 'r+' to change it",
|
||||||
|
}));
|
||||||
|
};
|
||||||
|
if matches!(self.source, Source::Remote { .. }) {
|
||||||
|
return Err(PyOSError::new_err("remote files are read-only"));
|
||||||
|
}
|
||||||
|
py.detach(|| {
|
||||||
|
let mut ed = editor.lock().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
let ed = ed
|
||||||
|
.as_mut()
|
||||||
|
.ok_or_else(|| PyOSError::new_err("the file is closed"))?;
|
||||||
|
let mut file = self.file.write().unwrap_or_else(PoisonError::into_inner);
|
||||||
|
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| f(ed)));
|
||||||
|
// Drop the old mapping and its chunk cache before reopening.
|
||||||
|
*file = None;
|
||||||
|
// Through the editor's file, not the path (see `open_editable`).
|
||||||
|
let reopened = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| ed.reader()));
|
||||||
|
self.generation.fetch_add(1, Ordering::AcqRel);
|
||||||
|
match reopened {
|
||||||
|
Ok(Ok(f)) => *file = Some(f),
|
||||||
|
Ok(Err(e)) => return Err(to_py_err(e)),
|
||||||
|
Err(_) => return Err(closed_after_failed_reopen()),
|
||||||
|
}
|
||||||
|
drop(file);
|
||||||
|
match result {
|
||||||
|
Ok(r) => r.map_err(to_py_err),
|
||||||
|
Err(p) => Err(crate::InternalError::new_err(format!(
|
||||||
|
"clawhdf5 internal error (please report it): {}",
|
||||||
|
panic_text(&*p)
|
||||||
|
))),
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A `clawhdf5_remote::Error` as a Python exception: the network side
|
||||||
|
/// (unreachable, a status, no range support, a changed file) is `OSError`,
|
||||||
|
/// a file that is not HDF5 is what `to_py_err` makes of it.
|
||||||
|
pub(crate) fn remote_err(e: clawhdf5_remote::Error) -> PyErr {
|
||||||
|
match e {
|
||||||
|
clawhdf5_remote::Error::Hdf5(e) => to_py_err(e),
|
||||||
|
other => PyOSError::new_err(other.to_string()),
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -7,13 +7,22 @@
|
|||||||
//!
|
//!
|
||||||
//! with clawhdf5.File('data.h5', 'r') as f:
|
//! with clawhdf5.File('data.h5', 'r') as f:
|
||||||
//! data = f['dataset_name'][:]
|
//! data = f['dataset_name'][:]
|
||||||
|
//!
|
||||||
|
//! with clawhdf5.File('http://host/data.h5') as f: # range requests
|
||||||
|
//! block = f['dataset_name'][10:20]
|
||||||
|
//!
|
||||||
|
//! with clawhdf5.File('data.h5', 'r+') as f: # in-place edits
|
||||||
|
//! f['dataset_name'][0] = 1.5
|
||||||
|
//! f.attrs['note'] = 'edited'
|
||||||
//! ```
|
//! ```
|
||||||
|
|
||||||
mod attrs;
|
mod attrs;
|
||||||
mod convert;
|
mod convert;
|
||||||
mod dataset;
|
mod dataset;
|
||||||
|
mod edit;
|
||||||
mod file;
|
mod file;
|
||||||
mod group;
|
mod group;
|
||||||
|
mod handle;
|
||||||
mod node;
|
mod node;
|
||||||
mod select;
|
mod select;
|
||||||
|
|
||||||
@@ -63,7 +72,7 @@ fn _panic_for_test() -> PyResult<()> {
|
|||||||
/// Convert a `clawhdf5_rs::Error` into a `PyErr`.
|
/// Convert a `clawhdf5_rs::Error` into a `PyErr`.
|
||||||
///
|
///
|
||||||
/// Maps different error variants to more specific Python exception types:
|
/// Maps different error variants to more specific Python exception types:
|
||||||
/// - I/O errors -> `PyIOError`
|
/// - I/O errors, and failed reads of a remote file -> `PyIOError`/`PyOSError`
|
||||||
/// - Format/parsing errors -> `PyValueError`
|
/// - Format/parsing errors -> `PyValueError`
|
||||||
/// - Missing dataset/path errors -> `PyKeyError`
|
/// - Missing dataset/path errors -> `PyKeyError`
|
||||||
/// - Invalid arguments -> `PyValueError`
|
/// - Invalid arguments -> `PyValueError`
|
||||||
@@ -73,6 +82,10 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr {
|
|||||||
use clawhdf5_rs::Error;
|
use clawhdf5_rs::Error;
|
||||||
match &e {
|
match &e {
|
||||||
Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()),
|
Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()),
|
||||||
|
// A failed read of the storage: a network error on a remote file.
|
||||||
|
Error::Format(clawhdf5_format::error::FormatError::Storage(_)) => {
|
||||||
|
PyErr::new::<pyo3::exceptions::PyOSError, _>(e.to_string())
|
||||||
|
}
|
||||||
Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()),
|
Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()),
|
||||||
Error::NotADataset(_) | Error::MissingMessage(_) => {
|
Error::NotADataset(_) | Error::MissingMessage(_) => {
|
||||||
PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string())
|
PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string())
|
||||||
|
|||||||
@@ -1,16 +1,24 @@
|
|||||||
//! Resolving paths to objects in a file opened for reading.
|
//! Resolving paths to objects in a file opened for reading.
|
||||||
|
//!
|
||||||
|
//! Everything here parses through `File::storage()` (the `clawhdf5_format`
|
||||||
|
//! `*_in` functions), never `File::as_bytes()`, so it works the same on a
|
||||||
|
//! memory-mapped local file and on a remote one; and it runs inside
|
||||||
|
//! `Handle::with`, without the GIL.
|
||||||
|
|
||||||
use std::sync::Arc;
|
use std::sync::Arc;
|
||||||
|
|
||||||
use clawhdf5_format::attribute::AttributeMessage;
|
use clawhdf5_format::attribute::AttributeMessage;
|
||||||
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
|
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
|
||||||
|
use clawhdf5_format::error::FormatError;
|
||||||
use clawhdf5_format::message_type::MessageType;
|
use clawhdf5_format::message_type::MessageType;
|
||||||
use clawhdf5_format::object_header::ObjectHeader;
|
use clawhdf5_format::object_header::ObjectHeader;
|
||||||
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
|
use clawhdf5_rs::File;
|
||||||
|
use pyo3::exceptions::{PyKeyError, PyOSError, PyTypeError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
|
|
||||||
use crate::dataset::PyDataset;
|
use crate::dataset::{DatasetMeta, PyDataset};
|
||||||
use crate::group::PyGroup;
|
use crate::group::PyGroup;
|
||||||
|
use crate::handle::Handle;
|
||||||
|
|
||||||
/// Join `key` onto the group path `base` the way h5py does: an absolute key
|
/// Join `key` onto the group path `base` the way h5py does: an absolute key
|
||||||
/// starts from the root, a relative one from `base`. Paths are kept without
|
/// starts from the root, a relative one from `base`. Paths are kept without
|
||||||
@@ -33,42 +41,46 @@ pub(crate) fn name(path: &str) -> String {
|
|||||||
format!("/{path}")
|
format!("/{path}")
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// A format error met at `path`: a failed read of the storage (a network
|
||||||
|
/// error on a remote file) is an `OSError`, anything else `other(message)`.
|
||||||
|
pub(crate) fn format_err(path: &str, e: FormatError, other: fn(String) -> PyErr) -> PyErr {
|
||||||
|
let msg = format!("{}: {e}", name(path));
|
||||||
|
match e {
|
||||||
|
FormatError::Storage(_) => PyOSError::new_err(msg),
|
||||||
|
_ => other(msg),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn value_err(msg: String) -> PyErr {
|
||||||
|
PyValueError::new_err(msg)
|
||||||
|
}
|
||||||
|
|
||||||
/// The address of the object at `path`, resolved from the root group.
|
/// The address of the object at `path`, resolved from the root group.
|
||||||
pub(crate) fn address(file: &clawhdf5_rs::File, path: &str) -> PyResult<u64> {
|
pub(crate) fn address(file: &File, path: &str) -> PyResult<u64> {
|
||||||
resolve_from(file, file.superblock().root_group_address, path, path)
|
resolve_from(file, file.superblock().root_group_address, path, path)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The address of `rel` resolved from the group at `group` (`full` is the
|
/// The address of `rel` resolved from the group at `group` (`full` is the
|
||||||
/// resulting path, for the error message).
|
/// resulting path, for the error message).
|
||||||
pub(crate) fn resolve_from(
|
pub(crate) fn resolve_from(file: &File, group: u64, rel: &str, full: &str) -> PyResult<u64> {
|
||||||
file: &clawhdf5_rs::File,
|
|
||||||
group: u64,
|
|
||||||
rel: &str,
|
|
||||||
full: &str,
|
|
||||||
) -> PyResult<u64> {
|
|
||||||
if rel.is_empty() {
|
if rel.is_empty() {
|
||||||
return Ok(group);
|
return Ok(group);
|
||||||
}
|
}
|
||||||
crate::no_panic(|| {
|
clawhdf5_format::group_v2::resolve_path_from_in(file.storage(), file.superblock(), group, rel)
|
||||||
clawhdf5_format::group_v2::resolve_path_from(file.as_bytes(), file.superblock(), group, rel)
|
.map_err(|e| match e {
|
||||||
.map_err(|e| {
|
FormatError::Storage(_) => format_err(full, e, value_err),
|
||||||
PyKeyError::new_err(format!(
|
e => PyKeyError::new_err(format!(
|
||||||
"Unable to open object (object '{}' doesn't exist): {e}",
|
"Unable to open object (object '{}' doesn't exist): {e}",
|
||||||
name(full)
|
name(full)
|
||||||
))
|
)),
|
||||||
})
|
})
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The object header at `addr` (the object at `path`).
|
/// The object header at `addr` (the object at `path`).
|
||||||
pub(crate) fn header_at(file: &clawhdf5_rs::File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
|
pub(crate) fn header_at(file: &File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
|
||||||
crate::no_panic(|| {
|
let sb = file.superblock();
|
||||||
let sb = file.superblock();
|
ObjectHeader::parse_in(file.storage(), addr, sb.offset_size, sb.length_size)
|
||||||
let at = usize::try_from(addr)
|
.map_err(|e| format_err(path, e, value_err))
|
||||||
.map_err(|_| PyValueError::new_err(format!("{}: address out of range", name(path))))?;
|
|
||||||
ObjectHeader::parse(file.as_bytes(), at, sb.offset_size, sb.length_size)
|
|
||||||
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// What an object header describes.
|
/// What an object header describes.
|
||||||
@@ -96,29 +108,50 @@ pub(crate) fn kind(hdr: &ObjectHeader) -> Option<Kind> {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The kind of the object at `addr`, from its header.
|
||||||
|
pub(crate) fn kind_at(file: &File, addr: u64, path: &str) -> PyResult<Option<Kind>> {
|
||||||
|
Ok(kind(&header_at(file, addr, path)?))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// What opening an object found, read without the GIL.
|
||||||
|
enum Found {
|
||||||
|
Dataset(DatasetMeta),
|
||||||
|
Group,
|
||||||
|
Datatype,
|
||||||
|
Other,
|
||||||
|
}
|
||||||
|
|
||||||
/// Open the object at `addr` (whose path is `path`) as a `Dataset` or
|
/// Open the object at `addr` (whose path is `path`) as a `Dataset` or
|
||||||
/// `Group`. Both keep the address, so later reads resolve nothing.
|
/// `Group`. Both keep the address, so later reads resolve nothing.
|
||||||
pub(crate) fn open(
|
pub(crate) fn open(
|
||||||
py: Python<'_>,
|
py: Python<'_>,
|
||||||
file: &Arc<clawhdf5_rs::File>,
|
handle: &Arc<Handle>,
|
||||||
path: String,
|
path: String,
|
||||||
addr: u64,
|
addr: u64,
|
||||||
) -> PyResult<Py<PyAny>> {
|
) -> PyResult<Py<PyAny>> {
|
||||||
let hdr = header_at(file, addr, &path)?;
|
let found = handle.with(py, |f| {
|
||||||
match kind(&hdr) {
|
let hdr = header_at(f, addr, &path)?;
|
||||||
Some(Kind::Dataset) => Ok(PyDataset::open(py, Arc::clone(file), path, addr, &hdr)?
|
Ok(match kind(&hdr) {
|
||||||
|
Some(Kind::Dataset) => Found::Dataset(DatasetMeta::load(f, addr, &hdr, &path)?),
|
||||||
|
Some(Kind::Group) => Found::Group,
|
||||||
|
Some(Kind::Datatype) => Found::Datatype,
|
||||||
|
None => Found::Other,
|
||||||
|
})
|
||||||
|
})?;
|
||||||
|
match found {
|
||||||
|
Found::Dataset(meta) => Ok(PyDataset::new(py, Arc::clone(handle), path, addr, meta)
|
||||||
.into_pyobject(py)?
|
.into_pyobject(py)?
|
||||||
.into_any()
|
.into_any()
|
||||||
.unbind()),
|
.unbind()),
|
||||||
Some(Kind::Group) => Ok(PyGroup::from_read(Arc::clone(file), path, addr)
|
Found::Group => Ok(PyGroup::from_read(Arc::clone(handle), path, addr)
|
||||||
.into_pyobject(py)?
|
.into_pyobject(py)?
|
||||||
.into_any()
|
.into_any()
|
||||||
.unbind()),
|
.unbind()),
|
||||||
Some(Kind::Datatype) => Err(PyTypeError::new_err(format!(
|
Found::Datatype => Err(PyTypeError::new_err(format!(
|
||||||
"{}: committed (named) datatypes are not supported by clawhdf5",
|
"{}: committed (named) datatypes are not supported by clawhdf5",
|
||||||
name(&path)
|
name(&path)
|
||||||
))),
|
))),
|
||||||
None => Err(PyValueError::new_err(format!(
|
Found::Other => Err(PyValueError::new_err(format!(
|
||||||
"{}: not a dataset, group or datatype",
|
"{}: not a dataset, group or datatype",
|
||||||
name(&path)
|
name(&path)
|
||||||
))),
|
))),
|
||||||
@@ -126,32 +159,26 @@ pub(crate) fn open(
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// The dataspace message of an object header.
|
/// The dataspace message of an object header.
|
||||||
pub(crate) fn dataspace(file: &clawhdf5_rs::File, hdr: &ObjectHeader) -> PyResult<Dataspace> {
|
pub(crate) fn dataspace(file: &File, hdr: &ObjectHeader, path: &str) -> PyResult<Dataspace> {
|
||||||
crate::no_panic(|| {
|
let sb = file.superblock();
|
||||||
let sb = file.superblock();
|
let msg = hdr
|
||||||
let msg = hdr
|
.messages
|
||||||
.messages
|
.iter()
|
||||||
.iter()
|
.find(|m| m.msg_type == MessageType::Dataspace)
|
||||||
.find(|m| m.msg_type == MessageType::Dataspace)
|
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
|
||||||
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
|
let data = clawhdf5_format::shared_message::message_data_in(
|
||||||
let data = clawhdf5_format::shared_message::message_data(
|
file.storage(),
|
||||||
file.as_bytes(),
|
msg,
|
||||||
msg,
|
sb.offset_size,
|
||||||
sb.offset_size,
|
sb.length_size,
|
||||||
sb.length_size,
|
)
|
||||||
)
|
.map_err(|e| format_err(path, e, value_err))?;
|
||||||
.map_err(|e| PyValueError::new_err(e.to_string()))?;
|
Dataspace::parse(&data, sb.length_size).map_err(|e| format_err(path, e, value_err))
|
||||||
Dataspace::parse(&data, sb.length_size).map_err(|e| PyValueError::new_err(e.to_string()))
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The chunk shape of a chunked dataset (one entry per dataset dimension),
|
/// The chunk shape of a chunked dataset (one entry per dataset dimension),
|
||||||
/// or `None` for other layouts or a layout message that does not parse.
|
/// or `None` for other layouts or a layout message that does not parse.
|
||||||
pub(crate) fn chunk_shape(
|
pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Option<Vec<u64>> {
|
||||||
file: &clawhdf5_rs::File,
|
|
||||||
hdr: &ObjectHeader,
|
|
||||||
rank: usize,
|
|
||||||
) -> Option<Vec<u64>> {
|
|
||||||
let sb = file.superblock();
|
let sb = file.superblock();
|
||||||
let msg = hdr
|
let msg = hdr
|
||||||
.messages
|
.messages
|
||||||
@@ -179,24 +206,18 @@ pub(crate) fn is_null(space: &Dataspace) -> bool {
|
|||||||
/// The attributes of the object at `addr` (whose path is `path`), sorted by
|
/// The attributes of the object at `addr` (whose path is `path`), sorted by
|
||||||
/// name (h5py's order). Attributes whose messages cannot be parsed are left
|
/// name (h5py's order). Attributes whose messages cannot be parsed are left
|
||||||
/// out, as the facade's `attrs()` does.
|
/// out, as the facade's `attrs()` does.
|
||||||
pub(crate) fn attributes(
|
pub(crate) fn attributes(file: &File, addr: u64, path: &str) -> PyResult<Vec<AttributeMessage>> {
|
||||||
file: &clawhdf5_rs::File,
|
|
||||||
addr: u64,
|
|
||||||
path: &str,
|
|
||||||
) -> PyResult<Vec<AttributeMessage>> {
|
|
||||||
let hdr = header_at(file, addr, path)?;
|
let hdr = header_at(file, addr, path)?;
|
||||||
crate::no_panic(|| {
|
let sb = file.superblock();
|
||||||
let sb = file.superblock();
|
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
|
||||||
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant(
|
file.storage(),
|
||||||
file.as_bytes(),
|
&hdr,
|
||||||
&hdr,
|
sb.offset_size,
|
||||||
sb.offset_size,
|
sb.length_size,
|
||||||
sb.length_size,
|
)
|
||||||
)
|
.map_err(|e| format_err(path, e, value_err))?;
|
||||||
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))?;
|
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
|
||||||
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
|
Ok(attrs)
|
||||||
Ok(attrs)
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
#[cfg(test)]
|
#[cfg(test)]
|
||||||
|
|||||||
@@ -7,11 +7,12 @@
|
|||||||
//! (negative from the end) drop their axis, slices must have a positive
|
//! (negative from the end) drop their axis, slices must have a positive
|
||||||
//! step, one `Ellipsis` fills the unmentioned axes, a single increasing list
|
//! step, one `Ellipsis` fills the unmentioned axes, a single increasing list
|
||||||
//! of integers may index one axis, and strings name compound fields.
|
//! of integers may index one axis, and strings name compound fields.
|
||||||
//! Everything else (`None`/`np.newaxis`, boolean masks, several index lists)
|
//! Everything else (`None`/`np.newaxis`, several index lists) is refused
|
||||||
//! is refused with the error h5py gives.
|
//! with the error h5py gives; boolean masks, which h5py supports, raise
|
||||||
|
//! `NotImplementedError`.
|
||||||
|
|
||||||
use clawhdf5_format::selection::Selection;
|
use clawhdf5_format::selection::Selection;
|
||||||
use pyo3::exceptions::{PyIndexError, PyTypeError, PyValueError};
|
use pyo3::exceptions::{PyIndexError, PyNotImplementedError, PyTypeError, PyValueError};
|
||||||
use pyo3::prelude::*;
|
use pyo3::prelude::*;
|
||||||
use pyo3::types::{PyEllipsis, PySlice, PyString, PyTuple};
|
use pyo3::types::{PyEllipsis, PySlice, PyString, PyTuple};
|
||||||
|
|
||||||
@@ -254,6 +255,17 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// A mask of the dataset's whole shape (`ds[ds[()] > 0]`).
|
||||||
|
if let [a] = args.as_slice() {
|
||||||
|
let np = key.py().import("numpy")?;
|
||||||
|
if a.is_instance(&np.getattr("ndarray")?)?
|
||||||
|
&& a.getattr("dtype")?.getattr("kind")?.extract::<String>()? == "b"
|
||||||
|
&& a.getattr("ndim")?.extract::<usize>()? > 1
|
||||||
|
&& a.getattr("shape")?.extract::<Vec<u64>>()? == dims
|
||||||
|
{
|
||||||
|
return Err(mask_unsupported());
|
||||||
|
}
|
||||||
|
}
|
||||||
if args.iter().any(|a| a.is_none()) {
|
if args.iter().any(|a| a.is_none()) {
|
||||||
return Err(PyTypeError::new_err(
|
return Err(PyTypeError::new_err(
|
||||||
"Indexing with None (or np.newaxis) is not supported",
|
"Indexing with None (or np.newaxis) is not supported",
|
||||||
@@ -332,6 +344,10 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
|
|||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn mask_unsupported() -> PyErr {
|
||||||
|
PyNotImplementedError::new_err("boolean mask indexing is not supported by clawhdf5")
|
||||||
|
}
|
||||||
|
|
||||||
fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
|
fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
|
||||||
if a.is_none() {
|
if a.is_none() {
|
||||||
return Err(PyTypeError::new_err(
|
return Err(PyTypeError::new_err(
|
||||||
@@ -379,8 +395,15 @@ fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
|
|||||||
let arr = np.call_method1("asarray", (a,))?;
|
let arr = np.call_method1("asarray", (a,))?;
|
||||||
let kind: String = arr.getattr("dtype")?.getattr("kind")?.extract()?;
|
let kind: String = arr.getattr("dtype")?.getattr("kind")?.extract()?;
|
||||||
if kind == "b" {
|
if kind == "b" {
|
||||||
|
// A mask along this axis (h5py supports them; clawhdf5 does
|
||||||
|
// not, for reads or writes: an unsupported operation). A mask
|
||||||
|
// of any other shape is a wrong key, as in h5py.
|
||||||
|
let shape: Vec<u64> = arr.getattr("shape")?.extract()?;
|
||||||
|
if shape == [n] {
|
||||||
|
return Err(mask_unsupported());
|
||||||
|
}
|
||||||
return Err(PyTypeError::new_err(
|
return Err(PyTypeError::new_err(
|
||||||
"Boolean mask indexing is not supported by clawhdf5",
|
"Boolean indexing array has incompatible shape",
|
||||||
));
|
));
|
||||||
}
|
}
|
||||||
let ndim: usize = arr.getattr("ndim")?.extract()?;
|
let ndim: usize = arr.getattr("ndim")?.extract()?;
|
||||||
|
|||||||
@@ -1,6 +1,10 @@
|
|||||||
"""Shared fixtures for the clawhdf5 Python binding tests."""
|
"""Shared fixtures for the clawhdf5 Python binding tests."""
|
||||||
|
|
||||||
import os
|
import os
|
||||||
|
import re
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
@@ -16,3 +20,134 @@ def h5py():
|
|||||||
pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable")
|
pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable")
|
||||||
pytest.skip("h5py not installed")
|
pytest.skip("h5py not installed")
|
||||||
return mod
|
return mod
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# An HTTP server for remote reads
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
|
||||||
|
|
||||||
|
|
||||||
|
class RangeServer:
|
||||||
|
"""A static file server on 127.0.0.1, in a thread of this process, that
|
||||||
|
answers `Range: bytes=a-b` with 206 and `Content-Range` (the way S3 and
|
||||||
|
common web servers do), sends an ETag and honours `If-Match`.
|
||||||
|
|
||||||
|
- `ranges=False`: ignores `Range` and answers 200 with the whole file,
|
||||||
|
like a server without range support.
|
||||||
|
- `down` (set by `close()`): hang up on every request.
|
||||||
|
- `delay`: seconds to wait before answering each request after the
|
||||||
|
first `delay_after` ones (a slow network).
|
||||||
|
- `log`: every request as `(method, path, range header)`.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, root, ranges=True):
|
||||||
|
self.root = str(root)
|
||||||
|
self.ranges = ranges
|
||||||
|
self.delay = 0.0
|
||||||
|
self.delay_after = 0
|
||||||
|
self.down = False
|
||||||
|
self.log = []
|
||||||
|
self._lock = threading.Lock()
|
||||||
|
server = self
|
||||||
|
|
||||||
|
class Handler(BaseHTTPRequestHandler):
|
||||||
|
protocol_version = "HTTP/1.1"
|
||||||
|
|
||||||
|
def log_message(self, *args): # quiet
|
||||||
|
pass
|
||||||
|
|
||||||
|
def do_HEAD(self):
|
||||||
|
self._serve(body=False)
|
||||||
|
|
||||||
|
def do_GET(self):
|
||||||
|
self._serve(body=True)
|
||||||
|
|
||||||
|
def _serve(self, body):
|
||||||
|
with server._lock:
|
||||||
|
server.log.append((self.command, self.path, self.headers.get("Range")))
|
||||||
|
n = len(server.log)
|
||||||
|
if server.delay and n > server.delay_after:
|
||||||
|
time.sleep(server.delay)
|
||||||
|
if server.down:
|
||||||
|
# Hang up without an answer (open keep-alive
|
||||||
|
# connections outlive shutdown(), so close() sets this).
|
||||||
|
self.close_connection = True
|
||||||
|
return
|
||||||
|
path = os.path.join(server.root, self.path.lstrip("/").split("?")[0])
|
||||||
|
if not os.path.isfile(path):
|
||||||
|
self.send_response(404)
|
||||||
|
self.send_header("Content-Length", "0")
|
||||||
|
self.end_headers()
|
||||||
|
return
|
||||||
|
with open(path, "rb") as fh:
|
||||||
|
data = fh.read()
|
||||||
|
st = os.stat(path)
|
||||||
|
etag = f'"{st.st_mtime_ns:x}-{st.st_size:x}"'
|
||||||
|
want = self.headers.get("If-Match")
|
||||||
|
if want is not None and want != etag and want != "*":
|
||||||
|
self.send_response(412)
|
||||||
|
self.send_header("Content-Length", "0")
|
||||||
|
self.end_headers()
|
||||||
|
return
|
||||||
|
rng = self.headers.get("Range") if server.ranges else None
|
||||||
|
m = _RANGE.match(rng.strip()) if rng else None
|
||||||
|
if m and (m.group(1) or m.group(2)):
|
||||||
|
size = len(data)
|
||||||
|
if m.group(1):
|
||||||
|
start = int(m.group(1))
|
||||||
|
end = int(m.group(2)) if m.group(2) else size - 1
|
||||||
|
else:
|
||||||
|
start = max(0, size - int(m.group(2)))
|
||||||
|
end = size - 1
|
||||||
|
if start >= size:
|
||||||
|
self.send_response(416)
|
||||||
|
self.send_header("Content-Range", f"bytes */{size}")
|
||||||
|
self.send_header("Content-Length", "0")
|
||||||
|
self.end_headers()
|
||||||
|
return
|
||||||
|
end = min(end, size - 1)
|
||||||
|
part = data[start : end + 1]
|
||||||
|
self.send_response(206)
|
||||||
|
self.send_header("Content-Range", f"bytes {start}-{end}/{size}")
|
||||||
|
else:
|
||||||
|
part = data
|
||||||
|
self.send_response(200)
|
||||||
|
if server.ranges:
|
||||||
|
self.send_header("Accept-Ranges", "bytes")
|
||||||
|
self.send_header("ETag", etag)
|
||||||
|
self.send_header("Content-Length", str(len(part)))
|
||||||
|
self.send_header("Content-Type", "application/x-hdf5")
|
||||||
|
self.end_headers()
|
||||||
|
if body:
|
||||||
|
try:
|
||||||
|
self.wfile.write(part)
|
||||||
|
except (BrokenPipeError, ConnectionResetError):
|
||||||
|
pass
|
||||||
|
|
||||||
|
self.httpd = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
|
||||||
|
self.httpd.daemon_threads = True
|
||||||
|
self.port = self.httpd.server_address[1]
|
||||||
|
self.thread = threading.Thread(target=self.httpd.serve_forever, daemon=True)
|
||||||
|
self.thread.start()
|
||||||
|
|
||||||
|
def url(self, name):
|
||||||
|
return f"http://127.0.0.1:{self.port}/{name}"
|
||||||
|
|
||||||
|
def requests(self):
|
||||||
|
with self._lock:
|
||||||
|
return len(self.log)
|
||||||
|
|
||||||
|
def close(self):
|
||||||
|
self.down = True
|
||||||
|
self.httpd.shutdown()
|
||||||
|
self.httpd.server_close()
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def range_server(tmp_path):
|
||||||
|
"""A range-capable server over `tmp_path`."""
|
||||||
|
server = RangeServer(tmp_path)
|
||||||
|
yield server
|
||||||
|
server.close()
|
||||||
|
|||||||
@@ -0,0 +1,940 @@
|
|||||||
|
"""In-place editing: clawhdf5.File(path, 'r+') against h5py.
|
||||||
|
|
||||||
|
Every edit is applied twice, to two copies of the same file: once through
|
||||||
|
h5py (libhdf5) and once through clawhdf5 (FileEditor). After every edit both
|
||||||
|
files are read back with h5py and must hold the same shapes, values and
|
||||||
|
attributes; clawhdf5's own view must agree; when h5py refuses an edit,
|
||||||
|
clawhdf5 must refuse it too and leave its file as it was. Files are written
|
||||||
|
by h5py (libver earliest and latest, so every chunk index kind) and by
|
||||||
|
clawhdf5; `h5dump` must read every result."""
|
||||||
|
|
||||||
|
import io
|
||||||
|
import os
|
||||||
|
import shutil
|
||||||
|
import subprocess
|
||||||
|
import threading
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
import clawhdf5
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Files
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
ENUM = {"RED": 0, "GREEN": 1, "BLUE": 7}
|
||||||
|
|
||||||
|
|
||||||
|
def _h5py_file(h5py, path, libver):
|
||||||
|
rng = np.random.default_rng(1)
|
||||||
|
with h5py.File(path, "w", libver=libver) as f:
|
||||||
|
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
|
||||||
|
f.create_dataset("be_i2", data=np.arange(24, dtype=">i2").reshape(4, 6))
|
||||||
|
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
|
||||||
|
f.create_dataset("f2", data=rng.standard_normal(12).astype("<f2"))
|
||||||
|
f.create_dataset("f4_2d", data=rng.standard_normal((8, 9)).astype("<f4"))
|
||||||
|
f.create_dataset("f8_3d", data=rng.standard_normal((4, 5, 6)))
|
||||||
|
f.create_dataset("u8", data=np.arange(10, dtype="<u8"))
|
||||||
|
f.create_dataset("c8", data=(np.arange(6) + 1j * np.arange(6)).astype("<c8"))
|
||||||
|
f.create_dataset("bool", data=np.array([True, False, True, True]))
|
||||||
|
f.create_dataset("enum", data=np.array([0, 1, 7, 0], dtype="i1"),
|
||||||
|
dtype=h5py.enum_dtype(ENUM, basetype="i1"))
|
||||||
|
f.create_dataset("s5", data=np.array([b"ab", b"cdefg", b""], dtype="S5"))
|
||||||
|
cmp_dt = np.dtype([("id", "<i4"), ("x", "<f8"), ("tag", "S3")])
|
||||||
|
f.create_dataset("cmp", data=np.array([(i, i / 2, b"t%d" % i) for i in range(5)], dtype=cmp_dt))
|
||||||
|
f.create_dataset("scalar", data=np.float64(3.5))
|
||||||
|
# Chunked: fixed maxshape (v4 fixed array under latest), one
|
||||||
|
# unlimited dimension (extensible array), two (v2 B-tree), one chunk.
|
||||||
|
f.create_dataset("chunk_fixed", data=np.arange(100, dtype="<i8").reshape(10, 10), chunks=(3, 4))
|
||||||
|
f.create_dataset("chunk_ext", data=rng.standard_normal((12, 7)), chunks=(5, 7), maxshape=(None, 7))
|
||||||
|
f.create_dataset("chunk_bt2", data=np.arange(30, dtype="<i4").reshape(5, 6), chunks=(2, 2),
|
||||||
|
maxshape=(None, None))
|
||||||
|
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20), chunks=(6, 6),
|
||||||
|
compression="gzip", maxshape=(40, 40))
|
||||||
|
f.create_dataset("chunk_single", data=np.arange(12, dtype="<u2").reshape(3, 4), chunks=(3, 4),
|
||||||
|
maxshape=(3, 4))
|
||||||
|
f.create_dataset("chunk_fill", shape=(8,), dtype="<i4", chunks=(3,), maxshape=(20,), fillvalue=-1)
|
||||||
|
f.create_dataset("vlen", data=["a", "bb"], dtype=h5py.string_dtype())
|
||||||
|
# Compact layout (low-level API).
|
||||||
|
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
|
||||||
|
dcpl.set_layout(h5py.h5d.COMPACT)
|
||||||
|
space = h5py.h5s.create_simple((7,))
|
||||||
|
dsid = h5py.h5d.create(f.id, b"compact", h5py.h5t.STD_I32LE, space, dcpl=dcpl)
|
||||||
|
dsid.write(h5py.h5s.ALL, h5py.h5s.ALL, np.arange(7, dtype="<i4"))
|
||||||
|
g = f.create_group("grp")
|
||||||
|
g.create_dataset("leaf", data=np.arange(5.0))
|
||||||
|
g.attrs["units"] = "m"
|
||||||
|
f.attrs["version"] = np.int32(1)
|
||||||
|
|
||||||
|
|
||||||
|
def _clawhdf5_file(path):
|
||||||
|
with clawhdf5.File(str(path), "w") as f:
|
||||||
|
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
|
||||||
|
f.create_dataset("f8", data=np.linspace(0, 1, 30).reshape(5, 6))
|
||||||
|
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20),
|
||||||
|
chunks=[6, 6], compression="gzip")
|
||||||
|
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
|
||||||
|
g = f.create_group("grp")
|
||||||
|
g.create_dataset("leaf", data=np.arange(5.0))
|
||||||
|
g.attrs["units"] = "m"
|
||||||
|
f.attrs["version"] = 1
|
||||||
|
|
||||||
|
|
||||||
|
# h5py's libver: "earliest" (v1 B-tree chunk indexes), "v114" (the 1.10+
|
||||||
|
# indexes: fixed and extensible arrays, v2 B-trees, single chunk) and
|
||||||
|
# "latest" (HDF5 2.0's newest format, which h5dump 1.14 cannot read).
|
||||||
|
SOURCES = ["h5py-earliest", "h5py-v114", "h5py-latest", "clawhdf5"]
|
||||||
|
|
||||||
|
|
||||||
|
def _make(h5py, tmp_path, source):
|
||||||
|
base = tmp_path / f"base-{source}.h5"
|
||||||
|
if source == "clawhdf5":
|
||||||
|
_clawhdf5_file(base)
|
||||||
|
else:
|
||||||
|
_h5py_file(h5py, str(base), source.split("-")[1])
|
||||||
|
theirs = tmp_path / f"theirs-{source}.h5"
|
||||||
|
ours = tmp_path / f"ours-{source}.h5"
|
||||||
|
shutil.copy(base, theirs)
|
||||||
|
shutil.copy(base, ours)
|
||||||
|
return str(theirs), str(ours), str(base)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Comparing files through h5py
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _norm_attr(v):
|
||||||
|
"""Attribute values comparable across the two writers: clawhdf5 stores
|
||||||
|
`str` as fixed-length UTF-8 (h5py reads bytes), h5py as variable-length
|
||||||
|
(h5py reads str)."""
|
||||||
|
if isinstance(v, bytes):
|
||||||
|
return ("str", v.decode("utf-8"))
|
||||||
|
if isinstance(v, str):
|
||||||
|
return ("str", v)
|
||||||
|
arr = np.asarray(v)
|
||||||
|
if arr.dtype.kind in "SO":
|
||||||
|
return ("strs", [x.decode() if isinstance(x, bytes) else x for x in arr.ravel().tolist()], arr.shape)
|
||||||
|
return (arr.dtype.str, arr.shape, arr.tobytes())
|
||||||
|
|
||||||
|
|
||||||
|
def snapshot(h5py, path):
|
||||||
|
"""What h5py sees in the file: every dataset's shape, dtype, bytes and
|
||||||
|
attributes (read without locking: clawhdf5 may hold the file open)."""
|
||||||
|
out = {}
|
||||||
|
with h5py.File(path, "r", locking=False) as f:
|
||||||
|
def visit(name, obj):
|
||||||
|
attrs = {k: _norm_attr(obj.attrs[k]) for k in obj.attrs}
|
||||||
|
if isinstance(obj, h5py.Dataset):
|
||||||
|
if obj.dtype.kind == "O":
|
||||||
|
data = [x for x in obj[...].ravel().tolist()]
|
||||||
|
else:
|
||||||
|
data = obj[()].tobytes() if obj.shape is not None else None
|
||||||
|
out[name] = (obj.shape, obj.dtype.str, obj.maxshape, data, attrs)
|
||||||
|
else:
|
||||||
|
out[name] = ("group", attrs)
|
||||||
|
visit("/", f)
|
||||||
|
f.visititems(visit)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def assert_same_files(h5py, theirs, ours, what):
|
||||||
|
a, b = snapshot(h5py, theirs), snapshot(h5py, ours)
|
||||||
|
assert a.keys() == b.keys(), what
|
||||||
|
for k in a:
|
||||||
|
assert a[k] == b[k], f"{what}: {k} differs\n h5py: {a[k]}\n clawhdf5: {b[k]}"
|
||||||
|
|
||||||
|
|
||||||
|
def assert_ours_reads_like_h5py(h5py, f, path, what):
|
||||||
|
"""clawhdf5's own view of the file it is editing matches h5py's."""
|
||||||
|
with h5py.File(path, "r", locking=False) as t:
|
||||||
|
for name in ["i4", "chunk_ext", "chunk_bt2", "chunk_gzip", "f8_3d", "cmp", "bool", "enum", "scalar"]:
|
||||||
|
if name not in t:
|
||||||
|
continue
|
||||||
|
o = f[name]
|
||||||
|
assert o.shape == t[name].shape, f"{what}: {name} shape"
|
||||||
|
assert o.maxshape == t[name].maxshape, f"{what}: {name} maxshape"
|
||||||
|
np.testing.assert_array_equal(o[()], t[name][()], err_msg=f"{what}: {name}")
|
||||||
|
for obj in ["/", "grp"]:
|
||||||
|
assert sorted(f[obj].attrs.keys()) == sorted(t[obj].attrs.keys()), what
|
||||||
|
for k in t[obj].attrs:
|
||||||
|
assert _norm_attr(f[obj].attrs[k]) == _norm_attr(t[obj].attrs[k]), f"{what}: {obj}.attrs[{k}]"
|
||||||
|
|
||||||
|
|
||||||
|
def h5dump_reads(path, base=None):
|
||||||
|
"""h5dump (libhdf5 1.14) reads every object and value of `path` — when it
|
||||||
|
reads the unedited `base` (it cannot read HDF5 2.0's newest format)."""
|
||||||
|
exe = shutil.which("h5dump")
|
||||||
|
if exe is None:
|
||||||
|
if os.environ.get("CLAWHDF5_REQUIRE_INTEROP") == "1":
|
||||||
|
pytest.fail("h5dump is required (CLAWHDF5_REQUIRE_INTEROP=1)")
|
||||||
|
return
|
||||||
|
h5rs = os.environ.get("CLAWHDF5_H5RS")
|
||||||
|
if h5rs:
|
||||||
|
# clawhdf5's structural and checksum validator (scripts/ci-test.sh
|
||||||
|
# points this at the h5rs it built).
|
||||||
|
r = subprocess.run([h5rs, "check", path], capture_output=True, text=True)
|
||||||
|
assert r.returncode == 0, (r.stdout + r.stderr)[-2000:]
|
||||||
|
if base is not None and subprocess.run([exe, "-H", base], capture_output=True).returncode != 0:
|
||||||
|
return
|
||||||
|
r = subprocess.run([exe, path], capture_output=True, text=True)
|
||||||
|
assert r.returncode == 0, r.stderr[-2000:]
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Applying one edit both ways
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _native_conversion(h5py, value, ds_dtype):
|
||||||
|
"""`value` as libhdf5 converts it to `ds_dtype` in native byte order.
|
||||||
|
|
||||||
|
libhdf5 2.0 (h5py 3.16) converts numbers differently when either side is
|
||||||
|
not in native byte order (its "soft" conversions): a float in (-1, 0)
|
||||||
|
becomes the integer type's minimum instead of 0, and an unsigned integer
|
||||||
|
too large for the signed type of the same size wraps instead of
|
||||||
|
saturating. clawhdf5 applies the native-order results to every byte
|
||||||
|
order, so the reference is h5py converting into a native dataset of the
|
||||||
|
same kind; the result then reaches the real dataset by a plain byte
|
||||||
|
swap."""
|
||||||
|
if not isinstance(value, np.ndarray):
|
||||||
|
return value
|
||||||
|
if value.dtype.kind not in "biuf" or ds_dtype.kind not in "biuf":
|
||||||
|
return value
|
||||||
|
if value.dtype.isnative and ds_dtype.isnative:
|
||||||
|
return value
|
||||||
|
with h5py.File(io.BytesIO(), "w") as tmp:
|
||||||
|
d = tmp.create_dataset("t", shape=value.shape, dtype=ds_dtype.newbyteorder("="))
|
||||||
|
d[...] = value.astype(value.dtype.newbyteorder("="))
|
||||||
|
return np.asarray(d[()])
|
||||||
|
|
||||||
|
|
||||||
|
def _apply(f, op, h5py=None):
|
||||||
|
"""Apply `op` to `f`; with `h5py`, `f` is an h5py file and a numpy array
|
||||||
|
value is first converted as libhdf5 converts in native byte order (see
|
||||||
|
`_native_conversion`)."""
|
||||||
|
kind = op[0]
|
||||||
|
if kind == "set":
|
||||||
|
_, name, key, value = op
|
||||||
|
if h5py is not None:
|
||||||
|
value = _native_conversion(h5py, value, f[name].dtype)
|
||||||
|
f[name][key] = value
|
||||||
|
elif kind == "resize":
|
||||||
|
_, name, size, axis = op
|
||||||
|
if axis is None:
|
||||||
|
f[name].resize(size)
|
||||||
|
else:
|
||||||
|
f[name].resize(size, axis=axis)
|
||||||
|
elif kind == "attr":
|
||||||
|
_, obj, name, value = op
|
||||||
|
f[obj].attrs[name] = value
|
||||||
|
else:
|
||||||
|
raise AssertionError(op)
|
||||||
|
|
||||||
|
|
||||||
|
def edit_both(h5py, theirs, ours_path, ours, op):
|
||||||
|
"""Apply `op` with h5py and with clawhdf5 (`ours`, open 'r+'); the two
|
||||||
|
files must then read the same through h5py. If h5py refuses, clawhdf5
|
||||||
|
must refuse and its file must be unchanged. Returns h5py's error."""
|
||||||
|
before = snapshot(h5py, ours_path)
|
||||||
|
try:
|
||||||
|
with h5py.File(theirs, "r+") as t:
|
||||||
|
_apply(t, op, h5py)
|
||||||
|
except Exception as e: # noqa: BLE001 - h5py refuses: so must we
|
||||||
|
try:
|
||||||
|
_apply(ours, op)
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
pass
|
||||||
|
else:
|
||||||
|
pytest.fail(f"{op!r:.300}: h5py refused ({type(e).__name__}: {e}), clawhdf5 did not")
|
||||||
|
assert snapshot(h5py, ours_path) == before, f"{op!r}: clawhdf5 changed the file while failing"
|
||||||
|
return e
|
||||||
|
_apply(ours, op)
|
||||||
|
assert_same_files(h5py, theirs, ours_path, repr(op)[:200])
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Tests
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("source", SOURCES)
|
||||||
|
def test_edit_sequence_matches_h5py(h5py, tmp_path, source):
|
||||||
|
theirs, ours_path, base = _make(h5py, tmp_path, source)
|
||||||
|
ops = [
|
||||||
|
("set", "i4", 0, 99),
|
||||||
|
("set", "i4", (slice(1, 5, 2), slice(None, None, 3)), np.array([[1.5, -2.5, 1e12, -1e12]])),
|
||||||
|
("set", "i4", (slice(None), 2), np.arange(6, dtype="<i8") * 1000),
|
||||||
|
("set", "i4", ([0, 2, 5], slice(4, 6)), np.array([[1, 2], [3, 4], [5, 6]])),
|
||||||
|
("set", "i4", (Ellipsis, -1), np.int8(-7)),
|
||||||
|
("set", "i4", (3, 3), 12345.9),
|
||||||
|
("set", "u1", slice(2, 8), np.array([-5, 0, 300, 255, 256, 1], dtype="<i4")),
|
||||||
|
("set", "u1", slice(0, 3), [1, 2, 3]),
|
||||||
|
("set", "chunk_gzip", (slice(0, 20, 7), slice(3, 17)), 42.25),
|
||||||
|
("set", "chunk_gzip", (slice(5, 11), slice(5, 11)), np.ones((6, 6), dtype="<f8") * np.pi),
|
||||||
|
("set", "grp/leaf", slice(None), np.array([1, 2, 3, 4, 5], dtype="<u2")),
|
||||||
|
("attr", "/", "version", np.int32(2)),
|
||||||
|
("attr", "/", "count", 7),
|
||||||
|
("attr", "grp", "scale", np.array([0.5, 0.25], dtype="<f4")),
|
||||||
|
("attr", "grp", "matrix", np.arange(6, dtype=">i8").reshape(2, 3)),
|
||||||
|
("attr", "i4", "flag", np.bool_(True)),
|
||||||
|
("attr", "i4", "z", np.complex64(1 - 2j)),
|
||||||
|
("attr", "i4", "raw", np.bytes_(b"abc")),
|
||||||
|
("set", "i4", slice(0, 2), np.zeros((3, 10))), # shape mismatch: refused by both
|
||||||
|
]
|
||||||
|
if source != "clawhdf5":
|
||||||
|
ops += [
|
||||||
|
("set", "be_i2", (slice(None), slice(1, 3)), np.array([70000, -70000], dtype="<i8")),
|
||||||
|
("set", "f2", slice(None, None, 4), np.array([1e6, -3.25, 0.1])),
|
||||||
|
("set", "f4_2d", (2, slice(None)), np.linspace(-1, 1, 9)),
|
||||||
|
("set", "f8_3d", (slice(1, 3), 2, slice(None, None, 2)), np.arange(3, dtype="<i2")),
|
||||||
|
("set", "u8", slice(None), np.array([-1, 0, 2**63, 1e30, -1e30, 5.5, 2, 3, 4, 5])),
|
||||||
|
("set", "c8", slice(1, 3), np.array([1 + 1j, 2 - 2j], dtype="<c16")),
|
||||||
|
("set", "c8", 0, np.float64(1.0)), # h5py: no conversion path
|
||||||
|
("set", "bool", slice(None), np.array([0, 3, 0, -1], dtype="<i4")),
|
||||||
|
("set", "bool", 1, np.array(True)),
|
||||||
|
("set", "enum", slice(0, 2), np.array([7, 1], dtype="<i4")),
|
||||||
|
("set", "s5", 0, np.bytes_(b"xyzuvw")),
|
||||||
|
("set", "s5", slice(1, 3), [b"q", b"rs"]),
|
||||||
|
("set", "s5", 2, np.array("uni")), # h5py: no conversion from 'U'
|
||||||
|
("set", "cmp", 2, np.array((9, 9.5, b"zz"), dtype=[("id", "<i4"), ("x", "<f8"), ("tag", "S3")])),
|
||||||
|
("set", "cmp", slice(3, 5), [(1, 0.5, b"a"), (2, 1.5, b"b")]),
|
||||||
|
("set", "scalar", (), 7.25),
|
||||||
|
("set", "scalar", Ellipsis, np.float32(-1.5)),
|
||||||
|
("set", "compact", slice(1, 6, 2), np.array([10, 20, 30])),
|
||||||
|
("set", "chunk_fixed", (slice(2, 9), slice(1, 10, 4)), np.arange(21).reshape(7, 3)),
|
||||||
|
("set", "chunk_single", (1, slice(None)), np.array([9, 8, 7, 6])),
|
||||||
|
("resize", "chunk_ext", (20, 7), None),
|
||||||
|
("set", "chunk_ext", slice(12, 20), np.full((8, 7), 2.5)),
|
||||||
|
("resize", "chunk_ext", 9, 0),
|
||||||
|
("resize", "chunk_ext", 16, 0),
|
||||||
|
("resize", "chunk_bt2", (9, 11), None),
|
||||||
|
("set", "chunk_bt2", (slice(4, 9), slice(5, 11)), np.arange(30).reshape(5, 6)),
|
||||||
|
("resize", "chunk_bt2", (3, 3), None),
|
||||||
|
("resize", "chunk_bt2", (7, 8), None),
|
||||||
|
("resize", "chunk_gzip", (40, 25), None),
|
||||||
|
("resize", "chunk_gzip", (41, 25), None), # beyond maxshape: refused
|
||||||
|
("resize", "chunk_fill", 15, None), # h5py: a size without axis must be a tuple
|
||||||
|
("resize", "chunk_fill", (15,), None),
|
||||||
|
("set", "chunk_fill", slice(10, 12), [1, 2]),
|
||||||
|
("resize", "chunk_fill", (4,), None),
|
||||||
|
("resize", "chunk_fill", (20,), None),
|
||||||
|
("resize", "i4", (7, 10), None), # not chunked: refused
|
||||||
|
]
|
||||||
|
with clawhdf5.File(ours_path, "r+") as ours:
|
||||||
|
assert ours.mode == "r+"
|
||||||
|
for i, op in enumerate(ops):
|
||||||
|
edit_both(h5py, theirs, ours_path, ours, op)
|
||||||
|
if i % 5 == 0:
|
||||||
|
assert_ours_reads_like_h5py(h5py, ours, ours_path, repr(op))
|
||||||
|
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
|
||||||
|
h5dump_reads(ours_path, base)
|
||||||
|
# Reopened, clawhdf5 reads what h5py reads.
|
||||||
|
with clawhdf5.File(ours_path, "r") as f:
|
||||||
|
assert f.mode == "r"
|
||||||
|
assert_ours_reads_like_h5py(h5py, f, ours_path, "reopened")
|
||||||
|
|
||||||
|
|
||||||
|
def _random_key(rng, shape):
|
||||||
|
key = []
|
||||||
|
for n in shape:
|
||||||
|
r = rng.random()
|
||||||
|
if n == 0 or r < 0.3:
|
||||||
|
start = int(rng.integers(0, n + 1)) if n else 0
|
||||||
|
stop = int(rng.integers(start, n + 1)) if n else 0
|
||||||
|
step = int(rng.integers(1, 4))
|
||||||
|
key.append(slice(start, stop, step))
|
||||||
|
elif r < 0.55:
|
||||||
|
key.append(int(rng.integers(-n, n)))
|
||||||
|
elif r < 0.7:
|
||||||
|
k = int(rng.integers(1, min(n, 4) + 1))
|
||||||
|
key.append(sorted(rng.choice(n, size=k, replace=False).tolist()))
|
||||||
|
else:
|
||||||
|
key.append(slice(None))
|
||||||
|
# Only one index list per key.
|
||||||
|
lists = [i for i, k in enumerate(key) if isinstance(k, list)]
|
||||||
|
for i in lists[1:]:
|
||||||
|
key[i] = slice(None)
|
||||||
|
return tuple(key)
|
||||||
|
|
||||||
|
|
||||||
|
def _selection_shape(key, shape):
|
||||||
|
out = []
|
||||||
|
fancy = False
|
||||||
|
for k, n in zip(key, shape):
|
||||||
|
if isinstance(k, slice):
|
||||||
|
out.append(len(range(*k.indices(n))))
|
||||||
|
elif isinstance(k, list):
|
||||||
|
out.append(len(k))
|
||||||
|
fancy = True
|
||||||
|
return tuple(out), fancy
|
||||||
|
|
||||||
|
|
||||||
|
def _random_value(rng, sel_shape, fancy, dtype):
|
||||||
|
r = rng.random()
|
||||||
|
if r < 0.2 or not sel_shape:
|
||||||
|
v = rng.standard_normal() * 1000
|
||||||
|
return np.float64(v) if rng.random() < 0.5 else int(v)
|
||||||
|
shape = list(sel_shape)
|
||||||
|
if not fancy and r < 0.35 and shape:
|
||||||
|
shape[0] = 1 # broadcast along the first axis
|
||||||
|
kind = rng.choice(["same", "f8", "i8", "u1", "f4"])
|
||||||
|
if kind == "same" and np.dtype(dtype).kind in "iuf":
|
||||||
|
dt = np.dtype(dtype)
|
||||||
|
else:
|
||||||
|
dt = np.dtype(str(kind) if kind != "same" else "f8")
|
||||||
|
base = rng.standard_normal(size=shape) * (10 ** rng.integers(0, 6))
|
||||||
|
with np.errstate(all="ignore"):
|
||||||
|
return base.astype(dt)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("source", SOURCES)
|
||||||
|
@pytest.mark.parametrize("seed", [0, 1, 2, 3])
|
||||||
|
def test_random_edits_match_h5py(h5py, tmp_path, source, seed):
|
||||||
|
"""Random writes (slices, steps, integers, index lists, broadcasts,
|
||||||
|
other dtypes and out-of-range values), resizes and attributes, each
|
||||||
|
compared with h5py applying the same edit."""
|
||||||
|
rng = np.random.default_rng(1000 * seed + SOURCES.index(source))
|
||||||
|
theirs, ours_path, base = _make(h5py, tmp_path, source)
|
||||||
|
with h5py.File(theirs, "r") as t:
|
||||||
|
names = [n for n in ["i4", "u1", "f8", "f4_2d", "f8_3d", "be_i2", "chunk_fixed", "chunk_ext",
|
||||||
|
"chunk_bt2", "chunk_gzip", "chunk_fill", "compact", "grp/leaf"] if n in t]
|
||||||
|
resizable = [n for n in ["chunk_ext", "chunk_bt2", "chunk_gzip", "chunk_fill"] if n in names]
|
||||||
|
refused = 0
|
||||||
|
with clawhdf5.File(ours_path, "r+") as ours:
|
||||||
|
for step in range(40):
|
||||||
|
r = rng.random()
|
||||||
|
if r < 0.15 and resizable:
|
||||||
|
name = str(rng.choice(resizable))
|
||||||
|
with h5py.File(theirs, "r") as t:
|
||||||
|
maxshape = t[name].maxshape
|
||||||
|
shape = t[name].shape
|
||||||
|
new = tuple(int(rng.integers(0, (m if m is not None else s + 10) + 1)) for m, s in zip(maxshape, shape))
|
||||||
|
op = ("resize", name, new, None)
|
||||||
|
elif r < 0.25:
|
||||||
|
obj = str(rng.choice(["/", "grp", names[0]]))
|
||||||
|
choices = [np.int16(rng.integers(-100, 100)), rng.standard_normal(3),
|
||||||
|
np.arange(int(rng.integers(1, 5)), dtype=">u4"), np.float32(0.5)]
|
||||||
|
value = choices[int(rng.integers(0, len(choices)))]
|
||||||
|
op = ("attr", obj, f"a{int(rng.integers(0, 4))}", value)
|
||||||
|
else:
|
||||||
|
name = str(rng.choice(names))
|
||||||
|
with h5py.File(theirs, "r") as t:
|
||||||
|
shape, dtype = t[name].shape, t[name].dtype
|
||||||
|
key = _random_key(rng, shape)
|
||||||
|
sel_shape, fancy = _selection_shape(key, shape)
|
||||||
|
op = ("set", name, key, _random_value(rng, sel_shape, fancy, dtype))
|
||||||
|
before = None
|
||||||
|
if op[0] == "resize":
|
||||||
|
with h5py.File(ours_path, "r", locking=False) as o:
|
||||||
|
before, fill = o[op[1]][()], o[op[1]].fillvalue
|
||||||
|
if edit_both(h5py, theirs, ours_path, ours, op) is not None:
|
||||||
|
refused += 1
|
||||||
|
elif before is not None:
|
||||||
|
# Independently of h5py (which a wrong index layout fools
|
||||||
|
# the same way): the kept elements keep their values.
|
||||||
|
want = resized_model(before, op[2], fill)
|
||||||
|
np.testing.assert_array_equal(ours[op[1]][()], want, err_msg=repr(op))
|
||||||
|
with h5py.File(ours_path, "r", locking=False) as o:
|
||||||
|
np.testing.assert_array_equal(o[op[1]][()], want, err_msg=repr(op))
|
||||||
|
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
|
||||||
|
assert refused < 30
|
||||||
|
h5dump_reads(ours_path, base)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("seed", range(10, 40))
|
||||||
|
def test_random_edits_on_clawhdf5_files(h5py, tmp_path, seed):
|
||||||
|
"""More random sequences on a clawhdf5-written file, whose resizes to
|
||||||
|
zero extents once left files `h5rs check` could not read."""
|
||||||
|
test_random_edits_match_h5py(h5py, tmp_path, "clawhdf5", seed)
|
||||||
|
|
||||||
|
|
||||||
|
def resized_model(before, shape, fill):
|
||||||
|
"""`before` resized to `shape` as HDF5 resizes: elements inside both
|
||||||
|
extents keep their values, the others read as the fill value."""
|
||||||
|
out = np.full(shape, fill, dtype=before.dtype)
|
||||||
|
common = tuple(slice(0, min(a, b)) for a, b in zip(before.shape, shape))
|
||||||
|
out[common] = before[common]
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
RESIZES = [(15, 15), (3, 2), (20, 20), (1, 1), (1, 0), (0, 0), (7, 20), (20, 13), (20, 20)]
|
||||||
|
|
||||||
|
|
||||||
|
def _check_resizes(h5py, path, name, orig, maxshape=(20, 20), base=None):
|
||||||
|
"""Resize `name` through RESIZES in 'r+', each checked against a numpy
|
||||||
|
model with clawhdf5 and h5py and the file with `h5rs check`."""
|
||||||
|
model = orig
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
ds = f[name]
|
||||||
|
for shape in RESIZES:
|
||||||
|
ds.resize(shape)
|
||||||
|
model = resized_model(model, shape, 0)
|
||||||
|
np.testing.assert_array_equal(ds[()], model, err_msg=f"{name} {shape}")
|
||||||
|
with h5py.File(path, "r", locking=False) as t:
|
||||||
|
np.testing.assert_array_equal(t[name][()], model, err_msg=f"h5py: {name} {shape}")
|
||||||
|
assert t[name].maxshape == maxshape
|
||||||
|
h5rs = os.environ.get("CLAWHDF5_H5RS")
|
||||||
|
if h5rs: # h5dump would wait for the editor's lock
|
||||||
|
r = subprocess.run([h5rs, "check", "--data", path], capture_output=True, text=True)
|
||||||
|
assert r.returncode == 0, f"{name} {shape}: " + (r.stdout + r.stderr)[-2000:]
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
ds.resize((maxshape[0] + 1, 20))
|
||||||
|
# Written values survive a shrink.
|
||||||
|
ds[...] = orig
|
||||||
|
ds.resize((15, 15))
|
||||||
|
with h5py.File(path, "r") as t:
|
||||||
|
np.testing.assert_array_equal(t[name][()], orig[:15, :15])
|
||||||
|
h5dump_reads(path, base)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("source", SOURCES)
|
||||||
|
def test_resizes_keep_values(h5py, tmp_path, source):
|
||||||
|
"""Shrinking, zero extents and growing back keep the values a numpy
|
||||||
|
model keeps, on every source (on clawhdf5's own files a shrink once
|
||||||
|
moved every chunk: h5py read the same wrong values)."""
|
||||||
|
_, ours, base = _make(h5py, tmp_path, source)
|
||||||
|
with clawhdf5.File(ours, "r+") as f:
|
||||||
|
f["chunk_gzip"].resize((20, 20))
|
||||||
|
maxshape = (20, 20) if source == "clawhdf5" else (40, 40)
|
||||||
|
_check_resizes(h5py, ours, "chunk_gzip", np.arange(400, dtype="<f4").reshape(20, 20), maxshape, base)
|
||||||
|
|
||||||
|
|
||||||
|
FIXTURES = os.path.join(os.path.dirname(__file__), "..", "..", "clawhdf5", "tests", "fixtures")
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("name", ["d", "z"])
|
||||||
|
def test_resizes_of_a_file_with_no_recorded_maxshape(h5py, tmp_path, name):
|
||||||
|
"""A file clawhdf5 2.7.0 wrote records no maximum dimensions for chunked
|
||||||
|
datasets; resizing it must keep the Fixed Array index's layout."""
|
||||||
|
path = str(tmp_path / "old.h5")
|
||||||
|
shutil.copy(os.path.join(FIXTURES, "chunked_no_maxshape_v2_7_0.h5"), path)
|
||||||
|
with h5py.File(path, "r") as t:
|
||||||
|
orig = t[name][()]
|
||||||
|
_check_resizes(h5py, path, name, orig)
|
||||||
|
|
||||||
|
|
||||||
|
NUMERIC = ["<i1", "<u1", "<i2", ">u2", "<i4", "<u4", ">i8", "<u8", "<f2", "<f4", ">f8"]
|
||||||
|
|
||||||
|
|
||||||
|
def _quiet(f):
|
||||||
|
with np.errstate(all="ignore"):
|
||||||
|
return f()
|
||||||
|
|
||||||
|
|
||||||
|
def _libhdf5_undefined(vals, target):
|
||||||
|
"""The values whose conversion to `target` libhdf5 2.0 (h5py 3.16) gets
|
||||||
|
wrong even in native byte order, where its C casts are undefined
|
||||||
|
behaviour; clawhdf5 saturates them as libhdf5's range handling
|
||||||
|
intends (docs/known-issues.md):
|
||||||
|
|
||||||
|
- half floats into unsigned integers: negatives wrap (-1 -> 65535) and
|
||||||
|
+inf becomes 0; into signed integers, +-inf becomes the minimum;
|
||||||
|
- a float equal to the integer maximum rounded up in the float's
|
||||||
|
precision (float32(2**31 - 1) == 2**31 -> int32, float64(2**64 - 1)
|
||||||
|
-> uint64) becomes the minimum (or 0);
|
||||||
|
- a double between 65504 and 65520 into a half float becomes infinity
|
||||||
|
(IEEE rounds it down to 65504, as numpy does).
|
||||||
|
"""
|
||||||
|
bad = np.zeros(vals.shape, dtype=bool)
|
||||||
|
if vals.dtype.kind != "f":
|
||||||
|
return bad
|
||||||
|
if target.kind in "iu":
|
||||||
|
if vals.dtype.itemsize == 2:
|
||||||
|
bad |= np.isinf(vals)
|
||||||
|
if target.kind == "u":
|
||||||
|
bad |= vals <= -1
|
||||||
|
top = _quiet(lambda: np.array(np.iinfo(target).max).astype(vals.dtype))
|
||||||
|
if float(top) > np.iinfo(target).max:
|
||||||
|
bad |= vals == top
|
||||||
|
if target.kind == "f" and target.itemsize == 2 and vals.dtype.itemsize > 2:
|
||||||
|
bad |= (np.abs(vals) > 65504) & (np.abs(vals) < 65520)
|
||||||
|
return bad
|
||||||
|
|
||||||
|
|
||||||
|
def test_numeric_conversions_match_h5py(h5py, tmp_path):
|
||||||
|
"""Every numeric source dtype into every numeric dataset dtype, with
|
||||||
|
values at and beyond the targets' limits, as libhdf5 converts them."""
|
||||||
|
edge = np.array([0, 1, -1, -0.3, 2.5, -2.5, 3.7, -3.7, 127.9, -128.9, 200.5, 255.5, 256, -129,
|
||||||
|
32767.5, 40000, 65504, 70000, -70000, 2**31 - 1, 2**31, -2**31 - 1,
|
||||||
|
4e9, 1e15, -1e15, 1e19, 1e300, -1e300, np.inf, -np.inf])
|
||||||
|
sources = {
|
||||||
|
"f8": edge,
|
||||||
|
"f4": _quiet(lambda: edge.astype("<f4")),
|
||||||
|
"f2": np.array([0, 1, -1, 2.5, -3.5, 65504, -65504, np.inf, -np.inf, 100.5], dtype="<f2"),
|
||||||
|
"i8": np.array([0, 1, -1, 127, 128, -129, 255, 256, 32768, -32769, 65536, 2**31, -2**31 - 1,
|
||||||
|
2**32, 2**62, -2**63, 2**63 - 1], dtype="<i8"),
|
||||||
|
"u8": np.array([0, 1, 127, 128, 255, 256, 65535, 65536, 2**31, 2**32, 2**63, 2**64 - 1], dtype="<u8"),
|
||||||
|
"i1": np.array([-128, -1, 0, 1, 127], dtype="i1"),
|
||||||
|
"u2": np.array([0, 255, 256, 65535], dtype=">u2"),
|
||||||
|
"b": np.array([True, False, True]),
|
||||||
|
}
|
||||||
|
for target in NUMERIC:
|
||||||
|
for sname, src in sources.items():
|
||||||
|
vals = src[~_libhdf5_undefined(src, np.dtype(target))]
|
||||||
|
path_t = str(tmp_path / f"t_{target[1:]}_{sname}.h5")
|
||||||
|
path_o = str(tmp_path / f"o_{target[1:]}_{sname}.h5")
|
||||||
|
with h5py.File(path_t, "w") as f:
|
||||||
|
f.create_dataset("d", shape=vals.shape, dtype=target)
|
||||||
|
shutil.copy(path_t, path_o)
|
||||||
|
what = f"{vals.dtype} -> {target}"
|
||||||
|
try:
|
||||||
|
with h5py.File(path_t, "r+") as f:
|
||||||
|
f["d"][...] = _native_conversion(h5py, vals, f["d"].dtype)
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
with clawhdf5.File(path_o, "r+") as f, pytest.raises(Exception):
|
||||||
|
f["d"][...] = vals
|
||||||
|
continue
|
||||||
|
with clawhdf5.File(path_o, "r+") as f:
|
||||||
|
f["d"][...] = vals
|
||||||
|
with h5py.File(path_t, "r") as a, h5py.File(path_o, "r") as b:
|
||||||
|
assert a["d"][...].tobytes() == b["d"][...].tobytes(), (
|
||||||
|
f"{what}: h5py {a['d'][...].tolist()} clawhdf5 {b['d'][...].tolist()}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_nan_into_an_integer_dataset_is_refused(h5py, tmp_path):
|
||||||
|
"""libhdf5 stores NaN as an arbitrary integer (0, the minimum or 2**63,
|
||||||
|
depending on the type); clawhdf5 refuses and writes nothing."""
|
||||||
|
path = str(tmp_path / "nan.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
with pytest.raises(ValueError, match="NaN"):
|
||||||
|
f["d"][...] = np.array([1.0, np.nan, 2.0, 3.0])
|
||||||
|
# A Python list goes through numpy, which refuses NaN too.
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
f["d"][0:2] = [np.nan, 1.0]
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
np.testing.assert_array_equal(f["d"][...], np.arange(4))
|
||||||
|
|
||||||
|
|
||||||
|
def test_unsupported_edits_are_clear_errors(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "u.h5")
|
||||||
|
_h5py_file(h5py, path, "earliest")
|
||||||
|
before = snapshot(h5py, path)
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
with pytest.raises(NotImplementedError, match="delet"):
|
||||||
|
del f.attrs["version"]
|
||||||
|
with pytest.raises(NotImplementedError, match="delet"):
|
||||||
|
del f["i4"]
|
||||||
|
with pytest.raises(NotImplementedError, match="delet"):
|
||||||
|
del f["grp"]["leaf"]
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f.create_dataset("new", data=np.arange(3.0))
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f.create_group("newgrp")
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f["grp"].create_dataset("new", data=np.arange(3.0))
|
||||||
|
with pytest.raises(NotImplementedError, match="variable-length"):
|
||||||
|
f["vlen"][0] = "x"
|
||||||
|
with pytest.raises(NotImplementedError, match="field"):
|
||||||
|
f["cmp"]["id"] = np.arange(5)
|
||||||
|
with pytest.raises(NotImplementedError):
|
||||||
|
f.attrs["empty"] = clawhdf5.Empty("f8")
|
||||||
|
with pytest.raises(TypeError, match="chunked"):
|
||||||
|
f["i4"].resize((7, 10))
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
f["chunk_gzip"].resize((41, 20))
|
||||||
|
with pytest.raises(ValueError, match="axis"):
|
||||||
|
f["chunk_ext"].resize(3, axis=2)
|
||||||
|
with pytest.raises(TypeError):
|
||||||
|
f["i4"][0] = np.array(["a"] * 10)
|
||||||
|
# h5py writes through boolean masks; clawhdf5 does not.
|
||||||
|
with pytest.raises(NotImplementedError, match="mask"):
|
||||||
|
f["u1"][np.arange(16) % 2 == 0] = 5
|
||||||
|
with pytest.raises(NotImplementedError, match="mask"):
|
||||||
|
f["i4"][f["i4"][()] > 30] = 0
|
||||||
|
with pytest.raises(NotImplementedError, match="mask"):
|
||||||
|
f["i4"][np.ones(6, dtype=bool), 2] = 0
|
||||||
|
assert snapshot(h5py, path) == before
|
||||||
|
h5dump_reads(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_read_only_files_and_modes(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "m.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"), chunks=(2,), maxshape=(None,))
|
||||||
|
with clawhdf5.File(path, "r") as f:
|
||||||
|
with pytest.raises(OSError, match="r\\+"):
|
||||||
|
f["d"][0] = 1
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
f["d"].resize((8,))
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
f.attrs["x"] = 1
|
||||||
|
with pytest.raises(NotImplementedError, match="does not exist"):
|
||||||
|
clawhdf5.File(str(tmp_path / "missing.h5"), "a")
|
||||||
|
with pytest.raises(ValueError, match="mode"):
|
||||||
|
clawhdf5.File(path, "rw")
|
||||||
|
with clawhdf5.File(path, "a") as f:
|
||||||
|
assert f.mode == "r+"
|
||||||
|
f["d"][1] = 10
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
assert f["d"][1] == 10
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_file_is_locked_while_open_for_editing(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "lock.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||||
|
f = clawhdf5.File(path, "r+")
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
clawhdf5.File(path, "r+")
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
h5py.File(path, "r+")
|
||||||
|
f.close()
|
||||||
|
with h5py.File(path, "r+") as g:
|
||||||
|
g["d"][0] = 5
|
||||||
|
with clawhdf5.File(path, "r+") as g:
|
||||||
|
g["d"][1] = 6
|
||||||
|
with h5py.File(path, "r") as g:
|
||||||
|
np.testing.assert_array_equal(g["d"][...], [5, 6, 2, 3])
|
||||||
|
|
||||||
|
|
||||||
|
def test_objects_see_edits_made_through_others(h5py, tmp_path):
|
||||||
|
"""A dataset or attrs object taken before an edit reports the file as
|
||||||
|
it is after it: the new shape, the new attribute."""
|
||||||
|
path = str(tmp_path / "live.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(6.0), chunks=(4,), maxshape=(None,))
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
d1 = f["d"]
|
||||||
|
d2 = f["d"]
|
||||||
|
attrs = d1.attrs
|
||||||
|
assert len(attrs) == 0 and "u" not in attrs
|
||||||
|
d2.resize((10,))
|
||||||
|
assert d1.shape == (10,) and d1.size == 10 and len(d1) == 10
|
||||||
|
np.testing.assert_array_equal(d1[6:], np.zeros(4))
|
||||||
|
d2.attrs["u"] = "m/s"
|
||||||
|
assert "u" in attrs and attrs["u"] == b"m/s" and len(attrs) == 1
|
||||||
|
f.attrs.create("shaped", np.arange(6), shape=(2, 3), dtype="<i2")
|
||||||
|
f.attrs.modify("shaped2", [1.5, 2.5])
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
assert f["d"].shape == (10,)
|
||||||
|
assert f["d"].attrs["u"] == b"m/s"
|
||||||
|
assert f.attrs["shaped"].dtype == np.dtype("<i2") and f.attrs["shaped"].shape == (2, 3)
|
||||||
|
np.testing.assert_array_equal(f.attrs["shaped2"], [1.5, 2.5])
|
||||||
|
|
||||||
|
|
||||||
|
def test_attribute_types_as_h5py_reads_them(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "attrs.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_group("g")
|
||||||
|
values = {
|
||||||
|
"i8": 5,
|
||||||
|
"f8": 2.5,
|
||||||
|
"i1": np.int8(-3),
|
||||||
|
"u8": np.uint64(2**64 - 1),
|
||||||
|
">f4": np.array([1.5, 2.5], dtype=">f4"),
|
||||||
|
"f2": np.float16(0.5),
|
||||||
|
"b": True,
|
||||||
|
"barr": np.array([True, False]),
|
||||||
|
"c16": np.complex128(1 + 2j),
|
||||||
|
"bytes": b"raw",
|
||||||
|
"sarr": np.array([b"a", b"bcd"]),
|
||||||
|
"str": "héllo",
|
||||||
|
"strs": ["x", "yz"],
|
||||||
|
"2d": np.arange(12, dtype="<u2").reshape(3, 4),
|
||||||
|
"empty": np.zeros((0,), dtype="<i4"),
|
||||||
|
}
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
for k, v in values.items():
|
||||||
|
f["g"].attrs[k] = v
|
||||||
|
# Replace one, with another type and size.
|
||||||
|
f["g"].attrs["i8"] = np.arange(100.0)
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
a = f["g"].attrs
|
||||||
|
np.testing.assert_array_equal(a["i8"], np.arange(100.0))
|
||||||
|
assert a["f8"] == 2.5 and a["f8"].dtype == np.float64
|
||||||
|
assert a["i1"] == -3 and a["i1"].dtype == np.int8
|
||||||
|
assert a["u8"] == 2**64 - 1 and a["u8"].dtype == np.uint64
|
||||||
|
assert a[">f4"].dtype == np.dtype(">f4")
|
||||||
|
assert a["f2"].dtype == np.float16
|
||||||
|
assert a["b"] is np.True_ or a["b"] == True # noqa: E712
|
||||||
|
assert a["barr"].dtype == np.bool_
|
||||||
|
assert a["c16"] == 1 + 2j
|
||||||
|
assert a["bytes"] == b"raw"
|
||||||
|
assert list(a["sarr"]) == [b"a", b"bcd"]
|
||||||
|
# str is stored as fixed-length UTF-8: h5py reads bytes.
|
||||||
|
assert a["str"].decode("utf-8") == "héllo"
|
||||||
|
assert [x.decode() for x in a["strs"]] == ["x", "yz"]
|
||||||
|
assert a["2d"].shape == (3, 4) and a["2d"].dtype == np.dtype("<u2")
|
||||||
|
assert a["empty"].shape == (0,)
|
||||||
|
with clawhdf5.File(path, "r") as f:
|
||||||
|
assert f["g"].attrs["c16"] == 1 + 2j
|
||||||
|
assert f["g"].attrs["str"].decode("utf-8") == "héllo"
|
||||||
|
h5dump_reads(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_many_attributes_move_to_dense_storage(h5py, tmp_path):
|
||||||
|
"""Past the compact limit (8 attributes under libver v114) the object's
|
||||||
|
attributes move to dense storage; h5py reads all of them."""
|
||||||
|
path = str(tmp_path / "dense.h5")
|
||||||
|
with h5py.File(path, "w", libver="v114") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(3))
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
for i in range(20):
|
||||||
|
f["d"].attrs[f"a{i:02d}"] = np.full(i + 1, i, dtype="<i2")
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
assert sorted(f["d"].attrs.keys()) == [f"a{i:02d}" for i in range(20)]
|
||||||
|
for i in range(20):
|
||||||
|
np.testing.assert_array_equal(f["d"].attrs[f"a{i:02d}"], np.full(i + 1, i))
|
||||||
|
h5dump_reads(path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_reads_never_see_a_half_written_edit(h5py, tmp_path):
|
||||||
|
"""Readers on other threads while one thread rewrites a dataset: every
|
||||||
|
read returns one whole version (all elements equal), never a mix."""
|
||||||
|
path = str(tmp_path / "race.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.zeros((64, 64)), chunks=(16, 16), compression="gzip")
|
||||||
|
errors = []
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
stop = threading.Event()
|
||||||
|
|
||||||
|
def read():
|
||||||
|
ds = f["d"]
|
||||||
|
while not stop.is_set():
|
||||||
|
a = ds[...]
|
||||||
|
if not (a == a.flat[0]).all():
|
||||||
|
errors.append(a)
|
||||||
|
return
|
||||||
|
|
||||||
|
readers = [threading.Thread(target=read) for _ in range(3)]
|
||||||
|
for t in readers:
|
||||||
|
t.start()
|
||||||
|
try:
|
||||||
|
for k in range(1, 25):
|
||||||
|
f["d"][...] = float(k)
|
||||||
|
finally:
|
||||||
|
stop.set()
|
||||||
|
for t in readers:
|
||||||
|
t.join()
|
||||||
|
np.testing.assert_array_equal(f["d"][...], np.full((64, 64), 24.0))
|
||||||
|
assert not errors, "a read saw a partly written dataset"
|
||||||
|
|
||||||
|
|
||||||
|
def test_close_releases_the_file(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "close.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||||
|
f = clawhdf5.File(path, "r+")
|
||||||
|
ds = f["d"]
|
||||||
|
f.close()
|
||||||
|
# The handle still reads the file as last written, but cannot edit it.
|
||||||
|
np.testing.assert_array_equal(ds[...], np.arange(4))
|
||||||
|
with pytest.raises(OSError, match="closed"):
|
||||||
|
ds[0] = 1
|
||||||
|
with h5py.File(path, "r+") as g:
|
||||||
|
g["d"][0] = 9
|
||||||
|
|
||||||
|
|
||||||
|
def _two_files(h5py, tmp_path):
|
||||||
|
"""d1/f.h5 and d2/f.h5: same name, different layouts (the review's
|
||||||
|
repro)."""
|
||||||
|
(tmp_path / "d1").mkdir()
|
||||||
|
(tmp_path / "d2").mkdir()
|
||||||
|
with h5py.File(tmp_path / "d1" / "f.h5", "w") as f:
|
||||||
|
f.create_dataset("x", data=np.arange(10, dtype="<i4"))
|
||||||
|
f.create_dataset("big", data=np.full(5000, 1.5))
|
||||||
|
with h5py.File(tmp_path / "d2" / "f.h5", "w") as f:
|
||||||
|
f.create_dataset("pad", data=np.full(3000, 2.5))
|
||||||
|
f.create_dataset("x", data=np.arange(10, dtype="<i4") + 500)
|
||||||
|
return tmp_path / "d1" / "f.h5", tmp_path / "d2" / "f.h5"
|
||||||
|
|
||||||
|
|
||||||
|
def _held_file_edited(h5py, held, other, other_bytes):
|
||||||
|
"""The edits went to `held`, planned from its own metadata; `other`
|
||||||
|
was not touched."""
|
||||||
|
assert other.read_bytes() == other_bytes, "the other file changed"
|
||||||
|
with h5py.File(held, "r") as f:
|
||||||
|
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
|
||||||
|
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
|
||||||
|
np.testing.assert_array_equal(f.attrs["note"], np.arange(50.0))
|
||||||
|
h5dump_reads(str(held))
|
||||||
|
|
||||||
|
|
||||||
|
def test_relative_path_and_chdir(h5py, tmp_path, monkeypatch):
|
||||||
|
"""A file opened by a relative path keeps being the file edited and read
|
||||||
|
after os.chdir (an edit was planned from the file the path named in the
|
||||||
|
new directory and written into the one open, corrupting it)."""
|
||||||
|
held, other = _two_files(h5py, tmp_path)
|
||||||
|
other_bytes = other.read_bytes()
|
||||||
|
monkeypatch.chdir(held.parent)
|
||||||
|
f = clawhdf5.File("f.h5", "r+")
|
||||||
|
ds = f["x"]
|
||||||
|
monkeypatch.chdir(other.parent)
|
||||||
|
ds[:] = np.full(10, 7, "<i4")
|
||||||
|
np.testing.assert_array_equal(ds[:5], np.full(5, 7)) # read from the held file
|
||||||
|
f.attrs["note"] = np.arange(50.0)
|
||||||
|
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
|
||||||
|
assert "pad" not in f
|
||||||
|
f.close()
|
||||||
|
_held_file_edited(h5py, held, other, other_bytes)
|
||||||
|
# 'w' writes where the path named when the file was opened.
|
||||||
|
monkeypatch.chdir(held.parent)
|
||||||
|
w = clawhdf5.File("new.h5", "w")
|
||||||
|
monkeypatch.chdir(other.parent)
|
||||||
|
w.create_dataset("d", data=np.arange(3))
|
||||||
|
w.close()
|
||||||
|
assert (held.parent / "new.h5").exists() and not (other.parent / "new.h5").exists()
|
||||||
|
|
||||||
|
|
||||||
|
def test_path_replaced_between_edits(h5py, tmp_path):
|
||||||
|
"""The path renamed away and another file put in its place between
|
||||||
|
edits: the edits go to the file held open, never mixed with the other."""
|
||||||
|
held, other = _two_files(h5py, tmp_path)
|
||||||
|
path = str(tmp_path / "f.h5")
|
||||||
|
os.replace(held, path)
|
||||||
|
moved = tmp_path / "moved.h5"
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
ds = f["x"]
|
||||||
|
ds[0] = 7
|
||||||
|
os.replace(path, moved)
|
||||||
|
shutil.copy(other, path)
|
||||||
|
other_bytes = open(path, "rb").read()
|
||||||
|
ds[:] = np.full(10, 7, "<i4")
|
||||||
|
f.attrs["note"] = np.arange(50.0)
|
||||||
|
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
|
||||||
|
assert "pad" not in f
|
||||||
|
from pathlib import Path
|
||||||
|
_held_file_edited(h5py, moved, Path(path), other_bytes)
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_edit_releases_the_gil(h5py, tmp_path):
|
||||||
|
"""Another Python thread keeps running while a large edit is written:
|
||||||
|
had the edit held the GIL, the other thread would stall for the whole
|
||||||
|
edit (Rust code never yields it)."""
|
||||||
|
import time
|
||||||
|
path = str(tmp_path / "big.h5")
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("d", shape=(1024, 1024), dtype="<f8", chunks=(64, 64), compression="gzip")
|
||||||
|
value = np.random.default_rng(0).standard_normal((1024, 1024))
|
||||||
|
stamps = []
|
||||||
|
stop = threading.Event()
|
||||||
|
|
||||||
|
def spin():
|
||||||
|
while not stop.is_set():
|
||||||
|
stamps.append(time.perf_counter())
|
||||||
|
|
||||||
|
with clawhdf5.File(path, "r+") as f:
|
||||||
|
ds = f["d"]
|
||||||
|
t = threading.Thread(target=spin)
|
||||||
|
t.start()
|
||||||
|
try:
|
||||||
|
time.sleep(0.05)
|
||||||
|
t0 = time.perf_counter()
|
||||||
|
ds[...] = value
|
||||||
|
t1 = time.perf_counter()
|
||||||
|
finally:
|
||||||
|
stop.set()
|
||||||
|
t.join()
|
||||||
|
np.testing.assert_array_equal(ds[...], value)
|
||||||
|
during = [x for x in stamps if t0 <= x <= t1]
|
||||||
|
gaps = np.diff([t0] + during + [t1])
|
||||||
|
assert t1 - t0 > 0.1, "the edit is too quick to tell"
|
||||||
|
assert gaps.max() < 0.5 * (t1 - t0), (
|
||||||
|
f"the other thread stalled for {gaps.max():.3f} s of a {t1 - t0:.3f} s edit")
|
||||||
@@ -178,15 +178,39 @@ def _write_fixture(h5py, path):
|
|||||||
g.attrs["depth"] = np.int8(3)
|
g.attrs["depth"] = np.int8(3)
|
||||||
|
|
||||||
|
|
||||||
@pytest.fixture(scope="module")
|
@pytest.fixture(scope="module", params=["local", "r+", "http", "http-1k-blocks"])
|
||||||
def pair(h5py, tmp_path_factory):
|
def pair(request, h5py, tmp_path_factory):
|
||||||
path = str(tmp_path_factory.mktemp("h5") / "fixture.h5")
|
"""The fixture file through h5py and through clawhdf5: opened locally
|
||||||
|
(read-only, and for editing: a copy, since editing locks the file), and
|
||||||
|
over HTTP range requests (a local server in this process) with the
|
||||||
|
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
|
||||||
|
through many small ranges."""
|
||||||
|
import shutil
|
||||||
|
|
||||||
|
from conftest import RangeServer
|
||||||
|
|
||||||
|
root = tmp_path_factory.mktemp("h5")
|
||||||
|
path = str(root / "fixture.h5")
|
||||||
_write_fixture(h5py, path)
|
_write_fixture(h5py, path)
|
||||||
theirs = h5py.File(path, "r")
|
theirs = h5py.File(path, "r")
|
||||||
ours = clawhdf5.File(path, "r")
|
server = None
|
||||||
|
if request.param == "local":
|
||||||
|
ours = clawhdf5.File(path, "r")
|
||||||
|
elif request.param == "r+":
|
||||||
|
copy = str(root / "editable.h5")
|
||||||
|
shutil.copy(path, copy)
|
||||||
|
ours = clawhdf5.File(copy, "r+")
|
||||||
|
else:
|
||||||
|
server = RangeServer(root)
|
||||||
|
if request.param == "http":
|
||||||
|
ours = clawhdf5.File(server.url("fixture.h5"))
|
||||||
|
else:
|
||||||
|
ours = clawhdf5.File.open_url(server.url("fixture.h5"), block_size=1024)
|
||||||
yield ours, theirs, path
|
yield ours, theirs, path
|
||||||
theirs.close()
|
theirs.close()
|
||||||
ours.close()
|
ours.close()
|
||||||
|
if server is not None:
|
||||||
|
server.close()
|
||||||
|
|
||||||
|
|
||||||
def _all_datasets(h5py, f):
|
def _all_datasets(h5py, f):
|
||||||
@@ -416,7 +440,7 @@ def test_unsupported_types_are_errors_not_data(pair):
|
|||||||
|
|
||||||
def test_boolean_masks_are_refused(pair):
|
def test_boolean_masks_are_refused(pair):
|
||||||
ours, _, _ = pair
|
ours, _, _ = pair
|
||||||
with pytest.raises(TypeError):
|
with pytest.raises(NotImplementedError, match="mask"):
|
||||||
ours["num/le_i4_1d"][np.ones(37, dtype=bool)]
|
ours["num/le_i4_1d"][np.ones(37, dtype=bool)]
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,244 @@
|
|||||||
|
"""Remote files: `clawhdf5.File(url)` / `File.open_url(url, ...)` read over
|
||||||
|
HTTP range requests (clawhdf5-remote's block cache), against a server in this
|
||||||
|
process (conftest.RangeServer). Values are compared with h5py reading the
|
||||||
|
same file locally; the rest checks what the server saw (only the blocks a
|
||||||
|
read needs are fetched), the failure modes (no range support, a missing
|
||||||
|
file, a file that changes, a server that goes away: errors, never wrong
|
||||||
|
data), and that the GIL is released while a read waits on the network."""
|
||||||
|
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
import clawhdf5
|
||||||
|
from conftest import RangeServer
|
||||||
|
|
||||||
|
|
||||||
|
def _write(h5py, path):
|
||||||
|
rng = np.random.default_rng(7)
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.create_dataset("contig", data=rng.standard_normal((400, 300)))
|
||||||
|
f.create_dataset(
|
||||||
|
"chunked",
|
||||||
|
data=rng.integers(0, 1000, size=(512, 512), dtype="<i4"),
|
||||||
|
chunks=(64, 64),
|
||||||
|
compression="gzip",
|
||||||
|
)
|
||||||
|
f.create_dataset("strings", data=["alpha", "beta", "gamma"], dtype=h5py.string_dtype())
|
||||||
|
g = f.create_group("grp")
|
||||||
|
g.create_dataset("small", data=np.arange(10, dtype="<u2"))
|
||||||
|
g.attrs["units"] = "m/s"
|
||||||
|
f.attrs["title"] = "remote test"
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def remote_file(h5py, tmp_path, range_server):
|
||||||
|
path = tmp_path / "remote.h5"
|
||||||
|
_write(h5py, str(path))
|
||||||
|
return path, range_server.url("remote.h5"), range_server
|
||||||
|
|
||||||
|
|
||||||
|
def test_remote_reads_match_h5py(h5py, remote_file):
|
||||||
|
path, url, server = remote_file
|
||||||
|
with h5py.File(path, "r") as theirs, clawhdf5.File(url) as ours:
|
||||||
|
assert ours.filename == url
|
||||||
|
assert list(ours.keys()) == list(theirs.keys())
|
||||||
|
assert "grp/small" in ours and "nope" not in ours
|
||||||
|
for name in ["contig", "chunked", "grp/small"]:
|
||||||
|
np.testing.assert_array_equal(ours[name][...], theirs[name][...])
|
||||||
|
np.testing.assert_array_equal(ours[name][3:7], theirs[name][3:7])
|
||||||
|
np.testing.assert_array_equal(ours["chunked"][100:130, 200:300:3], theirs["chunked"][100:130, 200:300:3])
|
||||||
|
np.testing.assert_array_equal(ours["chunked"][[1, 70, 300], 5], theirs["chunked"][[1, 70, 300], 5])
|
||||||
|
assert list(ours["strings"][...]) == list(theirs["strings"][...])
|
||||||
|
assert ours["grp"].attrs["units"] == theirs["grp"].attrs["units"]
|
||||||
|
assert ours.attrs["title"] == theirs.attrs["title"]
|
||||||
|
assert ours["chunked"].maxshape == theirs["chunked"].maxshape
|
||||||
|
assert server.requests() >= 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_small_read_fetches_only_its_blocks(h5py, remote_file):
|
||||||
|
"""With 4 KiB blocks, opening and reading one chunk of a 1 MB chunked
|
||||||
|
dataset costs a handful of requests and a few blocks, not the file."""
|
||||||
|
path, url, server = remote_file
|
||||||
|
size = os.path.getsize(path)
|
||||||
|
f = clawhdf5.File.open_url(url, block_size=4096)
|
||||||
|
opened = server.requests()
|
||||||
|
assert opened == 1, server.log
|
||||||
|
ds = f["chunked"]
|
||||||
|
got = ds[0:10, 0:10]
|
||||||
|
with h5py.File(path, "r") as theirs:
|
||||||
|
np.testing.assert_array_equal(got, theirs["chunked"][0:10, 0:10])
|
||||||
|
stats = f.remote_stats
|
||||||
|
assert stats["bytes_fetched"] < size / 4, (stats, size)
|
||||||
|
assert server.requests() - opened <= 12, server.log
|
||||||
|
# A second read of the same region is served by the cache.
|
||||||
|
before = server.requests()
|
||||||
|
ds[0:10, 0:10]
|
||||||
|
assert server.requests() == before
|
||||||
|
assert f.remote_stats["hits"] > stats["hits"]
|
||||||
|
assert clawhdf5.File(str(path), "r").remote_stats is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_server_without_range_support(h5py, tmp_path):
|
||||||
|
"""A server that ignores Range answers 200 with the whole file: that is
|
||||||
|
an OSError by default, and a whole download when allowed."""
|
||||||
|
path = tmp_path / "remote.h5"
|
||||||
|
_write(h5py, str(path))
|
||||||
|
server = RangeServer(tmp_path, ranges=False)
|
||||||
|
try:
|
||||||
|
url = server.url("remote.h5")
|
||||||
|
with pytest.raises(OSError, match="range"):
|
||||||
|
clawhdf5.File(url)
|
||||||
|
with clawhdf5.File.open_url(url, allow_full_download=True) as ours, h5py.File(path, "r") as theirs:
|
||||||
|
np.testing.assert_array_equal(ours["chunked"][...], theirs["chunked"][...])
|
||||||
|
np.testing.assert_array_equal(ours["contig"][5], theirs["contig"][5])
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
clawhdf5.File.open_url(url, allow_full_download=True, max_full_download=1000)
|
||||||
|
finally:
|
||||||
|
server.close()
|
||||||
|
|
||||||
|
|
||||||
|
def test_errors_are_oserrors(remote_file):
|
||||||
|
_, url, server = remote_file
|
||||||
|
with pytest.raises(OSError, match="404"):
|
||||||
|
clawhdf5.File(server.url("missing.h5"))
|
||||||
|
with pytest.raises(ValueError, match="read-only"):
|
||||||
|
clawhdf5.File(url, "r+")
|
||||||
|
with pytest.raises(ValueError, match="read-only"):
|
||||||
|
clawhdf5.File(url, "w")
|
||||||
|
with pytest.raises(OSError, match="unsupported URL"):
|
||||||
|
clawhdf5.File("nosuchscheme://x/y.h5")
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
clawhdf5.File.open_url(url, block_size=0)
|
||||||
|
with pytest.raises(TypeError):
|
||||||
|
clawhdf5.File.open_url(url, no_such_option=1)
|
||||||
|
|
||||||
|
|
||||||
|
def test_object_store_urls_need_their_features():
|
||||||
|
"""The default wheel has no S3/GCS/Azure clients (aws-lc-rs builds C):
|
||||||
|
such a URL is an OSError naming the build feature."""
|
||||||
|
for url, feature in [("s3://bucket/k.h5", "s3"), ("gs://b/k.h5", "gcs"), ("az://c/k.h5", "azure")]:
|
||||||
|
try:
|
||||||
|
clawhdf5.File(url)
|
||||||
|
except OSError as e:
|
||||||
|
if "feature" in str(e):
|
||||||
|
assert f"`{feature}`" in str(e), str(e)
|
||||||
|
else:
|
||||||
|
pytest.fail(f"{url} opened")
|
||||||
|
|
||||||
|
|
||||||
|
def test_https_needs_the_https_feature():
|
||||||
|
"""The default wheel has no TLS stack (rustls needs ring, which builds C):
|
||||||
|
an https URL is an OSError that names the build feature."""
|
||||||
|
with pytest.raises(OSError) as e:
|
||||||
|
clawhdf5.File("https://127.0.0.1:1/x.h5")
|
||||||
|
msg = str(e.value)
|
||||||
|
# Built with `--features https` the error is the refused connection.
|
||||||
|
assert "https" in msg or "connect" in msg.lower() or "refused" in msg.lower(), msg
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_changed_file_is_an_error_not_mixed_data(h5py, remote_file):
|
||||||
|
path, url, _ = remote_file
|
||||||
|
f = clawhdf5.File.open_url(url, block_size=1024)
|
||||||
|
first = f["grp/small"][...]
|
||||||
|
# Rewrite the file with other values: new ETag, same name.
|
||||||
|
time.sleep(0.01)
|
||||||
|
with h5py.File(path, "w") as g:
|
||||||
|
g.create_dataset("contig", data=np.zeros((400, 300)))
|
||||||
|
with pytest.raises(OSError, match="changed"):
|
||||||
|
f["contig"][...]
|
||||||
|
np.testing.assert_array_equal(first, np.arange(10, dtype="<u2"))
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_server_that_goes_away_is_an_error(h5py, tmp_path):
|
||||||
|
path = tmp_path / "remote.h5"
|
||||||
|
_write(h5py, str(path))
|
||||||
|
server = RangeServer(tmp_path)
|
||||||
|
f = clawhdf5.File.open_url(server.url("remote.h5"), block_size=1024, retries=0, timeout=2)
|
||||||
|
ds = f["contig"]
|
||||||
|
server.close()
|
||||||
|
with pytest.raises(OSError):
|
||||||
|
ds[...]
|
||||||
|
|
||||||
|
|
||||||
|
def test_threads_read_one_remote_file(h5py, remote_file):
|
||||||
|
path, url, _ = remote_file
|
||||||
|
f = clawhdf5.File.open_url(url, block_size=2048)
|
||||||
|
with h5py.File(path, "r") as theirs:
|
||||||
|
expected = theirs["chunked"][...]
|
||||||
|
errors = []
|
||||||
|
|
||||||
|
def work(i):
|
||||||
|
try:
|
||||||
|
rows = slice((i * 37) % 400, (i * 37) % 400 + 64)
|
||||||
|
np.testing.assert_array_equal(f["chunked"][rows], expected[rows])
|
||||||
|
except Exception as e: # noqa: BLE001
|
||||||
|
errors.append(e)
|
||||||
|
|
||||||
|
threads = [threading.Thread(target=work, args=(i,)) for i in range(16)]
|
||||||
|
for t in threads:
|
||||||
|
t.start()
|
||||||
|
for t in threads:
|
||||||
|
t.join()
|
||||||
|
assert not errors, errors[:3]
|
||||||
|
|
||||||
|
|
||||||
|
def test_remote_reads_release_the_gil(h5py, remote_file):
|
||||||
|
"""A read waiting on a slow server lets other Python threads run: a
|
||||||
|
thread counting in a loop keeps counting (and never stalls for long)
|
||||||
|
while the main thread reads through requests that each take 0.2 s."""
|
||||||
|
_, url, server = remote_file
|
||||||
|
f = clawhdf5.File.open_url(url, block_size=1024, max_parallel=1)
|
||||||
|
ds = f["contig"]
|
||||||
|
server.delay = 0.2
|
||||||
|
server.delay_after = server.requests()
|
||||||
|
old = sys.getswitchinterval()
|
||||||
|
sys.setswitchinterval(0.001)
|
||||||
|
stop = threading.Event()
|
||||||
|
progress = {"n": 0, "worst": 0.0}
|
||||||
|
|
||||||
|
def spin():
|
||||||
|
last = time.perf_counter()
|
||||||
|
while not stop.is_set():
|
||||||
|
now = time.perf_counter()
|
||||||
|
progress["worst"] = max(progress["worst"], now - last)
|
||||||
|
last = now
|
||||||
|
progress["n"] += 1
|
||||||
|
|
||||||
|
t = threading.Thread(target=spin)
|
||||||
|
try:
|
||||||
|
t.start()
|
||||||
|
time.sleep(0.02)
|
||||||
|
t0 = time.perf_counter()
|
||||||
|
before = server.requests()
|
||||||
|
ds[0:2]
|
||||||
|
took = time.perf_counter() - t0
|
||||||
|
stop.set()
|
||||||
|
t.join()
|
||||||
|
finally:
|
||||||
|
sys.setswitchinterval(old)
|
||||||
|
assert server.requests() > before
|
||||||
|
assert took >= 0.2, took
|
||||||
|
assert progress["n"] > 1000, progress
|
||||||
|
# Held across a 0.2 s request, the spinner would stall that long.
|
||||||
|
assert progress["worst"] < 0.1, (progress, took)
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_clawhdf5_written_file_reads_the_same_remotely(tmp_path, range_server):
|
||||||
|
path = tmp_path / "ours.h5"
|
||||||
|
data = np.arange(3000, dtype="<f8").reshape(100, 30)
|
||||||
|
with clawhdf5.File(str(path), "w") as f:
|
||||||
|
f.create_dataset("d", data=data, chunks=[10, 30], compression="gzip")
|
||||||
|
g = f.create_group("g")
|
||||||
|
g.create_dataset("i", data=np.arange(5, dtype="<i4"))
|
||||||
|
f.attrs["k"] = 3
|
||||||
|
with clawhdf5.File(range_server.url("ours.h5")) as f, clawhdf5.File(str(path)) as local:
|
||||||
|
np.testing.assert_array_equal(f["d"][...], data)
|
||||||
|
np.testing.assert_array_equal(f["d"][5:9, ::4], local["d"][5:9, ::4])
|
||||||
|
np.testing.assert_array_equal(f["g/i"][...], np.arange(5))
|
||||||
|
assert f.attrs["k"] == local.attrs["k"]
|
||||||
|
assert "file (read" in repr(f).lower() and "127.0.0.1" in repr(f)
|
||||||
@@ -1509,3 +1509,68 @@ fn unencodable_filters_are_unsupported() {
|
|||||||
"a refused edit changed the file"
|
"a refused edit changed the file"
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn fixture(dir: &Path, name: &str) -> std::path::PathBuf {
|
||||||
|
let path = dir.join(name);
|
||||||
|
std::fs::copy(
|
||||||
|
Path::new(env!("CARGO_MANIFEST_DIR"))
|
||||||
|
.join("../clawhdf5/tests/fixtures")
|
||||||
|
.join(name),
|
||||||
|
&path,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
path
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Zero extents on chunked datasets with no recorded maximum. The unfixed
|
||||||
|
/// editor left `chunk_zero_extent_no_maxshape.h5` (a 2.7.0-written file
|
||||||
|
/// resized to 1x0): a Fixed Array whose maximum, taken from the current
|
||||||
|
/// dimensions, has no chunks along one dimension, so every stride before it
|
||||||
|
/// is 0 — `h5rs check` panicked dividing by it and the next resize failed
|
||||||
|
/// with an internal error. Such a file must check clean and resize on; a
|
||||||
|
/// 2.7.0-written file taken through zero extents by the fixed editor must
|
||||||
|
/// check clean at every step and read the fill value where it grew.
|
||||||
|
#[test]
|
||||||
|
fn zero_extent_resizes_without_a_recorded_maximum() {
|
||||||
|
if !tools_ok() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = tmpdir();
|
||||||
|
let path = fixture(dir.path(), "chunk_zero_extent_no_maxshape.h5");
|
||||||
|
check_tools(&path, true);
|
||||||
|
let mut ed = FileEditor::open(&path).unwrap();
|
||||||
|
ed.resize("d", &[0, 0]).unwrap();
|
||||||
|
ed.resize("z", &[0, 0]).unwrap();
|
||||||
|
// Their maximum is now what the index was laid out by (1 x 0).
|
||||||
|
ed.resize("d", &[1, 0]).unwrap();
|
||||||
|
assert!(matches!(
|
||||||
|
ed.resize("d", &[1, 1]),
|
||||||
|
Err(Error::InvalidArgument(_))
|
||||||
|
));
|
||||||
|
drop(ed);
|
||||||
|
check_tools(&path, true);
|
||||||
|
|
||||||
|
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
|
||||||
|
for shape in [[15, 15], [3, 2], [1, 1], [1, 0], [0, 0], [0, 20], [20, 20]] {
|
||||||
|
let mut ed = FileEditor::open(&path).unwrap();
|
||||||
|
ed.resize("d", &shape).unwrap();
|
||||||
|
ed.resize("z", &shape).unwrap();
|
||||||
|
drop(ed);
|
||||||
|
check_tools(&path, true);
|
||||||
|
}
|
||||||
|
let f = File::open(&path).unwrap();
|
||||||
|
for name in ["d", "z"] {
|
||||||
|
let d = f.dataset(name).unwrap();
|
||||||
|
assert_eq!(d.shape().unwrap(), [20, 20]);
|
||||||
|
assert!(d.read_f32().unwrap().iter().all(|&v| v == 0.0), "{name}");
|
||||||
|
}
|
||||||
|
assert_eq!(
|
||||||
|
py(&format!(
|
||||||
|
"import h5py\n\
|
||||||
|
with h5py.File({:?}) as f:\n\
|
||||||
|
\x20 print(int(abs(f['d'][()]).sum() + abs(f['z'][()]).sum()), f['d'].maxshape)",
|
||||||
|
path.to_str().unwrap()
|
||||||
|
)),
|
||||||
|
"0 (20, 20)"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|||||||
@@ -65,7 +65,8 @@ const MSG_FLAG_DONTSHARE: u8 = 0x04;
|
|||||||
/// that may be mid-update.
|
/// that may be mid-update.
|
||||||
///
|
///
|
||||||
/// Every method is one self-contained edit: it re-reads the file's
|
/// Every method is one self-contained edit: it re-reads the file's
|
||||||
/// metadata, applies the change, and syncs the file before returning.
|
/// metadata (from the file it holds open, never by path), applies the
|
||||||
|
/// change, and syncs the file before returning.
|
||||||
///
|
///
|
||||||
/// # What it can change
|
/// # What it can change
|
||||||
///
|
///
|
||||||
@@ -750,9 +751,17 @@ impl FileEditor {
|
|||||||
/// consistent: a metadata cache image, paged or persistent free-space
|
/// consistent: a metadata cache image, paged or persistent free-space
|
||||||
/// management, a multi-file driver, a file another writer has marked
|
/// management, a multi-file driver, a file another writer has marked
|
||||||
/// open (superblock version 3 consistency flags).
|
/// open (superblock version 3 consistency flags).
|
||||||
|
///
|
||||||
|
/// The path is only used to open the file: every edit is planned from
|
||||||
|
/// and written to the file opened here, even if the path is renamed,
|
||||||
|
/// replaced or (relative) resolved from another working directory
|
||||||
|
/// later. [`path`](Self::path) is the absolute path it had at open.
|
||||||
pub fn open<P: AsRef<Path>>(path: P) -> Result<Self, Error> {
|
pub fn open<P: AsRef<Path>>(path: P) -> Result<Self, Error> {
|
||||||
let path = path.as_ref().to_path_buf();
|
let file = OpenOptions::new()
|
||||||
let file = OpenOptions::new().read(true).write(true).open(&path)?;
|
.read(true)
|
||||||
|
.write(true)
|
||||||
|
.open(path.as_ref())?;
|
||||||
|
let path = std::fs::canonicalize(path.as_ref())?;
|
||||||
match file.try_lock() {
|
match file.try_lock() {
|
||||||
Ok(()) => {}
|
Ok(()) => {}
|
||||||
Err(TryLockError::WouldBlock) => {
|
Err(TryLockError::WouldBlock) => {
|
||||||
@@ -768,16 +777,81 @@ impl FileEditor {
|
|||||||
file,
|
file,
|
||||||
free: FreeList::default(),
|
free: FreeList::default(),
|
||||||
};
|
};
|
||||||
let f = File::open(&ed.path)?;
|
let f = ed.plan_reader()?;
|
||||||
check_editable(&f)?;
|
check_editable(&f)?;
|
||||||
Ok(ed)
|
Ok(ed)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The file's path.
|
/// The file's absolute path when it was opened (it may have been
|
||||||
|
/// renamed since; the editor keeps editing the file it opened).
|
||||||
pub fn path(&self) -> &Path {
|
pub fn path(&self) -> &Path {
|
||||||
&self.path
|
&self.path
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// A reader over the file this editor holds, as last written: the file
|
||||||
|
/// opened by [`open`](Self::open), not whatever its path names now.
|
||||||
|
///
|
||||||
|
/// It opens the file anew (read-only), so it does not share the
|
||||||
|
/// editor's lock and stays usable after the editor is dropped: on Linux
|
||||||
|
/// through `/proc/self/fd`, which reaches the held file even after its
|
||||||
|
/// path was renamed or replaced; elsewhere by the path the file had at
|
||||||
|
/// open, refused with [`Error::Io`] when that path no longer names the
|
||||||
|
/// held file (on Unix, compared by device and inode; Windows cannot
|
||||||
|
/// check). With the `mmap` feature the reader maps the file: edits
|
||||||
|
/// through the editor change the bytes it sees, so take a new reader
|
||||||
|
/// after each edit rather than reading through an old one while an edit
|
||||||
|
/// runs.
|
||||||
|
pub fn reader(&self) -> Result<File, Error> {
|
||||||
|
let dir = self.path.parent().map(Path::to_path_buf);
|
||||||
|
File::from_std_file(self.reopen()?, dir)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A new read-only open file description of the held file (see
|
||||||
|
/// [`reader`](Self::reader)).
|
||||||
|
fn reopen(&self) -> Result<std::fs::File, Error> {
|
||||||
|
#[cfg(target_os = "linux")]
|
||||||
|
{
|
||||||
|
use std::os::fd::AsRawFd;
|
||||||
|
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
|
||||||
|
if let Ok(f) = std::fs::File::open(proc) {
|
||||||
|
return Ok(f);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let f = std::fs::File::open(&self.path)?;
|
||||||
|
#[cfg(unix)]
|
||||||
|
{
|
||||||
|
use std::os::unix::fs::MetadataExt;
|
||||||
|
let (a, b) = (self.file.metadata()?, f.metadata()?);
|
||||||
|
if (a.dev(), a.ino()) != (b.dev(), b.ino()) {
|
||||||
|
return Err(Error::Io(std::io::Error::other(format!(
|
||||||
|
"{} no longer names the file being edited (renamed or replaced)",
|
||||||
|
self.path.display()
|
||||||
|
))));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Ok(f)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A reader over the held file for planning an edit (dropped before the
|
||||||
|
/// edit writes). On Linux a new open file description through
|
||||||
|
/// `/proc/self/fd`: a mapping of a clone of the held descriptor would
|
||||||
|
/// share its `flock`, and a process forked meanwhile (any
|
||||||
|
/// `std::process::Command` on another thread) would briefly keep the
|
||||||
|
/// lock alive after the editor is dropped. Elsewhere a clone of the
|
||||||
|
/// held descriptor, which follows the file wherever its path goes.
|
||||||
|
fn plan_reader(&self) -> Result<File, Error> {
|
||||||
|
let dir = self.path.parent().map(Path::to_path_buf);
|
||||||
|
#[cfg(target_os = "linux")]
|
||||||
|
{
|
||||||
|
use std::os::fd::AsRawFd;
|
||||||
|
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
|
||||||
|
if let Ok(f) = std::fs::File::open(proc) {
|
||||||
|
return File::from_std_file(f, dir);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
File::from_std_file(self.file.try_clone()?, dir)
|
||||||
|
}
|
||||||
|
|
||||||
/// Bytes earlier edits of this editor freed that later ones can still
|
/// Bytes earlier edits of this editor freed that later ones can still
|
||||||
/// reuse.
|
/// reuse.
|
||||||
pub fn reusable_bytes(&self) -> u64 {
|
pub fn reusable_bytes(&self) -> u64 {
|
||||||
@@ -794,7 +868,7 @@ impl FileEditor {
|
|||||||
&mut self,
|
&mut self,
|
||||||
op: impl FnOnce(&File, &mut Image<'_>) -> Result<R, Error>,
|
op: impl FnOnce(&File, &mut Image<'_>) -> Result<R, Error>,
|
||||||
) -> Result<R, Error> {
|
) -> Result<R, Error> {
|
||||||
let f = File::open(&self.path)?;
|
let f = self.plan_reader()?;
|
||||||
check_editable(&f)?;
|
check_editable(&f)?;
|
||||||
let sb = f.superblock().clone();
|
let sb = f.superblock().clone();
|
||||||
let user_block = f.user_block_size();
|
let user_block = f.user_block_size();
|
||||||
@@ -877,7 +951,7 @@ impl FileEditor {
|
|||||||
/// the fill value (`H5D__chunk_prune_by_extent`).
|
/// the fill value (`H5D__chunk_prune_by_extent`).
|
||||||
pub fn resize(&mut self, path: &str, shape: &[u64]) -> Result<(), Error> {
|
pub fn resize(&mut self, path: &str, shape: &[u64]) -> Result<(), Error> {
|
||||||
self.edit(|f, img| {
|
self.edit(|f, img| {
|
||||||
let t = Target::load(f, path)?;
|
let mut t = Target::load(f, path)?;
|
||||||
let dims = t.dims().to_vec();
|
let dims = t.dims().to_vec();
|
||||||
if shape.len() != dims.len() {
|
if shape.len() != dims.len() {
|
||||||
return Err(Error::InvalidArgument(format!(
|
return Err(Error::InvalidArgument(format!(
|
||||||
@@ -889,7 +963,12 @@ impl FileEditor {
|
|||||||
if shape == dims.as_slice() {
|
if shape == dims.as_slice() {
|
||||||
return Ok(());
|
return Ok(());
|
||||||
}
|
}
|
||||||
let max = t.ds.max_dimensions.clone().unwrap_or_else(|| dims.clone());
|
// No maximum recorded means the current dimensions (see below).
|
||||||
|
let record_max = t.ds.max_dimensions.is_none();
|
||||||
|
let max =
|
||||||
|
t.ds.max_dimensions
|
||||||
|
.get_or_insert_with(|| dims.clone())
|
||||||
|
.clone();
|
||||||
for d in 0..dims.len() {
|
for d in 0..dims.len() {
|
||||||
if shape[d] > max[d] {
|
if shape[d] > max[d] {
|
||||||
return Err(Error::InvalidArgument(format!(
|
return Err(Error::InvalidArgument(format!(
|
||||||
@@ -928,7 +1007,36 @@ impl FileEditor {
|
|||||||
}
|
}
|
||||||
put_uint(&mut dims_bytes[d * ls..], n, img.ls);
|
put_uint(&mut dims_bytes[d * ls..], n, img.ls);
|
||||||
}
|
}
|
||||||
hdr.patch(img, i, first, &dims_bytes)?;
|
if !record_max {
|
||||||
|
hdr.patch(img, i, first, &dims_bytes)?;
|
||||||
|
} else {
|
||||||
|
// No maximum recorded (clawhdf5's writer, for a dataset
|
||||||
|
// created without a maxshape). libhdf5 never writes such a
|
||||||
|
// dataspace: `H5S_set_extent_simple` records the maximum,
|
||||||
|
// equal to the dimensions when none is given. Reading one,
|
||||||
|
// libhdf5 takes the maximum to be the *current* dimensions
|
||||||
|
// (`H5S_extent_get_dims`), so changing them would also
|
||||||
|
// change the maximum the chunk index was built with — the
|
||||||
|
// Fixed Array linearises chunks by it — and move every
|
||||||
|
// existing chunk. Record the maximum libhdf5 would have
|
||||||
|
// written, the dimensions before this resize, so the index
|
||||||
|
// keeps its layout and the dataset can grow back to them.
|
||||||
|
let body_len = first + dims.len() * ls;
|
||||||
|
if body.len() < body_len || body[2] & !0x01 != 0 {
|
||||||
|
return Err(Error::Unsupported("dataspace message layout".into()));
|
||||||
|
}
|
||||||
|
let mut new_body = body[..first].to_vec();
|
||||||
|
new_body[2] |= 0x01;
|
||||||
|
new_body.extend_from_slice(&dims_bytes);
|
||||||
|
let at = new_body.len();
|
||||||
|
new_body.resize(at + dims.len() * ls, 0);
|
||||||
|
for (d, &n) in dims.iter().enumerate() {
|
||||||
|
put_uint(&mut new_body[at + d * ls..], n, img.ls);
|
||||||
|
}
|
||||||
|
let (flags, corder) = (hdr.msgs[i].flags, hdr.msgs[i].corder);
|
||||||
|
hdr.delete(img, i)?;
|
||||||
|
hdr.insert(img, MSG_DATASPACE, flags, &new_body, corder)?;
|
||||||
|
}
|
||||||
let fill = fill_info(img, &hdr)?;
|
let fill = fill_info(img, &hdr)?;
|
||||||
hdr.finish(img)?;
|
hdr.finish(img)?;
|
||||||
let expand = shape.iter().zip(&dims).any(|(n, o)| n > o);
|
let expand = shape.iter().zip(&dims).any(|(n, o)| n > o);
|
||||||
|
|||||||
@@ -446,6 +446,38 @@ impl File {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// A reader over an already open file (the file itself, not whatever
|
||||||
|
/// its path names now): mapped with the `mmap` feature, else read into
|
||||||
|
/// memory. `base_dir` resolves external Virtual Dataset sources.
|
||||||
|
pub(crate) fn from_std_file(
|
||||||
|
file: std::fs::File,
|
||||||
|
base_dir: Option<std::path::PathBuf>,
|
||||||
|
) -> Result<Self, Error> {
|
||||||
|
#[cfg(feature = "mmap")]
|
||||||
|
let mut f = {
|
||||||
|
let reader = clawhdf5_io::MmapReader::from_file(file).map_err(Error::Io)?;
|
||||||
|
let (data, superblock) = FileData::new(Backing::Mmap(reader))?;
|
||||||
|
Self {
|
||||||
|
data,
|
||||||
|
superblock,
|
||||||
|
chunk_cache: ChunkCache::new(),
|
||||||
|
base_dir: None,
|
||||||
|
vds_resolver: None,
|
||||||
|
}
|
||||||
|
};
|
||||||
|
#[cfg(not(feature = "mmap"))]
|
||||||
|
let mut f = {
|
||||||
|
use std::io::{Read, Seek, SeekFrom};
|
||||||
|
let mut file = file;
|
||||||
|
let mut bytes = Vec::new();
|
||||||
|
file.seek(SeekFrom::Start(0)).map_err(Error::Io)?;
|
||||||
|
file.read_to_end(&mut bytes).map_err(Error::Io)?;
|
||||||
|
Self::from_bytes(bytes)?
|
||||||
|
};
|
||||||
|
f.base_dir = base_dir;
|
||||||
|
Ok(f)
|
||||||
|
}
|
||||||
|
|
||||||
/// Open an HDF5 file by reading it entirely into memory.
|
/// Open an HDF5 file by reading it entirely into memory.
|
||||||
///
|
///
|
||||||
/// This is the pre-mmap behaviour and is useful when memory-mapping is
|
/// This is the pre-mmap behaviour and is useful when memory-mapping is
|
||||||
|
|||||||
@@ -0,0 +1,287 @@
|
|||||||
|
//! `FileEditor::resize` on chunked datasets whose dataspace records no
|
||||||
|
//! maximum dimensions, as clawhdf5's writer stored a dataset created without
|
||||||
|
//! a `maxshape` up to 2.7.0 (`fixtures/chunked_no_maxshape_v2_7_0.h5`). libhdf5 never writes such a dataspace (`H5S_set_extent_simple`
|
||||||
|
//! always records the maximum, equal to the dimensions when none is given),
|
||||||
|
//! and its Fixed Array chunk index linearises chunks by the maximum
|
||||||
|
//! dimensions. The editor therefore records the maximum libhdf5 would have
|
||||||
|
//! written (the dimensions the index was built with) before it changes the
|
||||||
|
//! current ones, so existing chunks stay where the index put them and the
|
||||||
|
//! dataset can grow back to its original extent.
|
||||||
|
//!
|
||||||
|
//! The writer now records the maximum too, so h5py can resize what it writes.
|
||||||
|
//!
|
||||||
|
//! Checked against a model of the expected values, with our reader and with
|
||||||
|
//! h5py (`CLAWHDF5_PYTHON`; skipped without it unless
|
||||||
|
//! `CLAWHDF5_REQUIRE_INTEROP=1`), on files clawhdf5 (old and new) and h5py
|
||||||
|
//! wrote.
|
||||||
|
|
||||||
|
use std::path::{Path, PathBuf};
|
||||||
|
use std::process::Command;
|
||||||
|
|
||||||
|
use clawhdf5::{Error, File, FileBuilder, FileEditor};
|
||||||
|
|
||||||
|
fn python() -> String {
|
||||||
|
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn h5py_ok() -> bool {
|
||||||
|
let ok = Command::new(python())
|
||||||
|
.args(["-c", "import h5py, numpy"])
|
||||||
|
.output()
|
||||||
|
.is_ok_and(|o| o.status.success());
|
||||||
|
if !ok {
|
||||||
|
assert!(
|
||||||
|
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||||
|
"CLAWHDF5_REQUIRE_INTEROP=1 but h5py/numpy is not available"
|
||||||
|
);
|
||||||
|
eprintln!("SKIP (h5py part): h5py/numpy not available");
|
||||||
|
}
|
||||||
|
ok
|
||||||
|
}
|
||||||
|
|
||||||
|
fn py(script: &str) -> String {
|
||||||
|
let o = Command::new(python())
|
||||||
|
.args(["-c", script])
|
||||||
|
.output()
|
||||||
|
.expect("run python");
|
||||||
|
assert!(
|
||||||
|
o.status.success(),
|
||||||
|
"python failed:\n{script}\nSTDERR: {}",
|
||||||
|
String::from_utf8_lossy(&o.stderr)
|
||||||
|
);
|
||||||
|
String::from_utf8_lossy(&o.stdout).trim().to_string()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Row-major values of a 2-D model after resizing `data` (shape `old`) to
|
||||||
|
/// `new`: kept elements keep their values, new ones are 0 (the fill value).
|
||||||
|
fn resized(data: &[f32], old: [u64; 2], new: [u64; 2]) -> Vec<f32> {
|
||||||
|
let mut out = vec![0f32; (new[0] * new[1]) as usize];
|
||||||
|
for r in 0..old[0].min(new[0]) {
|
||||||
|
for c in 0..old[1].min(new[1]) {
|
||||||
|
out[(r * new[1] + c) as usize] = data[(r * old[1] + c) as usize];
|
||||||
|
}
|
||||||
|
}
|
||||||
|
out
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Our reader and (when available) h5py read `expect` at `shape`.
|
||||||
|
fn check(path: &Path, name: &str, shape: [u64; 2], expect: &[f32], with_h5py: bool) {
|
||||||
|
let f = File::open(path).unwrap();
|
||||||
|
let d = f.dataset(name).unwrap();
|
||||||
|
assert_eq!(d.shape().unwrap(), shape);
|
||||||
|
assert_eq!(
|
||||||
|
d.read_f32().unwrap(),
|
||||||
|
expect,
|
||||||
|
"{name}: our reader at {shape:?}"
|
||||||
|
);
|
||||||
|
if with_h5py {
|
||||||
|
let got = py(&format!(
|
||||||
|
"import h5py, numpy as np\n\
|
||||||
|
with h5py.File({p:?}, 'r') as f:\n\
|
||||||
|
\x20 d = f[{name:?}][()]\n\
|
||||||
|
print(d.shape, ','.join(repr(float(x)) for x in d.ravel()))",
|
||||||
|
p = path.to_str().unwrap()
|
||||||
|
));
|
||||||
|
let want = format!(
|
||||||
|
"({}, {}) {}",
|
||||||
|
shape[0],
|
||||||
|
shape[1],
|
||||||
|
expect
|
||||||
|
.iter()
|
||||||
|
.map(|x| format!("{:?}", f64::from(*x)))
|
||||||
|
.collect::<Vec<_>>()
|
||||||
|
.join(",")
|
||||||
|
);
|
||||||
|
assert_eq!(got, want.trim(), "{name}: h5py at {shape:?}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Resize `name` (20 x 20, values 0..400) through a sequence of shrinks,
|
||||||
|
/// zero extents and growth back, checking every step.
|
||||||
|
fn run(path: &Path, name: &str, with_h5py: bool) {
|
||||||
|
let orig: Vec<f32> = (0..400).map(|i| i as f32).collect();
|
||||||
|
let mut shape = [20u64, 20];
|
||||||
|
let mut data = orig.clone();
|
||||||
|
check(path, name, shape, &data, with_h5py);
|
||||||
|
for next in [
|
||||||
|
[15, 15],
|
||||||
|
[3, 2],
|
||||||
|
[20, 20],
|
||||||
|
[1, 1],
|
||||||
|
[1, 0],
|
||||||
|
[0, 0],
|
||||||
|
[7, 20],
|
||||||
|
[20, 13],
|
||||||
|
[20, 20],
|
||||||
|
] {
|
||||||
|
let mut ed = FileEditor::open(path).unwrap();
|
||||||
|
ed.resize(name, &next).unwrap();
|
||||||
|
drop(ed);
|
||||||
|
data = resized(&data, shape, next);
|
||||||
|
shape = next;
|
||||||
|
check(path, name, shape, &data, with_h5py);
|
||||||
|
}
|
||||||
|
// The maximum is the extent the dataset was created with.
|
||||||
|
let mut ed = FileEditor::open(path).unwrap();
|
||||||
|
assert!(matches!(
|
||||||
|
ed.resize(name, &[21, 20]),
|
||||||
|
Err(Error::InvalidArgument(_))
|
||||||
|
));
|
||||||
|
drop(ed);
|
||||||
|
let f = File::open(path).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
f.dataset(name).unwrap().max_dimensions().unwrap(),
|
||||||
|
Some(vec![20, 20])
|
||||||
|
);
|
||||||
|
drop(f);
|
||||||
|
// A shrink keeps the values it keeps.
|
||||||
|
let mut ed = FileEditor::open(path).unwrap();
|
||||||
|
let vals: Vec<f32> = orig.iter().map(|v| v + 0.5).collect();
|
||||||
|
ed.write_values(name, &clawhdf5::Selection::All, &vals)
|
||||||
|
.unwrap();
|
||||||
|
ed.resize(name, &[15, 15]).unwrap();
|
||||||
|
drop(ed);
|
||||||
|
check(
|
||||||
|
path,
|
||||||
|
name,
|
||||||
|
[15, 15],
|
||||||
|
&resized(&vals, [20, 20], [15, 15]),
|
||||||
|
with_h5py,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
fn fixture(dir: &Path, name: &str) -> PathBuf {
|
||||||
|
let path = dir.join(name);
|
||||||
|
std::fs::copy(
|
||||||
|
Path::new(env!("CARGO_MANIFEST_DIR"))
|
||||||
|
.join("tests/fixtures")
|
||||||
|
.join(name),
|
||||||
|
&path,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
path
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Files clawhdf5 2.7.0 wrote: a Fixed Array index (Single Chunk for `s`)
|
||||||
|
/// and a dataspace with no maximum. Shrinking scrambled the values
|
||||||
|
/// (released in 2.7.0's `FileEditor`, PR #18).
|
||||||
|
#[test]
|
||||||
|
fn resize_without_stored_maxshape_keeps_values() {
|
||||||
|
let with_h5py = h5py_ok();
|
||||||
|
for name in ["d", "z"] {
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
|
||||||
|
run(&path, name, with_h5py);
|
||||||
|
}
|
||||||
|
// A single-chunk dataset and an empty one keep their extents as maxima.
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
|
||||||
|
let mut ed = FileEditor::open(&path).unwrap();
|
||||||
|
ed.resize("s", &[2, 4]).unwrap();
|
||||||
|
ed.resize("s", &[3, 4]).unwrap();
|
||||||
|
assert!(matches!(
|
||||||
|
ed.resize("s", &[4, 4]),
|
||||||
|
Err(Error::InvalidArgument(_))
|
||||||
|
));
|
||||||
|
assert!(matches!(
|
||||||
|
ed.resize("e", &[1, 5]),
|
||||||
|
Err(Error::InvalidArgument(_))
|
||||||
|
));
|
||||||
|
ed.resize("e", &[0, 3]).unwrap();
|
||||||
|
drop(ed);
|
||||||
|
let f = File::open(&path).unwrap();
|
||||||
|
let s = f.dataset("s").unwrap();
|
||||||
|
assert_eq!(s.max_dimensions().unwrap(), Some(vec![3, 4]));
|
||||||
|
let mut want: Vec<i32> = (0..12).collect();
|
||||||
|
want[8..].fill(0);
|
||||||
|
assert_eq!(s.read_i32().unwrap(), want);
|
||||||
|
assert_eq!(
|
||||||
|
f.dataset("e").unwrap().max_dimensions().unwrap(),
|
||||||
|
Some(vec![0, 5])
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
fn written(path: &Path, deflate: bool) {
|
||||||
|
let data: Vec<f32> = (0..400).map(|i| i as f32).collect();
|
||||||
|
let mut b = FileBuilder::new();
|
||||||
|
let d = b
|
||||||
|
.create_dataset("d")
|
||||||
|
.with_f32_data(&data)
|
||||||
|
.with_shape(&[20, 20])
|
||||||
|
.with_chunks(&[6, 6]);
|
||||||
|
if deflate {
|
||||||
|
d.with_deflate(4);
|
||||||
|
}
|
||||||
|
b.write(path).unwrap();
|
||||||
|
}
|
||||||
|
|
||||||
|
/// clawhdf5's writer now records the maximum, as libhdf5 does.
|
||||||
|
#[test]
|
||||||
|
fn resize_file_written_without_maxshape_keeps_values() {
|
||||||
|
let with_h5py = h5py_ok();
|
||||||
|
for deflate in [false, true] {
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = dir.path().join("cw.h5");
|
||||||
|
written(&path, deflate);
|
||||||
|
let f = File::open(&path).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
f.dataset("d").unwrap().max_dimensions().unwrap(),
|
||||||
|
Some(vec![20, 20])
|
||||||
|
);
|
||||||
|
drop(f);
|
||||||
|
run(&path, "d", with_h5py);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// h5py resizing a file clawhdf5 wrote without a maxshape keeps its values
|
||||||
|
/// (it scrambled them while the writer recorded no maximum).
|
||||||
|
#[test]
|
||||||
|
fn h5py_resizes_what_clawhdf5_writes() {
|
||||||
|
if !h5py_ok() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
for deflate in [false, true] {
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = dir.path().join("cw.h5");
|
||||||
|
written(&path, deflate);
|
||||||
|
let out = py(&format!(
|
||||||
|
"import h5py, numpy as np\n\
|
||||||
|
exp = np.arange(400, dtype='f4').reshape(20, 20)\n\
|
||||||
|
with h5py.File({p:?}, 'r+') as f:\n\
|
||||||
|
\x20 f['d'].resize((15, 15))\n\
|
||||||
|
\x20 ok = np.array_equal(f['d'][()], exp[:15, :15])\n\
|
||||||
|
\x20 f['d'].resize((20, 20))\n\
|
||||||
|
\x20 back = f['d'][()]\n\
|
||||||
|
want = np.zeros((20, 20), 'f4'); want[:15, :15] = exp[:15, :15]\n\
|
||||||
|
print(ok and np.array_equal(back, want))",
|
||||||
|
p = path.to_str().unwrap()
|
||||||
|
));
|
||||||
|
assert_eq!(out, "True");
|
||||||
|
let mut want = vec![0f32; 400];
|
||||||
|
for r in 0..15 {
|
||||||
|
for c in 0..15 {
|
||||||
|
want[r * 20 + c] = (r * 20 + c) as f32;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
check(&path, "d", [20, 20], &want, false);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// h5py's files record the maximum; the same sequence must hold.
|
||||||
|
#[test]
|
||||||
|
fn resize_h5py_file_without_maxshape_keeps_values() {
|
||||||
|
if !h5py_ok() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
for libver in ["earliest", "v110", "latest"] {
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = dir.path().join("hp.h5");
|
||||||
|
py(&format!(
|
||||||
|
"import h5py, numpy as np\n\
|
||||||
|
with h5py.File({p:?}, 'w', libver=({libver:?}, 'latest')) as f:\n\
|
||||||
|
\x20 f.create_dataset('d', data=np.arange(400, dtype='f4').reshape(20, 20), chunks=(6, 6))",
|
||||||
|
p = path.to_str().unwrap()
|
||||||
|
));
|
||||||
|
run(&path, "d", true);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -162,3 +162,59 @@ fn shrink_then_grow_reads_fill() {
|
|||||||
raw[..4].fill(0.5);
|
raw[..4].fill(0.5);
|
||||||
assert_eq!(f.dataset("raw").unwrap().read_f64().unwrap(), raw);
|
assert_eq!(f.dataset("raw").unwrap().read_f64().unwrap(), raw);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The editor plans every edit from the file it holds open, never by
|
||||||
|
/// re-opening its path: with the path renamed away and another file put in
|
||||||
|
/// its place, edits go to the held file, planned from its own metadata,
|
||||||
|
/// and the file now at the path is untouched (planning from it and writing
|
||||||
|
/// into the held file corrupted the held one).
|
||||||
|
#[test]
|
||||||
|
fn edits_go_to_the_file_held_not_the_path() {
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let a = dir.path().join("a.h5");
|
||||||
|
let b = dir.path().join("b.h5");
|
||||||
|
let mut fb = FileBuilder::new();
|
||||||
|
fb.create_dataset("x")
|
||||||
|
.with_i32_data(&[0; 10])
|
||||||
|
.with_shape(&[10]);
|
||||||
|
fb.create_dataset("big")
|
||||||
|
.with_f64_data(&[1.5; 5000])
|
||||||
|
.with_shape(&[5000]);
|
||||||
|
fb.set_attr("title", AttrValue::String("a".into()));
|
||||||
|
fb.write(&a).unwrap();
|
||||||
|
let mut fb = FileBuilder::new();
|
||||||
|
fb.create_dataset("pad")
|
||||||
|
.with_f64_data(&[2.5; 3000])
|
||||||
|
.with_shape(&[3000]);
|
||||||
|
fb.create_dataset("x")
|
||||||
|
.with_i32_data(&[500; 10])
|
||||||
|
.with_shape(&[10]);
|
||||||
|
fb.write(&b).unwrap();
|
||||||
|
|
||||||
|
let mut ed = FileEditor::open(&a).unwrap();
|
||||||
|
assert!(ed.path().is_absolute());
|
||||||
|
let moved = dir.path().join("moved.h5");
|
||||||
|
std::fs::rename(&a, &moved).unwrap();
|
||||||
|
std::fs::rename(&b, &a).unwrap();
|
||||||
|
let other = std::fs::read(&a).unwrap();
|
||||||
|
|
||||||
|
ed.write_values("x", &Selection::All, &[7i32; 10]).unwrap();
|
||||||
|
let vals: Vec<f64> = (0..50).map(f64::from).collect();
|
||||||
|
ed.set_attr("/", "note", &AttrValue::F64Array(vals.clone()))
|
||||||
|
.unwrap();
|
||||||
|
// The editor's own reader sees the held file.
|
||||||
|
let r = ed.reader().unwrap();
|
||||||
|
assert_eq!(r.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
|
||||||
|
assert!(r.dataset("pad").is_err());
|
||||||
|
drop(r);
|
||||||
|
drop(ed);
|
||||||
|
|
||||||
|
assert!(
|
||||||
|
std::fs::read(&a).unwrap() == other,
|
||||||
|
"the file at the path changed"
|
||||||
|
);
|
||||||
|
let f = File::open(&moved).unwrap();
|
||||||
|
assert_eq!(f.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
|
||||||
|
assert_eq!(f.dataset("big").unwrap().read_f64().unwrap(), [1.5; 5000]);
|
||||||
|
assert!(matches!(f.root().attr("note").unwrap(), Some(AttrValue::F64Array(v)) if v == vals));
|
||||||
|
}
|
||||||
|
|||||||
Binary file not shown.
Binary file not shown.
@@ -966,7 +966,10 @@ fn skipped_optional_filters_are_masked_as_libhdf5_masks_them() {
|
|||||||
|
|
||||||
/// Files whose chunks all compress are written exactly as before optional
|
/// Files whose chunks all compress are written exactly as before optional
|
||||||
/// filters could be skipped: every mask is 0 and nothing else changed. The
|
/// filters could be skipped: every mask is 0 and nothing else changed. The
|
||||||
/// hashes are of the files the writer produced before that change.
|
/// hashes are of the files the writer produced before that change, except
|
||||||
|
/// that a chunked dataset without a maxshape now records its maximum
|
||||||
|
/// dimensions (8 bytes per dimension; `lzf_fixed`, `lzf_single`, and
|
||||||
|
/// `blosc_fixed`).
|
||||||
#[cfg(feature = "lzf")]
|
#[cfg(feature = "lzf")]
|
||||||
#[test]
|
#[test]
|
||||||
fn files_whose_chunks_all_compress_are_unchanged() {
|
fn files_whose_chunks_all_compress_are_unchanged() {
|
||||||
@@ -985,7 +988,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
|
|||||||
.with_chunks(&[500])
|
.with_chunks(&[500])
|
||||||
.with_lzf();
|
.with_lzf();
|
||||||
},
|
},
|
||||||
(3965, 449169442),
|
(3973, 452644487),
|
||||||
),
|
),
|
||||||
(
|
(
|
||||||
"lzf_ea_noshuffle",
|
"lzf_ea_noshuffle",
|
||||||
@@ -1017,7 +1020,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
|
|||||||
.with_chunks(&[3000])
|
.with_chunks(&[3000])
|
||||||
.with_lzf();
|
.with_lzf();
|
||||||
},
|
},
|
||||||
(546, 690805477),
|
(554, 1394027497),
|
||||||
),
|
),
|
||||||
];
|
];
|
||||||
#[cfg(feature = "blosc")]
|
#[cfg(feature = "blosc")]
|
||||||
@@ -1029,7 +1032,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
|
|||||||
.with_chunks(&[1024])
|
.with_chunks(&[1024])
|
||||||
.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
|
.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
|
||||||
},
|
},
|
||||||
(2776, 4278611376),
|
(2784, 1180133244),
|
||||||
));
|
));
|
||||||
for (name, build, want) in &cases {
|
for (name, build, want) in &cases {
|
||||||
let mut fb = clawhdf5::FileBuilder::new();
|
let mut fb = clawhdf5::FileBuilder::new();
|
||||||
|
|||||||
@@ -12,6 +12,10 @@ done on branch `feat/p3-m5-swmr-reader`, with its own design in
|
|||||||
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
|
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
|
||||||
next. Every count in §1–§2 was
|
next. Every count in §1–§2 was
|
||||||
|
|
||||||
|
object stores) and URLs in `h5rs` (see the M3 status below); the Python
|
||||||
|
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
|
||||||
|
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
|
||||||
|
|
||||||
change. Progress: M1, first part (the `Storage` trait and the metadata
|
change. Progress: M1, first part (the `Storage` trait and the metadata
|
||||||
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
||||||
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
||||||
@@ -477,7 +481,7 @@ fast path within benchmark noise.
|
|||||||
a request counter exposed for tests and users.
|
a request counter exposed for tests and users.
|
||||||
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
|
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
|
||||||
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
|
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
|
||||||
Python bindings, with these choices:
|
Python bindings (done 2026-09-27, below), with these choices:
|
||||||
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
|
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
|
||||||
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
|
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
|
||||||
sits below the facade.
|
sits below the facade.
|
||||||
@@ -514,6 +518,17 @@ fast path within benchmark noise.
|
|||||||
same work without a cache: 141 936 requests. Per file: A lists in 2
|
same work without a cache: 141 936 requests. Per file: A lists in 2
|
||||||
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
|
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
|
||||||
6.4 MB: 35 001 object headers spread over the file).
|
6.4 MB: 35 001 object headers spread over the file).
|
||||||
|
- *Status 2026-09-27, Python bindings:* done on branch
|
||||||
|
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
|
||||||
|
`File.open_url(url, **options)` (cache and HTTP options) go through
|
||||||
|
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
|
||||||
|
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
|
||||||
|
parsing (path lookups, headers, attributes, listings, the global heap)
|
||||||
|
moved from `File::as_bytes` to `File::storage()` and the `*_in`
|
||||||
|
functions, and every file access, metadata included, runs with the GIL
|
||||||
|
released. Checked by running the whole read-vs-h5py suite over an
|
||||||
|
in-process range server (1 MiB and 1 KiB blocks) and by request counts
|
||||||
|
in `crates/clawhdf5-py/tests/test_remote.py`.
|
||||||
|
|
||||||
**M4 — wasm lazy loading (1–2 weeks).**
|
**M4 — wasm lazy loading (1–2 weeks).**
|
||||||
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
|
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
|
||||||
|
|||||||
+87
-6
@@ -30,6 +30,33 @@ h5py's SWMR reader). A file still being written is read with
|
|||||||
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
|
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
|
||||||
at its length at open and is not meant for files that change while open.
|
at its length at open and is not meant for files that change while open.
|
||||||
|
|
||||||
|
## Shrinking a chunked dataset with no recorded maximum scrambled it
|
||||||
|
|
||||||
|
**Status:** fixed 2026-09-27, before any release (`FileEditor::resize`
|
||||||
|
shipped on main in PR #18, a4c2ace). Files the writer produced before the
|
||||||
|
fix still lack the maximum; see *Existing files*.
|
||||||
|
|
||||||
|
clawhdf5's writer stored no maximum dimensions for a chunked dataset
|
||||||
|
created without a `maxshape` (libhdf5 always stores one, equal to the
|
||||||
|
dimensions when none is given). With none recorded, the maximum is the
|
||||||
|
current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index
|
||||||
|
places chunks by the maximum. `FileEditor::resize` changed only the
|
||||||
|
current dimensions, so a shrink moved every existing chunk and every
|
||||||
|
reader returned wrong values; after a shrink the dataset could not grow
|
||||||
|
back. Found by the review of the Python editing work.
|
||||||
|
|
||||||
|
**Fix:** before the first resize of such a dataset the editor records the
|
||||||
|
maximum libhdf5 would have written (the dimensions the index was built
|
||||||
|
with); the writer now records it for every chunked dataset. **Test:**
|
||||||
|
`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:**
|
||||||
|
datasets written before the fix have no recorded maximum. The fixed
|
||||||
|
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
|
||||||
|
does not** — it scrambles them the same way, and lets them grow past
|
||||||
|
their Fixed Array. Resize them once with a fixed `FileEditor` (a resize
|
||||||
|
to the same shape changes nothing; shrink and grow back) before letting
|
||||||
|
libhdf5 resize them. A dataset already shrunk by the unfixed
|
||||||
|
editor holds misplaced chunks; rewrite it from a good copy.
|
||||||
|
|
||||||
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
|
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
|
||||||
|
|
||||||
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to
|
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to
|
||||||
@@ -139,6 +166,57 @@ space it leaves is too small for its next, larger version.
|
|||||||
**No journal.** A crash while an edit patches existing structures can leave
|
**No journal.** A crash while an edit patches existing structures can leave
|
||||||
the file inconsistent; see the `FileEditor` documentation.
|
the file inconsistent; see the `FileEditor` documentation.
|
||||||
|
|
||||||
|
**Renamed files outside Linux.** Edits always go to the file the editor
|
||||||
|
opened. `FileEditor::reader()` (which the Python `'r+'` handle reads
|
||||||
|
through) reopens it through `/proc/self/fd` on Linux; elsewhere it reopens
|
||||||
|
the path, and after the path was renamed or replaced it fails (on Unix;
|
||||||
|
Windows cannot tell and would read whatever the path names).
|
||||||
|
|
||||||
|
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
|
||||||
|
|
||||||
|
**Status:** open (added 2026-09-27). The Python bindings edit through
|
||||||
|
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
|
||||||
|
before anything is written. On top of them:
|
||||||
|
|
||||||
|
- **No new or deleted objects:** `create_dataset`/`create_group` in an
|
||||||
|
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
|
||||||
|
`NotImplementedError` (the editor changes values, shapes and
|
||||||
|
attributes only). Mode `'a'` works on an existing file only.
|
||||||
|
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
|
||||||
|
whole elements of the same structured dtype are), HDF5 array-type
|
||||||
|
elements, variable-length data, strings padded with spaces or
|
||||||
|
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
|
||||||
|
ones, h5py's, are writable), compounds containing such strings, null
|
||||||
|
dataspaces, index-list writes of more than 2²² elements (write them
|
||||||
|
in slices), and boolean-mask keys (`ds[mask] = v`, and mask reads;
|
||||||
|
h5py supports both).
|
||||||
|
- **`str` attributes are fixed-length UTF-8**, where h5py writes
|
||||||
|
variable-length strings: h5py reads them back as `bytes`
|
||||||
|
(`numpy.bytes_`), not `str`.
|
||||||
|
- **Numeric conversion follows libhdf5's native-order results, not its
|
||||||
|
bugs.** Arrays are converted as libhdf5 converts them (integers
|
||||||
|
saturate, floats are truncated toward zero and clipped), checked value by
|
||||||
|
value against h5py 3.16 / HDF5 2.0 in
|
||||||
|
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
|
||||||
|
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
|
||||||
|
differs from h5py on purpose:
|
||||||
|
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
|
||||||
|
minimum or 2⁶³ depending on the type);
|
||||||
|
- when the dataset or the array is not in native byte order, libhdf5's
|
||||||
|
"soft" conversions store a float in (-1, 0) as the integer minimum and
|
||||||
|
wrap an unsigned value too large for the signed type of the same size
|
||||||
|
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
|
||||||
|
native order;
|
||||||
|
- libhdf5's native casts that are undefined in C: half floats into
|
||||||
|
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
|
||||||
|
integers (→ minimum), a float equal to the integer maximum rounded up
|
||||||
|
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
|
||||||
|
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
|
||||||
|
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
|
||||||
|
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
|
||||||
|
- **Each edit reopens the file** (a new memory map) so that reads see it;
|
||||||
|
reads from other threads wait while an edit is written.
|
||||||
|
|
||||||
## Selection reads that decode more than the selection
|
## Selection reads that decode more than the selection
|
||||||
|
|
||||||
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
|
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
|
||||||
@@ -867,9 +945,9 @@ cache, but:
|
|||||||
`read_*_zerocopy`) need the file in memory and answer
|
`read_*_zerocopy`) need the file in memory and answer
|
||||||
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
|
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
|
||||||
panics for such a file (`File::contiguous_bytes()` is the fallible form).
|
panics for such a file (`File::contiguous_bytes()` is the fallible form).
|
||||||
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
|
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
|
||||||
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
|
(`h5rs` and the Python bindings read through `File::storage`, and take
|
||||||
`remote` feature).
|
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
|
||||||
- The file's length is read once, at open: a growing file (SWMR) is not
|
- The file's length is read once, at open: a growing file (SWMR) is not
|
||||||
followed (milestone M5). A remote file is pinned at open, so one that
|
followed (milestone M5). A remote file is pinned at open, so one that
|
||||||
grows is `RemoteError::FileChanged`.
|
grows is `RemoteError::FileChanged`.
|
||||||
@@ -885,9 +963,12 @@ cache, but:
|
|||||||
**Status:** open (added 2026-09-26, milestone M3 of
|
**Status:** open (added 2026-09-26, milestone M3 of
|
||||||
`docs/design/range-reads.md`).
|
`docs/design/range-reads.md`).
|
||||||
|
|
||||||
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
|
- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
|
||||||
parses through `File::as_bytes`, which a remote file does not have; the
|
milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
|
||||||
wasm reader's `openUrl` is milestone M4.
|
the default wheel reads plain `http://` only: `https://` needs a wheel
|
||||||
|
built with `--features https` (rustls with ring, which compiles C), and
|
||||||
|
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
|
||||||
|
The Python tests run against an in-process `http.server` only.
|
||||||
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
||||||
The design's policy of using a paged file's page size as the block size
|
The design's policy of using a paged file's page size as the block size
|
||||||
is not implemented, and only the first block is read ahead.
|
is not implemented, and only the first block is read ahead.
|
||||||
|
|||||||
+9
-3
@@ -131,14 +131,17 @@ run_step "cargo clippy (h5rs remote)" cargo clippy \
|
|||||||
# js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C.
|
# js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C.
|
||||||
# clawhdf5-remote is checked by default (plain HTTP) and with its
|
# clawhdf5-remote is checked by default (plain HTTP) and with its
|
||||||
# object-store feature, and h5rs with URL support (remote); the https
|
# object-store feature, and h5rs with URL support (remote); the https
|
||||||
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in.
|
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. The
|
||||||
|
# Python bindings (clawhdf5-py, remote reads over plain HTTP) are checked too:
|
||||||
|
# their https/s3/gcs/azure features are opt-in for the same reason.
|
||||||
no_c_in_default_build() {
|
no_c_in_default_build() {
|
||||||
local entry crate features found=0
|
local entry crate features found=0
|
||||||
for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \
|
for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \
|
||||||
clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \
|
clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \
|
||||||
clawhdf5-tools \
|
clawhdf5-tools \
|
||||||
clawhdf5-wasm \
|
clawhdf5-wasm \
|
||||||
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote; do
|
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote \
|
||||||
|
clawhdf5-py; do
|
||||||
crate=${entry%%:*}
|
crate=${entry%%:*}
|
||||||
features=()
|
features=()
|
||||||
[ "$entry" != "$crate" ] && features=(--features "${entry#*:}")
|
[ "$entry" != "$crate" ] && features=(--features "${entry#*:}")
|
||||||
@@ -269,7 +272,10 @@ python_package() {
|
|||||||
-i "$PYTHON" \
|
-i "$PYTHON" \
|
||||||
--out "$out/wheel" || return 1
|
--out "$out/wheel" || return 1
|
||||||
"$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1
|
"$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1
|
||||||
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
|
# The editing tests run `h5rs check` on every file they edit.
|
||||||
|
cargo build -q -p clawhdf5-tools || return 1
|
||||||
|
CLAWHDF5_H5RS="${CARGO_TARGET_DIR:-$root/target}/debug/h5rs" \
|
||||||
|
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
|
||||||
"$root/crates/clawhdf5-py/tests"
|
"$root/crates/clawhdf5-py/tests"
|
||||||
}
|
}
|
||||||
if "$PYTHON" -m maturin --version >/dev/null 2>&1 && "$PYTHON" -c "import pytest" >/dev/null 2>&1; then
|
if "$PYTHON" -m maturin --version >/dev/null 2>&1 && "$PYTHON" -c "import pytest" >/dev/null 2>&1; then
|
||||||
|
|||||||
Reference in New Issue
Block a user