py: remote files (clawhdf5.File(url), File.open_url) through File::storage()

The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.

- Every object of a File now shares one handle (src/handle.rs) that
  runs all file access, metadata included, with the GIL released and
  parses through File::storage() and the clawhdf5_format *_in functions.
  Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
  clawhdf5_remote::storage_for_url (read-only; another mode is a
  ValueError). File.open_url(url, **options) takes the cache and HTTP
  options (block_size, cache_size, headers, retries, timeout,
  allow_full_download, max_full_download, require_validator,
  max_redirects, max_parallel); File.remote_stats gives the block
  cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
  s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
  ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
  OSError, never KeyError/ValueError and never data; `key in group`
  raises it instead of answering False.

Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:40:44 -05:00
co-authored by Claude Opus 5.5
parent a4c2aced55
commit 910d81904c
18 changed files with 1146 additions and 299 deletions
+30
View File
@@ -2,6 +2,36 @@
## Unreleased ## Unreleased
### Python bindings: remote files (2026-09-27)
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
through the block cache, the whole read API (groups, attributes, every
dataset type and index the local reader handles). A URL is any
`scheme://…`; a remote file is read-only (another mode is a
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
`allow_full_download`, `max_full_download`, `require_validator`,
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
cache's counters. The default wheel builds plain HTTP only (no C: rustls
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
check now covers `clawhdf5-py`.
- **Every read parses through `File::storage()`** instead of
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
group listings, variable-length data through the global heap), inside one
shared file handle that releases the GIL for all file access, not only
dataset reads: a read waiting on the network lets other Python threads
run. A failed read of the storage (a network error, a file changed on
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
data; `key in group` raises it instead of answering `False`.
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
blocks and 1 KiB blocks) against a range-capable `http.server` in the
test process, plus `tests/test_remote.py`: request counts of a small
read, cache hits, a server without `Range` support (refused, or a
whole download when allowed), a file changed on the server, a server
that hangs up, 16 threads on one remote file, and a thread that keeps
running while a read waits on 0.2 s requests.
### Range reads, milestone M3: remote files (2026-09-26) ### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives - **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by a `clawhdf5::File` (through `File::open_storage`) that reads the file by
+12
View File
@@ -533,8 +533,20 @@ with clawhdf5.File("data.h5", "r") as f:
records = f["table"] # compound -> numpy structured array records = f["table"] # compound -> numpy structured array
ids = records["id"] # one field ids = records["id"] # one field
# A file on a web server: range requests through a block cache, nothing
# downloaded up front; the same read API. The GIL is released while waiting.
with clawhdf5.File("http://data.example.org/run42.h5") as f:
first = f["group/temperatures"][0]
f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024,
headers={"Authorization": "Bearer ..."})
``` ```
The default build reads `http://` URLs only; build with
`maturin develop --release --features https` (rustls with ring, which
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
object-store URLs.
Reads cover integers and IEEE floats of every width in either byte order, Reads cover integers and IEEE floats of every width in either byte order,
`bool`, enums, complex, fixed and variable-length strings, variable-length `bool`, enums, complex, fixed and variable-length strings, variable-length
sequences, opaque, HDF5 array types and compounds; other types (references, sequences, opaque, HDF5 array types and compounds; other types (references,
+9
View File
@@ -17,11 +17,20 @@ crate-type = ["cdylib", "rlib"]
[dependencies] [dependencies]
clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" } clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" }
clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" } clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
# Remote files (`clawhdf5.File(url)`): plain HTTP by default, which builds
# no C. HTTPS and the object stores are opt-in features below.
clawhdf5-remote = { path = "../clawhdf5-remote", version = "2.7.0" }
pyo3 = "0.29" pyo3 = "0.29"
numpy = "0.29" numpy = "0.29"
[features] [features]
extension-module = ["pyo3/extension-module"] extension-module = ["pyo3/extension-module"]
# https:// URLs (rustls with ring, which compiles C and assembly).
https = ["clawhdf5-remote/https"]
# s3://, gs://, az:// URLs (object_store; its cloud clients build aws-lc-rs, C).
s3 = ["clawhdf5-remote/s3"]
gcs = ["clawhdf5-remote/gcs"]
azure = ["clawhdf5-remote/azure"]
[package.metadata.docs.rs] [package.metadata.docs.rs]
features = [] features = []
+34 -1
View File
@@ -61,6 +61,36 @@ with clawhdf5.File("data.h5", "r") as f:
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null - Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
dataspace (h5py's `Empty`). dataspace (h5py's `Empty`).
## Remote files
A URL instead of a path reads the file where it is, through
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
```python
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
```
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise `OSError` instead of mixing versions.
Network failures are `OSError` too.
- Remote files are read-only.
- Schemes: the default build (no C) reads `http://`. `https://` needs
`maturin develop --release --features https` (rustls with ring, which
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
`azure` features (credentials from the environment; aws-lc-rs, C).
## Writing ## Writing
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array, `clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
@@ -76,7 +106,10 @@ pytest crates/clawhdf5-py/tests
``` ```
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py `tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
writes. `scripts/ci-test.sh` builds the wheel and runs these in CI. writes, opened locally and over HTTP (an in-process range server,
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
failures, the GIL). `scripts/ci-test.sh` builds the wheel and runs these
in CI.
## License ## License
+23 -19
View File
@@ -8,13 +8,14 @@ use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple}; use pyo3::types::{PyList, PyTuple};
use crate::convert::{Converter, Elements, resolve_vl}; use crate::convert::{Converter, Elements, resolve_vl};
use crate::handle::Handle;
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value}; use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value};
/// Backing storage for attributes. /// Backing storage for attributes.
enum AttrsInner { enum AttrsInner {
/// Attributes of an object in a file opened for reading, sorted by name. /// Attributes of an object in a file opened for reading, sorted by name.
Read { Read {
file: Arc<clawhdf5_rs::File>, handle: Arc<Handle>,
attrs: Vec<AttributeMessage>, attrs: Vec<AttributeMessage>,
}, },
/// Writable attribute list shared with a parent (PyFile or PyGroup). /// Writable attribute list shared with a parent (PyFile or PyGroup).
@@ -36,10 +37,15 @@ pub struct PyAttrs {
impl PyAttrs { impl PyAttrs {
/// The attributes of the object at `addr` (whose path is `path`) in a /// The attributes of the object at `addr` (whose path is `path`) in a
/// file opened for reading. /// file opened for reading.
pub(crate) fn read(file: Arc<clawhdf5_rs::File>, addr: u64, path: &str) -> PyResult<Self> { pub(crate) fn read(
let attrs = node::attributes(&file, addr, path)?; py: Python<'_>,
handle: Arc<Handle>,
addr: u64,
path: &str,
) -> PyResult<Self> {
let attrs = handle.with(py, |f| node::attributes(f, addr, path))?;
Ok(Self { Ok(Self {
inner: AttrsInner::Read { file, attrs }, inner: AttrsInner::Read { handle, attrs },
}) })
} }
@@ -55,8 +61,8 @@ impl PyAttrs {
impl PyAttrs { impl PyAttrs {
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> { fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
match &self.inner { match &self.inner {
AttrsInner::Read { file, attrs } => match attrs.iter().find(|a| a.name == key) { AttrsInner::Read { handle, attrs } => match attrs.iter().find(|a| a.name == key) {
Some(attr) => Ok(attr_to_py(py, file, attr)?.unbind()), Some(attr) => Ok(attr_to_py(py, handle, attr)?.unbind()),
None => Err(PyKeyError::new_err(format!( None => Err(PyKeyError::new_err(format!(
"Can't open attribute (can't locate attribute: '{key}')" "Can't open attribute (can't locate attribute: '{key}')"
))), ))),
@@ -146,9 +152,9 @@ impl PyAttrs {
/// Return attribute values as a list. /// Return attribute values as a list.
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let vals: Vec<Py<PyAny>> = match &self.inner { let vals: Vec<Py<PyAny>> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs AttrsInner::Read { handle, attrs } => attrs
.iter() .iter()
.map(|a| attr_to_py(py, file, a).map(Bound::unbind)) .map(|a| attr_to_py(py, handle, a).map(Bound::unbind))
.collect::<PyResult<_>>()?, .collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store AttrsInner::Write(store) => store
.lock() .lock()
@@ -167,9 +173,9 @@ impl PyAttrs {
/// Return attribute (key, value) pairs as a list of tuples. /// Return attribute (key, value) pairs as a list of tuples.
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner { let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs AttrsInner::Read { handle, attrs } => attrs
.iter() .iter()
.map(|a| Ok((a.name.clone(), attr_to_py(py, file, a)?.unbind()))) .map(|a| Ok((a.name.clone(), attr_to_py(py, handle, a)?.unbind())))
.collect::<PyResult<_>>()?, .collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store AttrsInner::Write(store) => store
.lock() .lock()
@@ -189,12 +195,11 @@ impl PyAttrs {
/// An attribute's value as h5py returns it. /// An attribute's value as h5py returns it.
fn attr_to_py<'py>( fn attr_to_py<'py>(
py: Python<'py>, py: Python<'py>,
file: &clawhdf5_rs::File, handle: &Handle,
attr: &AttributeMessage, attr: &AttributeMessage,
) -> PyResult<Bound<'py, PyAny>> { ) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| { crate::no_panic(|| {
let sb = file.superblock(); let conv = Converter::new(py, &attr.datatype, handle.offset_size)
let conv = Converter::new(py, &attr.datatype, sb.offset_size)
.map_err(|e| prefix_err(py, &attr.name, e))?; .map_err(|e| prefix_err(py, &attr.name, e))?;
if node::is_null(&attr.dataspace) { if node::is_null(&attr.dataspace) {
return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any()); return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any());
@@ -216,12 +221,11 @@ fn attr_to_py<'py>(
))); )));
} }
let raw = &attr.raw_data[..want]; let raw = &attr.raw_data[..want];
let file_data = file.as_bytes(); let (osz, lsz, unit) = (handle.offset_size, handle.length_size, conv.vl_unit);
let (osz, lsz, unit) = (sb.offset_size, sb.length_size, conv.vl_unit); let what = format!("attribute {}", attr.name);
Elements::Vl( Elements::Vl(handle.with(py, |f| {
py.detach(|| resolve_vl(file_data, raw, n, osz, lsz, unit)) resolve_vl(f.storage(), raw, n, osz, lsz, unit).map_err(|e| e.into_py(&what))
.map_err(|e| PyValueError::new_err(format!("attribute {}: {e}", attr.name)))?, })?)
)
} else { } else {
Elements::Bytes(attr.raw_data.clone()) Elements::Bytes(attr.raw_data.clone())
}; };
+49 -21
View File
@@ -17,6 +17,7 @@ use std::collections::HashMap;
use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder}; use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder};
use clawhdf5_format::global_heap::GlobalHeapCollection; use clawhdf5_format::global_heap::GlobalHeapCollection;
use clawhdf5_format::storage::Storage;
use numpy::PyArray1; use numpy::PyArray1;
use pyo3::exceptions::{PyTypeError, PyValueError}; use pyo3::exceptions::{PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
@@ -538,16 +539,16 @@ fn object_array<'py>(
/// their bytes: each element's stored length times `unit` (1 for strings, /// their bytes: each element's stored length times `unit` (1 for strings,
/// the base type's size for sequences). Pure Rust, so it runs without the /// the base type's size for sequences). Pure Rust, so it runs without the
/// GIL. /// GIL.
pub(crate) fn resolve_vl( pub(crate) fn resolve_vl<S: Storage + ?Sized>(
file_data: &[u8], file: &S,
raw: &[u8], raw: &[u8],
count: usize, count: usize,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
unit: usize, unit: usize,
) -> Result<Vec<Vec<u8>>, String> { ) -> Result<Vec<Vec<u8>>, VlError> {
let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size) let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size)
.map_err(|e| e.to_string())?; .map_err(|e| VlError::Invalid(e.to_string()))?;
let undefined = match offset_size { let undefined = match offset_size {
2 => 0xFFFF, 2 => 0xFFFF,
4 => 0xFFFF_FFFF, 4 => 0xFFFF_FFFF,
@@ -558,43 +559,70 @@ pub(crate) fn resolve_vl(
for vl in &refs { for vl in &refs {
if vl.collection_address == 0 || vl.collection_address == undefined { if vl.collection_address == 0 || vl.collection_address == undefined {
if vl.length != 0 { if vl.length != 0 {
return Err(format!( return Err(VlError::Invalid(format!(
"variable-length element of length {} has no heap address", "variable-length element of length {} has no heap address",
vl.length vl.length
)); )));
} }
out.push(Vec::new()); out.push(Vec::new());
continue; continue;
} }
let coll = match collections.entry(vl.collection_address) { let coll = match collections.entry(vl.collection_address) {
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(), std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
std::collections::hash_map::Entry::Vacant(e) => { std::collections::hash_map::Entry::Vacant(e) => e.insert(
let addr = usize::try_from(vl.collection_address) GlobalHeapCollection::parse_in(file, vl.collection_address, length_size)
.map_err(|_| "global heap address out of range".to_string())?; .map_err(VlError::from_format)?,
e.insert( ),
GlobalHeapCollection::parse(file_data, addr, length_size)
.map_err(|e| e.to_string())?,
)
}
}; };
let index = u16::try_from(vl.object_index) let index = u16::try_from(vl.object_index).map_err(|_| {
.map_err(|_| format!("global heap object index {} out of range", vl.object_index))?; VlError::Invalid(format!(
"global heap object index {} out of range",
vl.object_index
))
})?;
let obj = coll.get_object(index).ok_or_else(|| { let obj = coll.get_object(index).ok_or_else(|| {
format!( VlError::Invalid(format!(
"global heap object {index} not found in the collection at {}", "global heap object {index} not found in the collection at {}",
vl.collection_address vl.collection_address
) ))
})?; })?;
let need = (vl.length as usize) let need = (vl.length as usize)
.checked_mul(unit) .checked_mul(unit)
.ok_or("variable-length element too long")?; .ok_or_else(|| VlError::Invalid("variable-length element too long".into()))?;
if need > obj.data.len() { if need > obj.data.len() {
return Err(format!( return Err(VlError::Invalid(format!(
"variable-length element of {need} bytes in a {}-byte heap object", "variable-length element of {need} bytes in a {}-byte heap object",
obj.data.len() obj.data.len()
)); )));
} }
out.push(obj.data[..need].to_vec()); out.push(obj.data[..need].to_vec());
} }
Ok(out) Ok(out)
} }
/// Why variable-length elements could not be resolved.
#[derive(Debug)]
pub(crate) enum VlError {
/// Reading the file failed (a network error on a remote file).
Storage(String),
/// The references or the heap are not valid.
Invalid(String),
}
impl VlError {
fn from_format(e: clawhdf5_format::error::FormatError) -> Self {
match e {
clawhdf5_format::error::FormatError::Storage(_) => VlError::Storage(e.to_string()),
e => VlError::Invalid(e.to_string()),
}
}
/// As a Python exception, the message prefixed with `what`: a storage
/// failure is an `OSError`, anything else a `ValueError`.
pub(crate) fn into_py(self, what: &str) -> PyErr {
match self {
VlError::Storage(m) => pyo3::exceptions::PyOSError::new_err(format!("{what}: {m}")),
VlError::Invalid(m) => PyValueError::new_err(format!("{what}: {m}")),
}
}
}
+87 -57
View File
@@ -6,21 +6,53 @@
//! whole dataset instead); the //! whole dataset instead); the
//! bytes it returns become the numpy array's buffer without a copy (see //! bytes it returns become the numpy array's buffer without a copy (see
//! `convert`). All file access and decoding runs with the GIL released, so //! `convert`). All file access and decoding runs with the GIL released, so
//! Python threads reading the same or different datasets run in parallel. //! Python threads reading the same or different datasets run in parallel,
//! and a remote file's network reads never hold the GIL.
use std::sync::Arc; use std::sync::Arc;
use clawhdf5_format::datatype::Datatype; use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::object_header::ObjectHeader; use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_rs::File;
use pyo3::exceptions::{PyTypeError, PyValueError}; use pyo3::exceptions::{PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple}; use pyo3::types::{PyList, PyTuple};
use crate::attrs::PyAttrs; use crate::attrs::PyAttrs;
use crate::convert::{Converter, Elements, resolve_vl}; use crate::convert::{Converter, Elements, VlError, resolve_vl};
use crate::handle::Handle;
use crate::select::{self, Plan}; use crate::select::{self, Plan};
use crate::{PyEmpty, node, to_py_err}; use crate::{PyEmpty, node, to_py_err};
/// What opening a dataset reads from the file (without the GIL).
pub(crate) struct DatasetMeta {
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
shape: Option<Vec<u64>>,
chunks: Option<Vec<u64>>,
datatype: Datatype,
}
impl DatasetMeta {
pub(crate) fn load(f: &File, addr: u64, hdr: &ObjectHeader, path: &str) -> PyResult<Self> {
let null = node::is_null(&node::dataspace(f, hdr, path)?);
let ds = f.dataset_at(addr).map_err(to_py_err)?;
let shape = if null {
None
} else {
Some(ds.shape().map_err(to_py_err)?)
};
let datatype = ds.raw_datatype().map_err(to_py_err)?;
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(f, hdr, s.len()));
Ok(Self {
shape,
chunks,
datatype,
})
}
}
/// A dataset in a file opened for reading. /// A dataset in a file opened for reading.
/// ///
/// ```python /// ```python
@@ -30,7 +62,7 @@ use crate::{PyEmpty, node, to_py_err};
/// ``` /// ```
#[pyclass(name = "Dataset")] #[pyclass(name = "Dataset")]
pub struct PyDataset { pub struct PyDataset {
file: Arc<clawhdf5_rs::File>, handle: Arc<Handle>,
path: String, path: String,
/// Where the dataset's object header is: reads open it from here rather /// Where the dataset's object header is: reads open it from here rather
/// than resolve `path` again. /// than resolve `path` again.
@@ -45,39 +77,24 @@ pub struct PyDataset {
} }
impl PyDataset { impl PyDataset {
pub(crate) fn open( pub(crate) fn new(
py: Python<'_>, py: Python<'_>,
file: Arc<clawhdf5_rs::File>, handle: Arc<Handle>,
path: String, path: String,
addr: u64, addr: u64,
hdr: &ObjectHeader, meta: DatasetMeta,
) -> PyResult<Self> { ) -> Self {
crate::no_panic(|| { let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
let null = node::is_null(&node::dataspace(&file, hdr)?); .map_err(|e| e.value(py).to_string());
let (shape, datatype) = { Self {
let ds = file.dataset_at(addr).map_err(to_py_err)?; handle,
let shape = if null { path,
None addr,
} else { shape: meta.shape,
Some(ds.shape().map_err(to_py_err)?) chunks: meta.chunks,
}; datatype: meta.datatype,
(shape, ds.raw_datatype().map_err(to_py_err)?) conv,
}; }
let conv = Converter::new(py, &datatype, file.superblock().offset_size)
.map_err(|e| e.value(py).to_string());
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(&file, hdr, s.len()));
Ok(Self {
file,
path,
addr,
shape,
chunks,
datatype,
conv,
})
})
} }
fn converter(&self) -> PyResult<&Converter> { fn converter(&self) -> PyResult<&Converter> {
@@ -103,10 +120,10 @@ impl PyDataset {
}; };
let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size); let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size);
let read_shape = plan.read_shape(); let read_shape = plan.read_shape();
let file = &*self.file; let handle = &*self.handle;
let addr = self.addr; let addr = self.addr;
// Everything below touches only Rust data: release the GIL. // Everything below touches only Rust data: release the GIL.
let read = || -> Result<Elements, ReadError> { let read = |file: &File| -> Result<Elements, ReadError> {
let ds = file.dataset_at(addr)?; let ds = file.dataset_at(addr)?;
let mut blocks = Vec::with_capacity(reads.len()); let mut blocks = Vec::with_capacity(reads.len());
for read in reads { for read in reads {
@@ -147,7 +164,7 @@ impl PyDataset {
let sb = file.superblock(); let sb = file.superblock();
let n = read_shape.iter().product(); let n = read_shape.iter().product();
resolve_vl( resolve_vl(
file.as_bytes(), file.storage(),
&raw, &raw,
n, n,
sb.offset_size, sb.offset_size,
@@ -155,12 +172,20 @@ impl PyDataset {
unit, unit,
) )
.map(Elements::Vl) .map(Elements::Vl)
.map_err(ReadError::Other) .map_err(ReadError::Vl)
}; };
let data = py let data = py
.detach(|| { .detach(|| {
std::panic::catch_unwind(std::panic::AssertUnwindSafe(read)) handle
.unwrap_or_else(|p| Err(ReadError::Panic(crate::panic_text(&*p)))) .with_detached(|f| {
Ok(
std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| read(f)))
.unwrap_or_else(|p| {
Err(ReadError::Panic(crate::panic_text(&*p)))
}),
)
})
.unwrap_or_else(|e| Err(ReadError::Py(e)))
}) })
.map_err(|e| e.into_py(&self.path))?; .map_err(|e| e.into_py(&self.path))?;
let joined = conv.to_array(py, data, &read_shape, false)?; let joined = conv.to_array(py, data, &read_shape, false)?;
@@ -183,6 +208,8 @@ impl PyDataset {
/// An error from the read closure, turned into a Python error with the GIL. /// An error from the read closure, turned into a Python error with the GIL.
enum ReadError { enum ReadError {
Lib(clawhdf5_rs::Error), Lib(clawhdf5_rs::Error),
Vl(VlError),
Py(PyErr),
Other(String), Other(String),
Panic(String), Panic(String),
} }
@@ -197,6 +224,8 @@ impl ReadError {
fn into_py(self, path: &str) -> PyErr { fn into_py(self, path: &str) -> PyErr {
match self { match self {
ReadError::Lib(e) => to_py_err(e), ReadError::Lib(e) => to_py_err(e),
ReadError::Vl(e) => e.into_py(&node::name(path)),
ReadError::Py(e) => e,
ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))), ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))),
ReadError::Panic(msg) => crate::InternalError::new_err(format!( ReadError::Panic(msg) => crate::InternalError::new_err(format!(
"{}: clawhdf5 internal error (please report it): {msg}", "{}: clawhdf5 internal error (please report it): {msg}",
@@ -252,22 +281,23 @@ impl PyDataset {
/// The maximum shape (`None` per unlimited dimension), like h5py. /// The maximum shape (`None` per unlimited dimension), like h5py.
#[getter] #[getter]
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> { fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| { let Some(shape) = &self.shape else {
let Some(shape) = &self.shape else { return Ok(py.None().into_bound(py));
return Ok(py.None().into_bound(py)); };
}; let addr = self.addr;
let max = self let max = self
.file .handle
.dataset_at(self.addr) .with(py, |f| {
.and_then(|ds| ds.max_dimensions()) f.dataset_at(addr)
.map_err(to_py_err)? .and_then(|ds| ds.max_dimensions())
.unwrap_or_else(|| shape.clone()); .map_err(to_py_err)
let items: Vec<Option<u64>> = max })?
.into_iter() .unwrap_or_else(|| shape.clone());
.map(|d| (d != u64::MAX).then_some(d)) let items: Vec<Option<u64>> = max
.collect(); .into_iter()
Ok(PyTuple::new(py, items)?.into_any()) .map(|d| (d != u64::MAX).then_some(d))
}) .collect();
Ok(PyTuple::new(py, items)?.into_any())
} }
/// The dataset's numpy dtype, as h5py reports it. /// The dataset's numpy dtype, as h5py reports it.
@@ -295,8 +325,8 @@ impl PyDataset {
/// The dataset's attributes (read-only, dict-like). /// The dataset's attributes (read-only, dict-like).
#[getter] #[getter]
fn attrs(&self) -> PyResult<PyAttrs> { fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path) PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
} }
/// Read with h5py indexing: integers, slices with positive steps, /// Read with h5py indexing: integers, slices with positive steps,
+177 -29
View File
@@ -1,13 +1,17 @@
//! PyFile — the main entry point for opening and creating HDF5 files. //! PyFile — the main entry point for opening and creating HDF5 files.
use std::collections::HashMap;
use std::path::PathBuf; use std::path::PathBuf;
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use std::time::Duration;
use pyo3::exceptions::PyValueError;
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::PyList; use pyo3::types::{PyDict, PyList};
use crate::attrs::PyAttrs; use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group}; use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err}; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
/// Internal state for write mode. /// Internal state for write mode.
@@ -23,8 +27,10 @@ struct WriteState {
/// Mirrors the h5py.File interface: /// Mirrors the h5py.File interface:
/// ///
/// ```python /// ```python
/// # Reading /// # Reading, a local file or a URL (range requests, nothing downloaded
/// # up front)
/// f = clawhdf5.File('data.h5', 'r') /// f = clawhdf5.File('data.h5', 'r')
/// f = clawhdf5.File('https://example.org/data.h5')
/// ds = f['dataset'] /// ds = f['dataset']
/// f.close() /// f.close()
/// ///
@@ -44,27 +50,51 @@ enum FileInner {
Write(WriteState), Write(WriteState),
} }
/// Whether `s` is a URL (`scheme://…`) rather than a path: the scheme is a
/// letter followed by letters, digits, `+`, `-` or `.` (RFC 3986).
fn is_url(s: &str) -> bool {
let Some((scheme, _)) = s.split_once("://") else {
return false;
};
let mut chars = scheme.chars();
chars.next().is_some_and(|c| c.is_ascii_alphabetic())
&& chars.all(|c| c.is_ascii_alphanumeric() || matches!(c, '+' | '-' | '.'))
}
impl PyFile {
fn from_handle(handle: Arc<Handle>, filename: String) -> Self {
let root = handle.root;
Self {
inner: Some(FileInner::Read(ReadGroup::new(handle, String::new(), root))),
filename,
}
}
}
#[pymethods] #[pymethods]
impl PyFile { impl PyFile {
/// Open or create an HDF5 file. /// Open or create an HDF5 file.
/// ///
/// Parameters: /// Parameters:
/// path: file path /// path: file path, or a URL (`http://`, `https://`, `s3://`, `gs://`,
/// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write /// mode: 'r' for read (default), 'w' for write
#[new] #[new]
#[pyo3(signature = (path, mode="r"))] #[pyo3(signature = (path, mode="r"))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> { fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
let filename = path.to_string(); let filename = path.to_string();
match mode { if is_url(path) {
"r" => { if mode != "r" {
let file = py.detach(|| { return Err(PyValueError::new_err(format!(
crate::no_panic(|| clawhdf5_rs::File::open(path).map_err(to_py_err)) "remote files are read-only: mode '{mode}' is not supported for a URL"
})?; )));
Ok(Self {
inner: Some(FileInner::Read(root_group(Arc::new(file)))),
filename,
})
} }
let handle = Handle::open_url(py, path, &clawhdf5_remote::Options::default())?;
return Ok(Self::from_handle(handle, filename));
}
match mode {
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
"w" => Ok(Self { "w" => Ok(Self {
filename, filename,
inner: Some(FileInner::Write(WriteState { inner: Some(FileInner::Write(WriteState {
@@ -74,12 +104,122 @@ impl PyFile {
groups: Vec::new(), groups: Vec::new(),
})), })),
}), }),
other => Err(PyErr::new::<pyo3::exceptions::PyValueError, _>(format!( other => Err(PyValueError::new_err(format!(
"unsupported mode '{other}'; expected 'r' or 'w'" "unsupported mode '{other}'; expected 'r' or 'w'"
))), ))),
} }
} }
/// Open a remote file for reading, with options.
///
/// The file is read through a block cache with range requests: opening
/// costs one request (it also fetches the first block), and a read
/// fetches only the blocks it needs. The GIL is released while waiting
/// on the network.
///
/// Parameters (all optional):
/// block_size: bytes per cached block (default 1 MiB)
/// cache_size: byte budget of the block cache (default 64 MiB)
/// headers: dict of extra HTTP headers (e.g. Authorization), sent only
/// to the URL's own origin
/// retries: retries of a request that failed transiently (default 3)
/// timeout: seconds to connect and receive response headers (default 30)
/// allow_full_download: when the server ignores Range requests,
/// download the whole file once instead of failing (default False)
/// max_full_download: largest file such a download may fetch
/// (default 1 GiB)
/// require_validator: refuse a server that sends neither ETag nor
/// Last-Modified (default False)
/// max_redirects: redirects followed per request (default 5)
/// max_parallel: requests of one read in flight at once (default 8)
#[staticmethod]
#[allow(clippy::too_many_arguments)]
#[pyo3(signature = (url, *, block_size=None, cache_size=None, headers=None, retries=None,
timeout=None, allow_full_download=None, max_full_download=None,
require_validator=None, max_redirects=None, max_parallel=None))]
fn open_url(
py: Python<'_>,
url: &str,
block_size: Option<u64>,
cache_size: Option<u64>,
headers: Option<HashMap<String, String>>,
retries: Option<u32>,
timeout: Option<f64>,
allow_full_download: Option<bool>,
max_full_download: Option<u64>,
require_validator: Option<bool>,
max_redirects: Option<u32>,
max_parallel: Option<usize>,
) -> PyResult<Self> {
let mut options = clawhdf5_remote::Options::default();
if let Some(b) = block_size {
if b == 0 {
return Err(PyValueError::new_err("block_size must be positive"));
}
options.cache.block_size = b;
options.cache.coalesce_gap = b;
// The opening request fetches the first block, not 1 MiB.
options.http.first_request = b;
}
if let Some(c) = cache_size {
options.cache.capacity = c;
}
let http = &mut options.http;
if let Some(h) = headers {
http.headers = h.into_iter().collect();
}
if let Some(r) = retries {
http.retries = r;
}
if let Some(t) = timeout {
if !(t.is_finite() && t > 0.0) {
return Err(PyValueError::new_err("timeout must be a positive number"));
}
http.timeout = Duration::from_secs_f64(t);
}
if let Some(a) = allow_full_download {
http.allow_full_download = a;
}
if let Some(m) = max_full_download {
http.max_full_download = m;
}
if let Some(v) = require_validator {
http.require_validator = v;
}
if let Some(r) = max_redirects {
http.max_redirects = r;
}
if let Some(p) = max_parallel {
if p == 0 {
return Err(PyValueError::new_err("max_parallel must be positive"));
}
http.max_parallel = p;
}
let handle = Handle::open_url(py, url, &options)?;
Ok(Self::from_handle(handle, url.to_string()))
}
/// For a remote file, what its block cache has done so far (reads,
/// hits, misses, requests, bytes fetched, ...); `None` for a local file.
#[getter]
fn remote_stats<'py>(&self, py: Python<'py>) -> PyResult<Option<Bound<'py, PyDict>>> {
let Some(storage) = self.read_file()?.handle.remote_storage() else {
return Ok(None);
};
let s = storage.stats();
let d = PyDict::new(py);
d.set_item("reads", s.reads)?;
d.set_item("hits", s.hits)?;
d.set_item("misses", s.misses)?;
d.set_item("waits", s.waits)?;
d.set_item("requests", s.requests)?;
d.set_item("fetch_calls", s.fetch_calls)?;
d.set_item("bytes_fetched", s.bytes_fetched)?;
d.set_item("evictions", s.evictions)?;
d.set_item("cached_bytes", s.cached_bytes)?;
Ok(Some(d))
}
/// Close the file. In write mode, this finalizes and writes the file. /// Close the file. In write mode, this finalizes and writes the file.
fn close(&mut self) -> PyResult<()> { fn close(&mut self) -> PyResult<()> {
let inner = self.inner.take().ok_or_else(|| { let inner = self.inner.take().ok_or_else(|| {
@@ -121,7 +261,7 @@ impl PyFile {
/// List the names of all children in the root group. /// List the names of all children in the root group.
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let names = self.read_file()?.member_names()?; let names = self.read_file()?.member_names(py)?;
Ok(PyList::new(py, names)?.into_any().unbind()) Ok(PyList::new(py, names)?.into_any().unbind())
} }
@@ -139,8 +279,8 @@ impl PyFile {
self.keys(py)?.call_method0(py, "__iter__") self.keys(py)?.call_method0(py, "__iter__")
} }
fn __len__(&self) -> PyResult<usize> { fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
Ok(self.read_file()?.member_names()?.len()) Ok(self.read_file()?.member_names(py)?.len())
} }
/// The root group's name, `/`. /// The root group's name, `/`.
@@ -149,7 +289,7 @@ impl PyFile {
"/" "/"
} }
/// The path the file was opened with. /// The path (or URL) the file was opened with.
#[getter] #[getter]
fn filename(&self) -> &str { fn filename(&self) -> &str {
&self.filename &self.filename
@@ -204,9 +344,9 @@ impl PyFile {
/// Attribute access. In read mode, returns attributes of the root group. /// Attribute access. In read mode, returns attributes of the root group.
/// In write mode, returns a writable attrs handle. /// In write mode, returns a writable attrs handle.
#[getter] #[getter]
fn attrs(&self) -> PyResult<PyAttrs> { fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match self.inner.as_ref() { match self.inner.as_ref() {
Some(FileInner::Read(root)) => root.attrs(), Some(FileInner::Read(root)) => root.attrs(py),
Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))), Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))),
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>( None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"file is closed", "file is closed",
@@ -216,9 +356,10 @@ impl PyFile {
fn __repr__(&self) -> String { fn __repr__(&self) -> String {
match &self.inner { match &self.inner {
Some(FileInner::Read(root)) => { Some(FileInner::Read(root)) => match root.handle.redacted_url() {
format!("<HDF5 File (read, {} bytes)>", root.file.as_bytes().len()) Some(url) => format!("<HDF5 File (read, \"{url}\")>"),
} None => format!("<HDF5 File (read, \"{}\")>", self.filename),
},
Some(FileInner::Write(s)) => { Some(FileInner::Write(s)) => {
format!("<HDF5 File (write, \"{}\")>", s.path.display()) format!("<HDF5 File (write, \"{}\")>", s.path.display())
} }
@@ -226,8 +367,8 @@ impl PyFile {
} }
} }
fn __contains__(&self, key: &str) -> PyResult<bool> { fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
Ok(self.read_file()?.contains(key)) self.read_file()?.contains(py, key)
} }
} }
@@ -271,11 +412,6 @@ fn parse_compression(
} }
} }
fn root_group(file: Arc<clawhdf5_rs::File>) -> ReadGroup {
let root = file.superblock().root_group_address;
ReadGroup::new(file, String::new(), root)
}
/// Build and write the HDF5 file from accumulated write state. /// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> { fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| { crate::no_panic(|| {
@@ -309,6 +445,18 @@ fn finalize_write(state: WriteState) -> PyResult<()> {
mod tests { mod tests {
use super::*; use super::*;
#[test]
fn urls_and_paths() {
assert!(is_url("http://h/f.h5"));
assert!(is_url("s3://bucket/key.h5"));
assert!(is_url("git+https://x"));
assert!(!is_url("data.h5"));
assert!(!is_url("/tmp/a://b.h5"));
assert!(!is_url("dir/x://y"));
assert!(!is_url("1http://x"));
assert!(!is_url("://x"));
}
#[test] #[test]
fn parse_gzip_compression() { fn parse_gzip_compression() {
assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6)); assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6));
+67 -84
View File
@@ -3,11 +3,12 @@
use std::collections::HashMap; use std::collections::HashMap;
use std::sync::{Arc, Mutex, OnceLock}; use std::sync::{Arc, Mutex, OnceLock};
use pyo3::exceptions::{PyIOError, PyKeyError, PyValueError}; use pyo3::exceptions::{PyIOError, PyKeyError, PyOSError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use pyo3::types::PyList; use pyo3::types::PyList;
use crate::attrs::PyAttrs; use crate::attrs::PyAttrs;
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node}; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
/// Shared state for a group being written. /// Shared state for a group being written.
@@ -34,9 +35,9 @@ enum GroupInner {
} }
impl PyGroup { impl PyGroup {
pub(crate) fn from_read(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self { pub(crate) fn from_read(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self { Self {
inner: GroupInner::Read(ReadGroup::new(file, path, addr)), inner: GroupInner::Read(ReadGroup::new(handle, path, addr)),
} }
} }
@@ -60,9 +61,10 @@ impl PyGroup {
/// h5py). It keeps its own address and, once listed, its links, so looking /// h5py). It keeps its own address and, once listed, its links, so looking
/// up a child neither resolves the path from the root nor scans the group's /// up a child neither resolves the path from the root nor scans the group's
/// links again: visiting every member of a large group is linear, not /// links again: visiting every member of a large group is linear, not
/// quadratic. /// quadratic. (Edits never add or remove links, so these stay valid in a
/// file open for editing.)
pub(crate) struct ReadGroup { pub(crate) struct ReadGroup {
pub file: Arc<clawhdf5_rs::File>, pub handle: Arc<Handle>,
pub path: String, pub path: String,
pub addr: u64, pub addr: u64,
/// Link name -> object address (soft links resolved), filled on first use. /// Link name -> object address (soft links resolved), filled on first use.
@@ -72,9 +74,9 @@ pub(crate) struct ReadGroup {
} }
impl ReadGroup { impl ReadGroup {
pub(crate) fn new(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self { pub(crate) fn new(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self { Self {
file, handle,
path, path,
addr, addr,
links: OnceLock::new(), links: OnceLock::new(),
@@ -82,17 +84,14 @@ impl ReadGroup {
} }
} }
fn links(&self) -> PyResult<&HashMap<String, u64>> { fn links(&self, py: Python<'_>) -> PyResult<&HashMap<String, u64>> {
if let Some(links) = self.links.get() { if let Some(links) = self.links.get() {
return Ok(links); return Ok(links);
} }
let entries = crate::no_panic(|| { let (addr, path) = (self.addr, &self.path);
clawhdf5_format::group_v2::resolve_group_children( let entries = self.handle.with(py, |f| {
self.file.as_bytes(), clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr)
self.file.superblock(), .map_err(|e| node::format_err(path, e, PyValueError::new_err))
self.addr,
)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", node::name(&self.path))))
})?; })?;
let map = entries let map = entries
.into_iter() .into_iter()
@@ -102,7 +101,7 @@ impl ReadGroup {
} }
/// The path and address of `key` (a name, a relative or an absolute path). /// The path and address of `key` (a name, a relative or an absolute path).
fn locate(&self, key: &str) -> PyResult<(String, u64)> { fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> {
let path = node::join(&self.path, key); let path = node::join(&self.path, key);
let rel = if self.path.is_empty() { let rel = if self.path.is_empty() {
Some(path.as_str()) Some(path.as_str())
@@ -112,24 +111,24 @@ impl ReadGroup {
path.strip_prefix(self.path.as_str()) path.strip_prefix(self.path.as_str())
.and_then(|r| r.strip_prefix('/')) .and_then(|r| r.strip_prefix('/'))
}; };
let addr = match rel { // A direct child: the link table, when it has the name.
// A direct child: the link table, when it has the name. if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/'))
Some(name) if !name.is_empty() && !name.contains('/') => { && let Some(&a) = self.links(py)?.get(name)
match self.links()?.get(name) { {
Some(&a) => a, return Ok((path, a));
None => node::resolve_from(&self.file, self.addr, name, &path)?, }
} let addr = self.addr;
} let found = self.handle.with(py, |f| match rel {
Some(rel) => node::resolve_from(&self.file, self.addr, rel, &path)?, Some(rel) => node::resolve_from(f, addr, rel, &path),
None => node::address(&self.file, &path)?, None => node::address(f, &path),
}; })?;
Ok((path, addr)) Ok((path, found))
} }
/// `group[key]`. /// `group[key]`.
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> { pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
let (path, addr) = self.locate(key)?; let (path, addr) = self.locate(py, key)?;
node::open(py, &self.file, path, addr) node::open(py, &self.handle, path, addr)
} }
/// `group.get(key, default)`. /// `group.get(key, default)`.
@@ -148,48 +147,57 @@ impl ReadGroup {
} }
/// Names of the group's datasets and subgroups, sorted (h5py's order). /// Names of the group's datasets and subgroups, sorted (h5py's order).
pub(crate) fn member_names(&self) -> PyResult<&[String]> { pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> {
if let Some(m) = self.members.get() { if let Some(m) = self.members.get() {
return Ok(m); return Ok(m);
} }
let mut names = Vec::new(); let links = self.links(py)?;
for (name, &addr) in self.links()? { let path = &self.path;
let hdr = node::header_at(&self.file, addr, &node::join(&self.path, name))?; let mut names = self.handle.with(py, |f| {
if matches!( let mut names = Vec::new();
node::kind(&hdr), for (name, &addr) in links {
Some(node::Kind::Dataset | node::Kind::Group) if matches!(
) { node::kind_at(f, addr, &node::join(path, name))?,
names.push(name.clone()); Some(node::Kind::Dataset | node::Kind::Group)
) {
names.push(name.clone());
}
} }
} Ok(names)
})?;
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes())); names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
Ok(self.members.get_or_init(|| names)) Ok(self.members.get_or_init(|| names))
} }
pub(crate) fn contains(&self, key: &str) -> bool { /// `key in group`: whether `key` names a dataset or group. A failed
self.locate(key) /// read of the file (a network error) is raised, not `False`.
.and_then(|(path, addr)| node::header_at(&self.file, addr, &path)) pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
.ok() let found = self
.and_then(|h| node::kind(&h)) .locate(py, key)
.is_some_and(|k| k != node::Kind::Datatype) .and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path)));
match found {
Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)),
Err(e) if e.is_instance_of::<PyOSError>(py) => Err(e),
Err(_) => Ok(false),
}
} }
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> { pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
self.member_names()? self.member_names(py)?
.iter() .iter()
.map(|n| self.get_item(py, n)) .map(|n| self.get_item(py, n))
.collect() .collect()
} }
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> { pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
self.member_names()? self.member_names(py)?
.iter() .iter()
.map(|n| Ok((n.clone(), self.get_item(py, n)?))) .map(|n| Ok((n.clone(), self.get_item(py, n)?)))
.collect() .collect()
} }
pub(crate) fn attrs(&self) -> PyResult<PyAttrs> { pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path) PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
} }
} }
@@ -210,7 +218,7 @@ impl PyGroup {
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> { fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => { GroupInner::Read(g) => {
let list = PyList::new(py, g.member_names()?)?; let list = PyList::new(py, g.member_names(py)?)?;
Ok(list.into_any().unbind()) Ok(list.into_any().unbind())
} }
GroupInner::Write(state) => { GroupInner::Write(state) => {
@@ -236,9 +244,9 @@ impl PyGroup {
self.keys(py)?.call_method0(py, "__iter__") self.keys(py)?.call_method0(py, "__iter__")
} }
fn __len__(&self) -> PyResult<usize> { fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => Ok(g.member_names()?.len()), GroupInner::Read(g) => Ok(g.member_names(py)?.len()),
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()), GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
} }
} }
@@ -301,9 +309,9 @@ impl PyGroup {
/// Attribute access. /// Attribute access.
#[getter] #[getter]
fn attrs(&self) -> PyResult<PyAttrs> { fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => g.attrs(), GroupInner::Read(g) => g.attrs(py),
GroupInner::Write(state) => { GroupInner::Write(state) => {
let store = Arc::clone(&state.lock().unwrap().attrs); let store = Arc::clone(&state.lock().unwrap().attrs);
Ok(PyAttrs::from_write(store)) Ok(PyAttrs::from_write(store))
@@ -311,10 +319,10 @@ impl PyGroup {
} }
} }
fn __repr__(&self) -> String { fn __repr__(&self, py: Python<'_>) -> String {
match &self.inner { match &self.inner {
GroupInner::Read(g) => { GroupInner::Read(g) => {
let n = g.member_names().map_or(0, |m| m.len()); let n = g.member_names(py).map_or(0, |m| m.len());
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path)) format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
} }
GroupInner::Write(state) => { GroupInner::Write(state) => {
@@ -324,9 +332,9 @@ impl PyGroup {
} }
} }
fn __contains__(&self, key: &str) -> PyResult<bool> { fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
match &self.inner { match &self.inner {
GroupInner::Read(g) => Ok(g.contains(key)), GroupInner::Read(g) => g.contains(py, key),
GroupInner::Write(state) => { GroupInner::Write(state) => {
let guard = state.lock().unwrap(); let guard = state.lock().unwrap();
Ok(guard.datasets.iter().any(|d| d.name == key)) Ok(guard.datasets.iter().any(|d| d.name == key))
@@ -357,31 +365,6 @@ pub(crate) fn finalize_write_group(
mod tests { mod tests {
use super::*; use super::*;
#[test]
fn member_names_are_sorted() {
let mut b = clawhdf5_rs::FileBuilder::new();
b.create_dataset("zeta").with_f64_data(&[1.0]);
b.create_dataset("alpha").with_f64_data(&[1.0]);
let mut g = b.create_group("mid");
g.create_dataset("x").with_f64_data(&[1.0]);
let finished = g.finish();
b.add_group(finished);
let bytes = b.finish().unwrap();
let file = Arc::new(clawhdf5_rs::File::from_bytes(bytes).unwrap());
let root = file.superblock().root_group_address;
let top = ReadGroup::new(Arc::clone(&file), String::new(), root);
assert_eq!(top.member_names().unwrap(), ["alpha", "mid", "zeta"]);
let (path, addr) = top.locate("mid").unwrap();
assert_eq!(path, "mid");
let mid = ReadGroup::new(Arc::clone(&file), path, addr);
assert_eq!(mid.member_names().unwrap(), ["x"]);
assert!(top.contains("mid/x"));
assert!(mid.contains("/alpha"));
assert!(mid.contains("x") && mid.contains("./x"));
assert!(!top.contains("nope"));
assert!(!mid.contains("alpha"));
}
#[test] #[test]
fn finalize_group() { fn finalize_group() {
let state = WriteGroupState { let state = WriteGroupState {
+130
View File
@@ -0,0 +1,130 @@
//! The open file every object of a `File` shares.
//!
//! Every read goes through [`Handle::with`], which releases the GIL and
//! parses through `File::storage()`, so the same code serves a local file
//! (memory-mapped) and a remote one (`clawhdf5-remote`: range requests
//! through a block cache, so a network read never holds the GIL).
//!
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
//! with the GIL released, and code that holds it never touches Python.
use std::path::PathBuf;
use std::sync::{Arc, PoisonError, RwLock};
use clawhdf5_rs::File;
use pyo3::exceptions::PyOSError;
use pyo3::prelude::*;
use crate::to_py_err;
/// Where the file's bytes come from.
pub(crate) enum Source {
/// A local path (memory-mapped).
Local(#[allow(dead_code)] PathBuf),
/// A URL, read through `clawhdf5-remote`'s block cache.
Remote {
url: String,
storage: Arc<clawhdf5_remote::RemoteStorage>,
},
}
pub(crate) struct Handle {
file: RwLock<Option<File>>,
source: Source,
pub offset_size: u8,
pub length_size: u8,
pub root: u64,
}
fn closed_after_failed_reopen() -> PyErr {
PyOSError::new_err("the file could not be reopened after an edit; open it again")
}
impl Handle {
fn new(file: File, source: Source) -> Arc<Self> {
let sb = file.superblock();
let (offset_size, length_size, root) =
(sb.offset_size, sb.length_size, sb.root_group_address);
Arc::new(Self {
file: RwLock::new(Some(file)),
source,
offset_size,
length_size,
root,
})
}
/// A local file, read-only.
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
Ok(Self::new(file, Source::Local(PathBuf::from(path))))
}
/// A remote file (`http(s)://`, `s3://`, ...).
pub(crate) fn open_url(
py: Python<'_>,
url: &str,
options: &clawhdf5_remote::Options,
) -> PyResult<Arc<Self>> {
let (file, storage) = py.detach(|| {
crate::no_panic(|| {
let storage = clawhdf5_remote::storage_for_url(url, options).map_err(remote_err)?;
let file = File::open_storage(storage.clone()).map_err(to_py_err)?;
Ok((file, storage))
})
})?;
Ok(Self::new(
file,
Source::Remote {
url: url.to_string(),
storage,
},
))
}
/// Run `f` on the file with the GIL released (a remote read may wait
/// on the network; other Python threads run meanwhile). `f` must not
/// touch Python.
pub(crate) fn with<R: Send>(
&self,
py: Python<'_>,
f: impl FnOnce(&File) -> PyResult<R> + Send,
) -> PyResult<R> {
py.detach(|| self.with_detached(f))
}
/// [`with`](Self::with) for code that already runs without the GIL.
pub(crate) fn with_detached<R>(&self, f: impl FnOnce(&File) -> PyResult<R>) -> PyResult<R> {
crate::no_panic(|| {
let guard = self.file.read().unwrap_or_else(PoisonError::into_inner);
let file = guard.as_ref().ok_or_else(closed_after_failed_reopen)?;
f(file)
})
}
/// The remote file's block cache.
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
match &self.source {
Source::Remote { storage, .. } => Some(storage),
Source::Local(_) => None,
}
}
/// The URL of a remote file, credentials and query values redacted.
pub(crate) fn redacted_url(&self) -> Option<String> {
match &self.source {
Source::Remote { url, .. } => Some(clawhdf5_remote::redact_url(url)),
Source::Local(_) => None,
}
}
}
/// A `clawhdf5_remote::Error` as a Python exception: the network side
/// (unreachable, a status, no range support, a changed file) is `OSError`,
/// a file that is not HDF5 is what `to_py_err` makes of it.
pub(crate) fn remote_err(e: clawhdf5_remote::Error) -> PyErr {
match e {
clawhdf5_remote::Error::Hdf5(e) => to_py_err(e),
other => PyOSError::new_err(other.to_string()),
}
}
+6 -1
View File
@@ -14,6 +14,7 @@ mod convert;
mod dataset; mod dataset;
mod file; mod file;
mod group; mod group;
mod handle;
mod node; mod node;
mod select; mod select;
@@ -63,7 +64,7 @@ fn _panic_for_test() -> PyResult<()> {
/// Convert a `clawhdf5_rs::Error` into a `PyErr`. /// Convert a `clawhdf5_rs::Error` into a `PyErr`.
/// ///
/// Maps different error variants to more specific Python exception types: /// Maps different error variants to more specific Python exception types:
/// - I/O errors -> `PyIOError` /// - I/O errors, and failed reads of a remote file -> `PyIOError`/`PyOSError`
/// - Format/parsing errors -> `PyValueError` /// - Format/parsing errors -> `PyValueError`
/// - Missing dataset/path errors -> `PyKeyError` /// - Missing dataset/path errors -> `PyKeyError`
/// - Invalid arguments -> `PyValueError` /// - Invalid arguments -> `PyValueError`
@@ -73,6 +74,10 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr {
use clawhdf5_rs::Error; use clawhdf5_rs::Error;
match &e { match &e {
Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()), Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()),
// A failed read of the storage: a network error on a remote file.
Error::Format(clawhdf5_format::error::FormatError::Storage(_)) => {
PyErr::new::<pyo3::exceptions::PyOSError, _>(e.to_string())
}
Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()), Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()),
Error::NotADataset(_) | Error::MissingMessage(_) => { Error::NotADataset(_) | Error::MissingMessage(_) => {
PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string()) PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string())
+93 -72
View File
@@ -1,16 +1,24 @@
//! Resolving paths to objects in a file opened for reading. //! Resolving paths to objects in a file opened for reading.
//!
//! Everything here parses through `File::storage()` (the `clawhdf5_format`
//! `*_in` functions), never `File::as_bytes()`, so it works the same on a
//! memory-mapped local file and on a remote one; and it runs inside
//! `Handle::with`, without the GIL.
use std::sync::Arc; use std::sync::Arc;
use clawhdf5_format::attribute::AttributeMessage; use clawhdf5_format::attribute::AttributeMessage;
use clawhdf5_format::dataspace::{Dataspace, DataspaceType}; use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
use clawhdf5_format::error::FormatError;
use clawhdf5_format::message_type::MessageType; use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader; use clawhdf5_format::object_header::ObjectHeader;
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError}; use clawhdf5_rs::File;
use pyo3::exceptions::{PyKeyError, PyOSError, PyTypeError, PyValueError};
use pyo3::prelude::*; use pyo3::prelude::*;
use crate::dataset::PyDataset; use crate::dataset::{DatasetMeta, PyDataset};
use crate::group::PyGroup; use crate::group::PyGroup;
use crate::handle::Handle;
/// Join `key` onto the group path `base` the way h5py does: an absolute key /// Join `key` onto the group path `base` the way h5py does: an absolute key
/// starts from the root, a relative one from `base`. Paths are kept without /// starts from the root, a relative one from `base`. Paths are kept without
@@ -33,42 +41,46 @@ pub(crate) fn name(path: &str) -> String {
format!("/{path}") format!("/{path}")
} }
/// A format error met at `path`: a failed read of the storage (a network
/// error on a remote file) is an `OSError`, anything else `other(message)`.
pub(crate) fn format_err(path: &str, e: FormatError, other: fn(String) -> PyErr) -> PyErr {
let msg = format!("{}: {e}", name(path));
match e {
FormatError::Storage(_) => PyOSError::new_err(msg),
_ => other(msg),
}
}
fn value_err(msg: String) -> PyErr {
PyValueError::new_err(msg)
}
/// The address of the object at `path`, resolved from the root group. /// The address of the object at `path`, resolved from the root group.
pub(crate) fn address(file: &clawhdf5_rs::File, path: &str) -> PyResult<u64> { pub(crate) fn address(file: &File, path: &str) -> PyResult<u64> {
resolve_from(file, file.superblock().root_group_address, path, path) resolve_from(file, file.superblock().root_group_address, path, path)
} }
/// The address of `rel` resolved from the group at `group` (`full` is the /// The address of `rel` resolved from the group at `group` (`full` is the
/// resulting path, for the error message). /// resulting path, for the error message).
pub(crate) fn resolve_from( pub(crate) fn resolve_from(file: &File, group: u64, rel: &str, full: &str) -> PyResult<u64> {
file: &clawhdf5_rs::File,
group: u64,
rel: &str,
full: &str,
) -> PyResult<u64> {
if rel.is_empty() { if rel.is_empty() {
return Ok(group); return Ok(group);
} }
crate::no_panic(|| { clawhdf5_format::group_v2::resolve_path_from_in(file.storage(), file.superblock(), group, rel)
clawhdf5_format::group_v2::resolve_path_from(file.as_bytes(), file.superblock(), group, rel) .map_err(|e| match e {
.map_err(|e| { FormatError::Storage(_) => format_err(full, e, value_err),
PyKeyError::new_err(format!( e => PyKeyError::new_err(format!(
"Unable to open object (object '{}' doesn't exist): {e}", "Unable to open object (object '{}' doesn't exist): {e}",
name(full) name(full)
)) )),
}) })
})
} }
/// The object header at `addr` (the object at `path`). /// The object header at `addr` (the object at `path`).
pub(crate) fn header_at(file: &clawhdf5_rs::File, addr: u64, path: &str) -> PyResult<ObjectHeader> { pub(crate) fn header_at(file: &File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
crate::no_panic(|| { let sb = file.superblock();
let sb = file.superblock(); ObjectHeader::parse_in(file.storage(), addr, sb.offset_size, sb.length_size)
let at = usize::try_from(addr) .map_err(|e| format_err(path, e, value_err))
.map_err(|_| PyValueError::new_err(format!("{}: address out of range", name(path))))?;
ObjectHeader::parse(file.as_bytes(), at, sb.offset_size, sb.length_size)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))
})
} }
/// What an object header describes. /// What an object header describes.
@@ -96,29 +108,50 @@ pub(crate) fn kind(hdr: &ObjectHeader) -> Option<Kind> {
} }
} }
/// The kind of the object at `addr`, from its header.
pub(crate) fn kind_at(file: &File, addr: u64, path: &str) -> PyResult<Option<Kind>> {
Ok(kind(&header_at(file, addr, path)?))
}
/// What opening an object found, read without the GIL.
enum Found {
Dataset(DatasetMeta),
Group,
Datatype,
Other,
}
/// Open the object at `addr` (whose path is `path`) as a `Dataset` or /// Open the object at `addr` (whose path is `path`) as a `Dataset` or
/// `Group`. Both keep the address, so later reads resolve nothing. /// `Group`. Both keep the address, so later reads resolve nothing.
pub(crate) fn open( pub(crate) fn open(
py: Python<'_>, py: Python<'_>,
file: &Arc<clawhdf5_rs::File>, handle: &Arc<Handle>,
path: String, path: String,
addr: u64, addr: u64,
) -> PyResult<Py<PyAny>> { ) -> PyResult<Py<PyAny>> {
let hdr = header_at(file, addr, &path)?; let found = handle.with(py, |f| {
match kind(&hdr) { let hdr = header_at(f, addr, &path)?;
Some(Kind::Dataset) => Ok(PyDataset::open(py, Arc::clone(file), path, addr, &hdr)? Ok(match kind(&hdr) {
Some(Kind::Dataset) => Found::Dataset(DatasetMeta::load(f, addr, &hdr, &path)?),
Some(Kind::Group) => Found::Group,
Some(Kind::Datatype) => Found::Datatype,
None => Found::Other,
})
})?;
match found {
Found::Dataset(meta) => Ok(PyDataset::new(py, Arc::clone(handle), path, addr, meta)
.into_pyobject(py)? .into_pyobject(py)?
.into_any() .into_any()
.unbind()), .unbind()),
Some(Kind::Group) => Ok(PyGroup::from_read(Arc::clone(file), path, addr) Found::Group => Ok(PyGroup::from_read(Arc::clone(handle), path, addr)
.into_pyobject(py)? .into_pyobject(py)?
.into_any() .into_any()
.unbind()), .unbind()),
Some(Kind::Datatype) => Err(PyTypeError::new_err(format!( Found::Datatype => Err(PyTypeError::new_err(format!(
"{}: committed (named) datatypes are not supported by clawhdf5", "{}: committed (named) datatypes are not supported by clawhdf5",
name(&path) name(&path)
))), ))),
None => Err(PyValueError::new_err(format!( Found::Other => Err(PyValueError::new_err(format!(
"{}: not a dataset, group or datatype", "{}: not a dataset, group or datatype",
name(&path) name(&path)
))), ))),
@@ -126,32 +159,26 @@ pub(crate) fn open(
} }
/// The dataspace message of an object header. /// The dataspace message of an object header.
pub(crate) fn dataspace(file: &clawhdf5_rs::File, hdr: &ObjectHeader) -> PyResult<Dataspace> { pub(crate) fn dataspace(file: &File, hdr: &ObjectHeader, path: &str) -> PyResult<Dataspace> {
crate::no_panic(|| { let sb = file.superblock();
let sb = file.superblock(); let msg = hdr
let msg = hdr .messages
.messages .iter()
.iter() .find(|m| m.msg_type == MessageType::Dataspace)
.find(|m| m.msg_type == MessageType::Dataspace) .ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?; let data = clawhdf5_format::shared_message::message_data_in(
let data = clawhdf5_format::shared_message::message_data( file.storage(),
file.as_bytes(), msg,
msg, sb.offset_size,
sb.offset_size, sb.length_size,
sb.length_size, )
) .map_err(|e| format_err(path, e, value_err))?;
.map_err(|e| PyValueError::new_err(e.to_string()))?; Dataspace::parse(&data, sb.length_size).map_err(|e| format_err(path, e, value_err))
Dataspace::parse(&data, sb.length_size).map_err(|e| PyValueError::new_err(e.to_string()))
})
} }
/// The chunk shape of a chunked dataset (one entry per dataset dimension), /// The chunk shape of a chunked dataset (one entry per dataset dimension),
/// or `None` for other layouts or a layout message that does not parse. /// or `None` for other layouts or a layout message that does not parse.
pub(crate) fn chunk_shape( pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Option<Vec<u64>> {
file: &clawhdf5_rs::File,
hdr: &ObjectHeader,
rank: usize,
) -> Option<Vec<u64>> {
let sb = file.superblock(); let sb = file.superblock();
let msg = hdr let msg = hdr
.messages .messages
@@ -179,24 +206,18 @@ pub(crate) fn is_null(space: &Dataspace) -> bool {
/// The attributes of the object at `addr` (whose path is `path`), sorted by /// The attributes of the object at `addr` (whose path is `path`), sorted by
/// name (h5py's order). Attributes whose messages cannot be parsed are left /// name (h5py's order). Attributes whose messages cannot be parsed are left
/// out, as the facade's `attrs()` does. /// out, as the facade's `attrs()` does.
pub(crate) fn attributes( pub(crate) fn attributes(file: &File, addr: u64, path: &str) -> PyResult<Vec<AttributeMessage>> {
file: &clawhdf5_rs::File,
addr: u64,
path: &str,
) -> PyResult<Vec<AttributeMessage>> {
let hdr = header_at(file, addr, path)?; let hdr = header_at(file, addr, path)?;
crate::no_panic(|| { let sb = file.superblock();
let sb = file.superblock(); let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant( file.storage(),
file.as_bytes(), &hdr,
&hdr, sb.offset_size,
sb.offset_size, sb.length_size,
sb.length_size, )
) .map_err(|e| format_err(path, e, value_err))?;
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))?; attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes())); Ok(attrs)
Ok(attrs)
})
} }
#[cfg(test)] #[cfg(test)]
+135
View File
@@ -1,6 +1,10 @@
"""Shared fixtures for the clawhdf5 Python binding tests.""" """Shared fixtures for the clawhdf5 Python binding tests."""
import os import os
import re
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import pytest import pytest
@@ -16,3 +20,134 @@ def h5py():
pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable") pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable")
pytest.skip("h5py not installed") pytest.skip("h5py not installed")
return mod return mod
# ---------------------------------------------------------------------------
# An HTTP server for remote reads
# ---------------------------------------------------------------------------
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
class RangeServer:
"""A static file server on 127.0.0.1, in a thread of this process, that
answers `Range: bytes=a-b` with 206 and `Content-Range` (the way S3 and
common web servers do), sends an ETag and honours `If-Match`.
- `ranges=False`: ignores `Range` and answers 200 with the whole file,
like a server without range support.
- `down` (set by `close()`): hang up on every request.
- `delay`: seconds to wait before answering each request after the
first `delay_after` ones (a slow network).
- `log`: every request as `(method, path, range header)`.
"""
def __init__(self, root, ranges=True):
self.root = str(root)
self.ranges = ranges
self.delay = 0.0
self.delay_after = 0
self.down = False
self.log = []
self._lock = threading.Lock()
server = self
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def log_message(self, *args): # quiet
pass
def do_HEAD(self):
self._serve(body=False)
def do_GET(self):
self._serve(body=True)
def _serve(self, body):
with server._lock:
server.log.append((self.command, self.path, self.headers.get("Range")))
n = len(server.log)
if server.delay and n > server.delay_after:
time.sleep(server.delay)
if server.down:
# Hang up without an answer (open keep-alive
# connections outlive shutdown(), so close() sets this).
self.close_connection = True
return
path = os.path.join(server.root, self.path.lstrip("/").split("?")[0])
if not os.path.isfile(path):
self.send_response(404)
self.send_header("Content-Length", "0")
self.end_headers()
return
with open(path, "rb") as fh:
data = fh.read()
st = os.stat(path)
etag = f'"{st.st_mtime_ns:x}-{st.st_size:x}"'
want = self.headers.get("If-Match")
if want is not None and want != etag and want != "*":
self.send_response(412)
self.send_header("Content-Length", "0")
self.end_headers()
return
rng = self.headers.get("Range") if server.ranges else None
m = _RANGE.match(rng.strip()) if rng else None
if m and (m.group(1) or m.group(2)):
size = len(data)
if m.group(1):
start = int(m.group(1))
end = int(m.group(2)) if m.group(2) else size - 1
else:
start = max(0, size - int(m.group(2)))
end = size - 1
if start >= size:
self.send_response(416)
self.send_header("Content-Range", f"bytes */{size}")
self.send_header("Content-Length", "0")
self.end_headers()
return
end = min(end, size - 1)
part = data[start : end + 1]
self.send_response(206)
self.send_header("Content-Range", f"bytes {start}-{end}/{size}")
else:
part = data
self.send_response(200)
if server.ranges:
self.send_header("Accept-Ranges", "bytes")
self.send_header("ETag", etag)
self.send_header("Content-Length", str(len(part)))
self.send_header("Content-Type", "application/x-hdf5")
self.end_headers()
if body:
try:
self.wfile.write(part)
except (BrokenPipeError, ConnectionResetError):
pass
self.httpd = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
self.httpd.daemon_threads = True
self.port = self.httpd.server_address[1]
self.thread = threading.Thread(target=self.httpd.serve_forever, daemon=True)
self.thread.start()
def url(self, name):
return f"http://127.0.0.1:{self.port}/{name}"
def requests(self):
with self._lock:
return len(self.log)
def close(self):
self.down = True
self.httpd.shutdown()
self.httpd.server_close()
@pytest.fixture
def range_server(tmp_path):
"""A range-capable server over `tmp_path`."""
server = RangeServer(tmp_path)
yield server
server.close()
+21 -4
View File
@@ -178,15 +178,32 @@ def _write_fixture(h5py, path):
g.attrs["depth"] = np.int8(3) g.attrs["depth"] = np.int8(3)
@pytest.fixture(scope="module") @pytest.fixture(scope="module", params=["local", "http", "http-1k-blocks"])
def pair(h5py, tmp_path_factory): def pair(request, h5py, tmp_path_factory):
path = str(tmp_path_factory.mktemp("h5") / "fixture.h5") """The fixture file through h5py and through clawhdf5: opened locally,
and over HTTP range requests (a local server in this process) with the
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
through many small ranges."""
from conftest import RangeServer
root = tmp_path_factory.mktemp("h5")
path = str(root / "fixture.h5")
_write_fixture(h5py, path) _write_fixture(h5py, path)
theirs = h5py.File(path, "r") theirs = h5py.File(path, "r")
ours = clawhdf5.File(path, "r") server = None
if request.param == "local":
ours = clawhdf5.File(path, "r")
else:
server = RangeServer(root)
if request.param == "http":
ours = clawhdf5.File(server.url("fixture.h5"))
else:
ours = clawhdf5.File.open_url(server.url("fixture.h5"), block_size=1024)
yield ours, theirs, path yield ours, theirs, path
theirs.close() theirs.close()
ours.close() ours.close()
if server is not None:
server.close()
def _all_datasets(h5py, f): def _all_datasets(h5py, f):
+244
View File
@@ -0,0 +1,244 @@
"""Remote files: `clawhdf5.File(url)` / `File.open_url(url, ...)` read over
HTTP range requests (clawhdf5-remote's block cache), against a server in this
process (conftest.RangeServer). Values are compared with h5py reading the
same file locally; the rest checks what the server saw (only the blocks a
read needs are fetched), the failure modes (no range support, a missing
file, a file that changes, a server that goes away: errors, never wrong
data), and that the GIL is released while a read waits on the network."""
import os
import sys
import threading
import time
import numpy as np
import pytest
import clawhdf5
from conftest import RangeServer
def _write(h5py, path):
rng = np.random.default_rng(7)
with h5py.File(path, "w") as f:
f.create_dataset("contig", data=rng.standard_normal((400, 300)))
f.create_dataset(
"chunked",
data=rng.integers(0, 1000, size=(512, 512), dtype="<i4"),
chunks=(64, 64),
compression="gzip",
)
f.create_dataset("strings", data=["alpha", "beta", "gamma"], dtype=h5py.string_dtype())
g = f.create_group("grp")
g.create_dataset("small", data=np.arange(10, dtype="<u2"))
g.attrs["units"] = "m/s"
f.attrs["title"] = "remote test"
@pytest.fixture
def remote_file(h5py, tmp_path, range_server):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
return path, range_server.url("remote.h5"), range_server
def test_remote_reads_match_h5py(h5py, remote_file):
path, url, server = remote_file
with h5py.File(path, "r") as theirs, clawhdf5.File(url) as ours:
assert ours.filename == url
assert list(ours.keys()) == list(theirs.keys())
assert "grp/small" in ours and "nope" not in ours
for name in ["contig", "chunked", "grp/small"]:
np.testing.assert_array_equal(ours[name][...], theirs[name][...])
np.testing.assert_array_equal(ours[name][3:7], theirs[name][3:7])
np.testing.assert_array_equal(ours["chunked"][100:130, 200:300:3], theirs["chunked"][100:130, 200:300:3])
np.testing.assert_array_equal(ours["chunked"][[1, 70, 300], 5], theirs["chunked"][[1, 70, 300], 5])
assert list(ours["strings"][...]) == list(theirs["strings"][...])
assert ours["grp"].attrs["units"] == theirs["grp"].attrs["units"]
assert ours.attrs["title"] == theirs.attrs["title"]
assert ours["chunked"].maxshape == theirs["chunked"].maxshape
assert server.requests() >= 2
def test_a_small_read_fetches_only_its_blocks(h5py, remote_file):
"""With 4 KiB blocks, opening and reading one chunk of a 1 MB chunked
dataset costs a handful of requests and a few blocks, not the file."""
path, url, server = remote_file
size = os.path.getsize(path)
f = clawhdf5.File.open_url(url, block_size=4096)
opened = server.requests()
assert opened == 1, server.log
ds = f["chunked"]
got = ds[0:10, 0:10]
with h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(got, theirs["chunked"][0:10, 0:10])
stats = f.remote_stats
assert stats["bytes_fetched"] < size / 4, (stats, size)
assert server.requests() - opened <= 12, server.log
# A second read of the same region is served by the cache.
before = server.requests()
ds[0:10, 0:10]
assert server.requests() == before
assert f.remote_stats["hits"] > stats["hits"]
assert clawhdf5.File(str(path), "r").remote_stats is None
def test_server_without_range_support(h5py, tmp_path):
"""A server that ignores Range answers 200 with the whole file: that is
an OSError by default, and a whole download when allowed."""
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path, ranges=False)
try:
url = server.url("remote.h5")
with pytest.raises(OSError, match="range"):
clawhdf5.File(url)
with clawhdf5.File.open_url(url, allow_full_download=True) as ours, h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(ours["chunked"][...], theirs["chunked"][...])
np.testing.assert_array_equal(ours["contig"][5], theirs["contig"][5])
with pytest.raises(OSError):
clawhdf5.File.open_url(url, allow_full_download=True, max_full_download=1000)
finally:
server.close()
def test_errors_are_oserrors(remote_file):
_, url, server = remote_file
with pytest.raises(OSError, match="404"):
clawhdf5.File(server.url("missing.h5"))
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "r+")
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "w")
with pytest.raises(OSError, match="unsupported URL"):
clawhdf5.File("nosuchscheme://x/y.h5")
with pytest.raises(ValueError):
clawhdf5.File.open_url(url, block_size=0)
with pytest.raises(TypeError):
clawhdf5.File.open_url(url, no_such_option=1)
def test_object_store_urls_need_their_features():
"""The default wheel has no S3/GCS/Azure clients (aws-lc-rs builds C):
such a URL is an OSError naming the build feature."""
for url, feature in [("s3://bucket/k.h5", "s3"), ("gs://b/k.h5", "gcs"), ("az://c/k.h5", "azure")]:
try:
clawhdf5.File(url)
except OSError as e:
if "feature" in str(e):
assert f"`{feature}`" in str(e), str(e)
else:
pytest.fail(f"{url} opened")
def test_https_needs_the_https_feature():
"""The default wheel has no TLS stack (rustls needs ring, which builds C):
an https URL is an OSError that names the build feature."""
with pytest.raises(OSError) as e:
clawhdf5.File("https://127.0.0.1:1/x.h5")
msg = str(e.value)
# Built with `--features https` the error is the refused connection.
assert "https" in msg or "connect" in msg.lower() or "refused" in msg.lower(), msg
def test_a_changed_file_is_an_error_not_mixed_data(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=1024)
first = f["grp/small"][...]
# Rewrite the file with other values: new ETag, same name.
time.sleep(0.01)
with h5py.File(path, "w") as g:
g.create_dataset("contig", data=np.zeros((400, 300)))
with pytest.raises(OSError, match="changed"):
f["contig"][...]
np.testing.assert_array_equal(first, np.arange(10, dtype="<u2"))
def test_a_server_that_goes_away_is_an_error(h5py, tmp_path):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path)
f = clawhdf5.File.open_url(server.url("remote.h5"), block_size=1024, retries=0, timeout=2)
ds = f["contig"]
server.close()
with pytest.raises(OSError):
ds[...]
def test_threads_read_one_remote_file(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=2048)
with h5py.File(path, "r") as theirs:
expected = theirs["chunked"][...]
errors = []
def work(i):
try:
rows = slice((i * 37) % 400, (i * 37) % 400 + 64)
np.testing.assert_array_equal(f["chunked"][rows], expected[rows])
except Exception as e: # noqa: BLE001
errors.append(e)
threads = [threading.Thread(target=work, args=(i,)) for i in range(16)]
for t in threads:
t.start()
for t in threads:
t.join()
assert not errors, errors[:3]
def test_remote_reads_release_the_gil(h5py, remote_file):
"""A read waiting on a slow server lets other Python threads run: a
thread counting in a loop keeps counting (and never stalls for long)
while the main thread reads through requests that each take 0.2 s."""
_, url, server = remote_file
f = clawhdf5.File.open_url(url, block_size=1024, max_parallel=1)
ds = f["contig"]
server.delay = 0.2
server.delay_after = server.requests()
old = sys.getswitchinterval()
sys.setswitchinterval(0.001)
stop = threading.Event()
progress = {"n": 0, "worst": 0.0}
def spin():
last = time.perf_counter()
while not stop.is_set():
now = time.perf_counter()
progress["worst"] = max(progress["worst"], now - last)
last = now
progress["n"] += 1
t = threading.Thread(target=spin)
try:
t.start()
time.sleep(0.02)
t0 = time.perf_counter()
before = server.requests()
ds[0:2]
took = time.perf_counter() - t0
stop.set()
t.join()
finally:
sys.setswitchinterval(old)
assert server.requests() > before
assert took >= 0.2, took
assert progress["n"] > 1000, progress
# Held across a 0.2 s request, the spinner would stall that long.
assert progress["worst"] < 0.1, (progress, took)
def test_a_clawhdf5_written_file_reads_the_same_remotely(tmp_path, range_server):
path = tmp_path / "ours.h5"
data = np.arange(3000, dtype="<f8").reshape(100, 30)
with clawhdf5.File(str(path), "w") as f:
f.create_dataset("d", data=data, chunks=[10, 30], compression="gzip")
g = f.create_group("g")
g.create_dataset("i", data=np.arange(5, dtype="<i4"))
f.attrs["k"] = 3
with clawhdf5.File(range_server.url("ours.h5")) as f, clawhdf5.File(str(path)) as local:
np.testing.assert_array_equal(f["d"][...], data)
np.testing.assert_array_equal(f["d"][5:9, ::4], local["d"][5:9, ::4])
np.testing.assert_array_equal(f["g/i"][...], np.arange(5))
assert f.attrs["k"] == local.attrs["k"]
assert "file (read" in repr(f).lower() and "127.0.0.1" in repr(f)
+15 -3
View File
@@ -7,8 +7,9 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
`File::open_storage` gives the facade's read API over any `Storage` (see `File::open_storage` gives the facade's read API over any `Storage` (see
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch `CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S), `feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is object stores) and URLs in `h5rs` (see the M3 status below); the Python
next. Every count in §1–§2 was bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
change. Progress: M1, first part (the `Storage` trait and the metadata change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done; parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
@@ -475,7 +476,7 @@ fast path within benchmark noise.
a request counter exposed for tests and users. a request counter exposed for tests and users.
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it. - Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the - *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
Python bindings, with these choices: Python bindings (done 2026-09-27, below), with these choices:
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of - A new crate, `clawhdf5-remote`, instead of a `remote` feature of
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io` `clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
sits below the facade. sits below the facade.
@@ -512,6 +513,17 @@ fast path within benchmark noise.
same work without a cache: 141 936 requests. Per file: A lists in 2 same work without a cache: 141 936 requests. Per file: A lists in 2
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
6.4 MB: 35 001 object headers spread over the file). 6.4 MB: 35 001 object headers spread over the file).
- *Status 2026-09-27, Python bindings:* done on branch
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
`File.open_url(url, **options)` (cache and HTTP options) go through
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
parsing (path lookups, headers, attributes, listings, the global heap)
moved from `File::as_bytes` to `File::storage()` and the `*_in`
functions, and every file access, metadata included, runs with the GIL
released. Checked by running the whole read-vs-h5py suite over an
in-process range server (1 MiB and 1 KiB blocks) and by request counts
in `crates/clawhdf5-py/tests/test_remote.py`.
**M4 — wasm lazy loading (1–2 weeks).** **M4 — wasm lazy loading (1–2 weeks).**
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a - `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
+9 -6
View File
@@ -844,9 +844,9 @@ cache, but:
`read_*_zerocopy`) need the file in memory and answer `read_*_zerocopy`) need the file in memory and answer
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()` `FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
panics for such a file (`File::contiguous_bytes()` is the fallible form). panics for such a file (`File::contiguous_bytes()` is the fallible form).
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a `LazyFile`, `MmapFile` and the wasm bindings still read a whole file
whole file (`h5rs` reads through `File::storage`, and takes URLs with its (`h5rs` and the Python bindings read through `File::storage`, and take
`remote` feature). URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
- The file's length is read once, at open: a growing file (SWMR) is not - The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`. grows is `RemoteError::FileChanged`.
@@ -862,9 +862,12 @@ cache, but:
**Status:** open (added 2026-09-26, milestone M3 of **Status:** open (added 2026-09-26, milestone M3 of
`docs/design/range-reads.md`). `docs/design/range-reads.md`).
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3) - **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
parses through `File::as_bytes`, which a remote file does not have; the milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
wasm reader's `openUrl` is milestone M4. the default wheel reads plain `http://` only: `https://` needs a wheel
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise). - **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead. is not implemented, and only the first block is read ahead.
+5 -2
View File
@@ -131,14 +131,17 @@ run_step "cargo clippy (h5rs remote)" cargo clippy \
# js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C. # js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C.
# clawhdf5-remote is checked by default (plain HTTP) and with its # clawhdf5-remote is checked by default (plain HTTP) and with its
# object-store feature, and h5rs with URL support (remote); the https # object-store feature, and h5rs with URL support (remote); the https
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. # (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. The
# Python bindings (clawhdf5-py, remote reads over plain HTTP) are checked too:
# their https/s3/gcs/azure features are opt-in for the same reason.
no_c_in_default_build() { no_c_in_default_build() {
local entry crate features found=0 local entry crate features found=0
for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \ for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \
clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \ clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \
clawhdf5-tools \ clawhdf5-tools \
clawhdf5-wasm \ clawhdf5-wasm \
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote; do clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote \
clawhdf5-py; do
crate=${entry%%:*} crate=${entry%%:*}
features=() features=()
[ "$entry" != "$crate" ] && features=(--features "${entry#*:}") [ "$entry" != "$crate" ] && features=(--features "${entry#*:}")