py: remote files (clawhdf5.File(url), File.open_url) through File::storage()

The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.

- Every object of a File now shares one handle (src/handle.rs) that
  runs all file access, metadata included, with the GIL released and
  parses through File::storage() and the clawhdf5_format *_in functions.
  Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
  clawhdf5_remote::storage_for_url (read-only; another mode is a
  ValueError). File.open_url(url, **options) takes the cache and HTTP
  options (block_size, cache_size, headers, retries, timeout,
  allow_full_download, max_full_download, require_validator,
  max_redirects, max_parallel); File.remote_stats gives the block
  cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
  s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
  ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
  OSError, never KeyError/ValueError and never data; `key in group`
  raises it instead of answering False.

Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:40:44 -05:00
co-authored by Claude Opus 5.5
parent a4c2aced55
commit 910d81904c
18 changed files with 1146 additions and 299 deletions
+30
View File
@@ -2,6 +2,36 @@
## Unreleased
### Python bindings: remote files (2026-09-27)
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
through the block cache, the whole read API (groups, attributes, every
dataset type and index the local reader handles). A URL is any
`scheme://…`; a remote file is read-only (another mode is a
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
`allow_full_download`, `max_full_download`, `require_validator`,
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
cache's counters. The default wheel builds plain HTTP only (no C: rustls
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
check now covers `clawhdf5-py`.
- **Every read parses through `File::storage()`** instead of
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
group listings, variable-length data through the global heap), inside one
shared file handle that releases the GIL for all file access, not only
dataset reads: a read waiting on the network lets other Python threads
run. A failed read of the storage (a network error, a file changed on
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
data; `key in group` raises it instead of answering `False`.
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
blocks and 1 KiB blocks) against a range-capable `http.server` in the
test process, plus `tests/test_remote.py`: request counts of a small
read, cache hits, a server without `Range` support (refused, or a
whole download when allowed), a file changed on the server, a server
that hangs up, 16 threads on one remote file, and a thread that keeps
running while a read waits on 0.2 s requests.
### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
+12
View File
@@ -533,8 +533,20 @@ with clawhdf5.File("data.h5", "r") as f:
records = f["table"] # compound -> numpy structured array
ids = records["id"] # one field
# A file on a web server: range requests through a block cache, nothing
# downloaded up front; the same read API. The GIL is released while waiting.
with clawhdf5.File("http://data.example.org/run42.h5") as f:
first = f["group/temperatures"][0]
f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024,
headers={"Authorization": "Bearer ..."})
```
The default build reads `http://` URLs only; build with
`maturin develop --release --features https` (rustls with ring, which
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
object-store URLs.
Reads cover integers and IEEE floats of every width in either byte order,
`bool`, enums, complex, fixed and variable-length strings, variable-length
sequences, opaque, HDF5 array types and compounds; other types (references,
+9
View File
@@ -17,11 +17,20 @@ crate-type = ["cdylib", "rlib"]
[dependencies]
clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" }
clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
# Remote files (`clawhdf5.File(url)`): plain HTTP by default, which builds
# no C. HTTPS and the object stores are opt-in features below.
clawhdf5-remote = { path = "../clawhdf5-remote", version = "2.7.0" }
pyo3 = "0.29"
numpy = "0.29"
[features]
extension-module = ["pyo3/extension-module"]
# https:// URLs (rustls with ring, which compiles C and assembly).
https = ["clawhdf5-remote/https"]
# s3://, gs://, az:// URLs (object_store; its cloud clients build aws-lc-rs, C).
s3 = ["clawhdf5-remote/s3"]
gcs = ["clawhdf5-remote/gcs"]
azure = ["clawhdf5-remote/azure"]
[package.metadata.docs.rs]
features = []
+34 -1
View File
@@ -61,6 +61,36 @@ with clawhdf5.File("data.h5", "r") as f:
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
dataspace (h5py's `Empty`).
## Remote files
A URL instead of a path reads the file where it is, through
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
```python
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
```
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise `OSError` instead of mixing versions.
Network failures are `OSError` too.
- Remote files are read-only.
- Schemes: the default build (no C) reads `http://`. `https://` needs
`maturin develop --release --features https` (rustls with ring, which
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
`azure` features (credentials from the environment; aws-lc-rs, C).
## Writing
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
@@ -76,7 +106,10 @@ pytest crates/clawhdf5-py/tests
```
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
writes. `scripts/ci-test.sh` builds the wheel and runs these in CI.
writes, opened locally and over HTTP (an in-process range server,
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
failures, the GIL). `scripts/ci-test.sh` builds the wheel and runs these
in CI.
## License
+23 -19
View File
@@ -8,13 +8,14 @@ use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple};
use crate::convert::{Converter, Elements, resolve_vl};
use crate::handle::Handle;
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value};
/// Backing storage for attributes.
enum AttrsInner {
/// Attributes of an object in a file opened for reading, sorted by name.
Read {
file: Arc<clawhdf5_rs::File>,
handle: Arc<Handle>,
attrs: Vec<AttributeMessage>,
},
/// Writable attribute list shared with a parent (PyFile or PyGroup).
@@ -36,10 +37,15 @@ pub struct PyAttrs {
impl PyAttrs {
/// The attributes of the object at `addr` (whose path is `path`) in a
/// file opened for reading.
pub(crate) fn read(file: Arc<clawhdf5_rs::File>, addr: u64, path: &str) -> PyResult<Self> {
let attrs = node::attributes(&file, addr, path)?;
pub(crate) fn read(
py: Python<'_>,
handle: Arc<Handle>,
addr: u64,
path: &str,
) -> PyResult<Self> {
let attrs = handle.with(py, |f| node::attributes(f, addr, path))?;
Ok(Self {
inner: AttrsInner::Read { file, attrs },
inner: AttrsInner::Read { handle, attrs },
})
}
@@ -55,8 +61,8 @@ impl PyAttrs {
impl PyAttrs {
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
match &self.inner {
AttrsInner::Read { file, attrs } => match attrs.iter().find(|a| a.name == key) {
Some(attr) => Ok(attr_to_py(py, file, attr)?.unbind()),
AttrsInner::Read { handle, attrs } => match attrs.iter().find(|a| a.name == key) {
Some(attr) => Ok(attr_to_py(py, handle, attr)?.unbind()),
None => Err(PyKeyError::new_err(format!(
"Can't open attribute (can't locate attribute: '{key}')"
))),
@@ -146,9 +152,9 @@ impl PyAttrs {
/// Return attribute values as a list.
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let vals: Vec<Py<PyAny>> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs
AttrsInner::Read { handle, attrs } => attrs
.iter()
.map(|a| attr_to_py(py, file, a).map(Bound::unbind))
.map(|a| attr_to_py(py, handle, a).map(Bound::unbind))
.collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store
.lock()
@@ -167,9 +173,9 @@ impl PyAttrs {
/// Return attribute (key, value) pairs as a list of tuples.
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs
AttrsInner::Read { handle, attrs } => attrs
.iter()
.map(|a| Ok((a.name.clone(), attr_to_py(py, file, a)?.unbind())))
.map(|a| Ok((a.name.clone(), attr_to_py(py, handle, a)?.unbind())))
.collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store
.lock()
@@ -189,12 +195,11 @@ impl PyAttrs {
/// An attribute's value as h5py returns it.
fn attr_to_py<'py>(
py: Python<'py>,
file: &clawhdf5_rs::File,
handle: &Handle,
attr: &AttributeMessage,
) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| {
let sb = file.superblock();
let conv = Converter::new(py, &attr.datatype, sb.offset_size)
let conv = Converter::new(py, &attr.datatype, handle.offset_size)
.map_err(|e| prefix_err(py, &attr.name, e))?;
if node::is_null(&attr.dataspace) {
return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any());
@@ -216,12 +221,11 @@ fn attr_to_py<'py>(
)));
}
let raw = &attr.raw_data[..want];
let file_data = file.as_bytes();
let (osz, lsz, unit) = (sb.offset_size, sb.length_size, conv.vl_unit);
Elements::Vl(
py.detach(|| resolve_vl(file_data, raw, n, osz, lsz, unit))
.map_err(|e| PyValueError::new_err(format!("attribute {}: {e}", attr.name)))?,
)
let (osz, lsz, unit) = (handle.offset_size, handle.length_size, conv.vl_unit);
let what = format!("attribute {}", attr.name);
Elements::Vl(handle.with(py, |f| {
resolve_vl(f.storage(), raw, n, osz, lsz, unit).map_err(|e| e.into_py(&what))
})?)
} else {
Elements::Bytes(attr.raw_data.clone())
};
+49 -21
View File
@@ -17,6 +17,7 @@ use std::collections::HashMap;
use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder};
use clawhdf5_format::global_heap::GlobalHeapCollection;
use clawhdf5_format::storage::Storage;
use numpy::PyArray1;
use pyo3::exceptions::{PyTypeError, PyValueError};
use pyo3::prelude::*;
@@ -538,16 +539,16 @@ fn object_array<'py>(
/// their bytes: each element's stored length times `unit` (1 for strings,
/// the base type's size for sequences). Pure Rust, so it runs without the
/// GIL.
pub(crate) fn resolve_vl(
file_data: &[u8],
pub(crate) fn resolve_vl<S: Storage + ?Sized>(
file: &S,
raw: &[u8],
count: usize,
offset_size: u8,
length_size: u8,
unit: usize,
) -> Result<Vec<Vec<u8>>, String> {
) -> Result<Vec<Vec<u8>>, VlError> {
let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size)
.map_err(|e| e.to_string())?;
.map_err(|e| VlError::Invalid(e.to_string()))?;
let undefined = match offset_size {
2 => 0xFFFF,
4 => 0xFFFF_FFFF,
@@ -558,43 +559,70 @@ pub(crate) fn resolve_vl(
for vl in &refs {
if vl.collection_address == 0 || vl.collection_address == undefined {
if vl.length != 0 {
return Err(format!(
return Err(VlError::Invalid(format!(
"variable-length element of length {} has no heap address",
vl.length
));
)));
}
out.push(Vec::new());
continue;
}
let coll = match collections.entry(vl.collection_address) {
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
std::collections::hash_map::Entry::Vacant(e) => {
let addr = usize::try_from(vl.collection_address)
.map_err(|_| "global heap address out of range".to_string())?;
e.insert(
GlobalHeapCollection::parse(file_data, addr, length_size)
.map_err(|e| e.to_string())?,
)
}
std::collections::hash_map::Entry::Vacant(e) => e.insert(
GlobalHeapCollection::parse_in(file, vl.collection_address, length_size)
.map_err(VlError::from_format)?,
),
};
let index = u16::try_from(vl.object_index)
.map_err(|_| format!("global heap object index {} out of range", vl.object_index))?;
let index = u16::try_from(vl.object_index).map_err(|_| {
VlError::Invalid(format!(
"global heap object index {} out of range",
vl.object_index
))
})?;
let obj = coll.get_object(index).ok_or_else(|| {
format!(
VlError::Invalid(format!(
"global heap object {index} not found in the collection at {}",
vl.collection_address
)
))
})?;
let need = (vl.length as usize)
.checked_mul(unit)
.ok_or("variable-length element too long")?;
.ok_or_else(|| VlError::Invalid("variable-length element too long".into()))?;
if need > obj.data.len() {
return Err(format!(
return Err(VlError::Invalid(format!(
"variable-length element of {need} bytes in a {}-byte heap object",
obj.data.len()
));
)));
}
out.push(obj.data[..need].to_vec());
}
Ok(out)
}
/// Why variable-length elements could not be resolved.
#[derive(Debug)]
pub(crate) enum VlError {
/// Reading the file failed (a network error on a remote file).
Storage(String),
/// The references or the heap are not valid.
Invalid(String),
}
impl VlError {
fn from_format(e: clawhdf5_format::error::FormatError) -> Self {
match e {
clawhdf5_format::error::FormatError::Storage(_) => VlError::Storage(e.to_string()),
e => VlError::Invalid(e.to_string()),
}
}
/// As a Python exception, the message prefixed with `what`: a storage
/// failure is an `OSError`, anything else a `ValueError`.
pub(crate) fn into_py(self, what: &str) -> PyErr {
match self {
VlError::Storage(m) => pyo3::exceptions::PyOSError::new_err(format!("{what}: {m}")),
VlError::Invalid(m) => PyValueError::new_err(format!("{what}: {m}")),
}
}
}
+72 -42
View File
@@ -6,21 +6,53 @@
//! whole dataset instead); the
//! bytes it returns become the numpy array's buffer without a copy (see
//! `convert`). All file access and decoding runs with the GIL released, so
//! Python threads reading the same or different datasets run in parallel.
//! Python threads reading the same or different datasets run in parallel,
//! and a remote file's network reads never hold the GIL.
use std::sync::Arc;
use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_rs::File;
use pyo3::exceptions::{PyTypeError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple};
use crate::attrs::PyAttrs;
use crate::convert::{Converter, Elements, resolve_vl};
use crate::convert::{Converter, Elements, VlError, resolve_vl};
use crate::handle::Handle;
use crate::select::{self, Plan};
use crate::{PyEmpty, node, to_py_err};
/// What opening a dataset reads from the file (without the GIL).
pub(crate) struct DatasetMeta {
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
shape: Option<Vec<u64>>,
chunks: Option<Vec<u64>>,
datatype: Datatype,
}
impl DatasetMeta {
pub(crate) fn load(f: &File, addr: u64, hdr: &ObjectHeader, path: &str) -> PyResult<Self> {
let null = node::is_null(&node::dataspace(f, hdr, path)?);
let ds = f.dataset_at(addr).map_err(to_py_err)?;
let shape = if null {
None
} else {
Some(ds.shape().map_err(to_py_err)?)
};
let datatype = ds.raw_datatype().map_err(to_py_err)?;
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(f, hdr, s.len()));
Ok(Self {
shape,
chunks,
datatype,
})
}
}
/// A dataset in a file opened for reading.
///
/// ```python
@@ -30,7 +62,7 @@ use crate::{PyEmpty, node, to_py_err};
/// ```
#[pyclass(name = "Dataset")]
pub struct PyDataset {
file: Arc<clawhdf5_rs::File>,
handle: Arc<Handle>,
path: String,
/// Where the dataset's object header is: reads open it from here rather
/// than resolve `path` again.
@@ -45,39 +77,24 @@ pub struct PyDataset {
}
impl PyDataset {
pub(crate) fn open(
pub(crate) fn new(
py: Python<'_>,
file: Arc<clawhdf5_rs::File>,
handle: Arc<Handle>,
path: String,
addr: u64,
hdr: &ObjectHeader,
) -> PyResult<Self> {
crate::no_panic(|| {
let null = node::is_null(&node::dataspace(&file, hdr)?);
let (shape, datatype) = {
let ds = file.dataset_at(addr).map_err(to_py_err)?;
let shape = if null {
None
} else {
Some(ds.shape().map_err(to_py_err)?)
};
(shape, ds.raw_datatype().map_err(to_py_err)?)
};
let conv = Converter::new(py, &datatype, file.superblock().offset_size)
meta: DatasetMeta,
) -> Self {
let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
.map_err(|e| e.value(py).to_string());
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(&file, hdr, s.len()));
Ok(Self {
file,
Self {
handle,
path,
addr,
shape,
chunks,
datatype,
shape: meta.shape,
chunks: meta.chunks,
datatype: meta.datatype,
conv,
})
})
}
}
fn converter(&self) -> PyResult<&Converter> {
@@ -103,10 +120,10 @@ impl PyDataset {
};
let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size);
let read_shape = plan.read_shape();
let file = &*self.file;
let handle = &*self.handle;
let addr = self.addr;
// Everything below touches only Rust data: release the GIL.
let read = || -> Result<Elements, ReadError> {
let read = |file: &File| -> Result<Elements, ReadError> {
let ds = file.dataset_at(addr)?;
let mut blocks = Vec::with_capacity(reads.len());
for read in reads {
@@ -147,7 +164,7 @@ impl PyDataset {
let sb = file.superblock();
let n = read_shape.iter().product();
resolve_vl(
file.as_bytes(),
file.storage(),
&raw,
n,
sb.offset_size,
@@ -155,12 +172,20 @@ impl PyDataset {
unit,
)
.map(Elements::Vl)
.map_err(ReadError::Other)
.map_err(ReadError::Vl)
};
let data = py
.detach(|| {
std::panic::catch_unwind(std::panic::AssertUnwindSafe(read))
.unwrap_or_else(|p| Err(ReadError::Panic(crate::panic_text(&*p))))
handle
.with_detached(|f| {
Ok(
std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| read(f)))
.unwrap_or_else(|p| {
Err(ReadError::Panic(crate::panic_text(&*p)))
}),
)
})
.unwrap_or_else(|e| Err(ReadError::Py(e)))
})
.map_err(|e| e.into_py(&self.path))?;
let joined = conv.to_array(py, data, &read_shape, false)?;
@@ -183,6 +208,8 @@ impl PyDataset {
/// An error from the read closure, turned into a Python error with the GIL.
enum ReadError {
Lib(clawhdf5_rs::Error),
Vl(VlError),
Py(PyErr),
Other(String),
Panic(String),
}
@@ -197,6 +224,8 @@ impl ReadError {
fn into_py(self, path: &str) -> PyErr {
match self {
ReadError::Lib(e) => to_py_err(e),
ReadError::Vl(e) => e.into_py(&node::name(path)),
ReadError::Py(e) => e,
ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))),
ReadError::Panic(msg) => crate::InternalError::new_err(format!(
"{}: clawhdf5 internal error (please report it): {msg}",
@@ -252,22 +281,23 @@ impl PyDataset {
/// The maximum shape (`None` per unlimited dimension), like h5py.
#[getter]
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| {
let Some(shape) = &self.shape else {
return Ok(py.None().into_bound(py));
};
let addr = self.addr;
let max = self
.file
.dataset_at(self.addr)
.handle
.with(py, |f| {
f.dataset_at(addr)
.and_then(|ds| ds.max_dimensions())
.map_err(to_py_err)?
.map_err(to_py_err)
})?
.unwrap_or_else(|| shape.clone());
let items: Vec<Option<u64>> = max
.into_iter()
.map(|d| (d != u64::MAX).then_some(d))
.collect();
Ok(PyTuple::new(py, items)?.into_any())
})
}
/// The dataset's numpy dtype, as h5py reports it.
@@ -295,8 +325,8 @@ impl PyDataset {
/// The dataset's attributes (read-only, dict-like).
#[getter]
fn attrs(&self) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
}
/// Read with h5py indexing: integers, slices with positive steps,
+177 -29
View File
@@ -1,13 +1,17 @@
//! PyFile — the main entry point for opening and creating HDF5 files.
use std::collections::HashMap;
use std::path::PathBuf;
use std::sync::{Arc, Mutex};
use std::time::Duration;
use pyo3::exceptions::PyValueError;
use pyo3::prelude::*;
use pyo3::types::PyList;
use pyo3::types::{PyDict, PyList};
use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
/// Internal state for write mode.
@@ -23,8 +27,10 @@ struct WriteState {
/// Mirrors the h5py.File interface:
///
/// ```python
/// # Reading
/// # Reading, a local file or a URL (range requests, nothing downloaded
/// # up front)
/// f = clawhdf5.File('data.h5', 'r')
/// f = clawhdf5.File('https://example.org/data.h5')
/// ds = f['dataset']
/// f.close()
///
@@ -44,27 +50,51 @@ enum FileInner {
Write(WriteState),
}
/// Whether `s` is a URL (`scheme://…`) rather than a path: the scheme is a
/// letter followed by letters, digits, `+`, `-` or `.` (RFC 3986).
fn is_url(s: &str) -> bool {
let Some((scheme, _)) = s.split_once("://") else {
return false;
};
let mut chars = scheme.chars();
chars.next().is_some_and(|c| c.is_ascii_alphabetic())
&& chars.all(|c| c.is_ascii_alphanumeric() || matches!(c, '+' | '-' | '.'))
}
impl PyFile {
fn from_handle(handle: Arc<Handle>, filename: String) -> Self {
let root = handle.root;
Self {
inner: Some(FileInner::Read(ReadGroup::new(handle, String::new(), root))),
filename,
}
}
}
#[pymethods]
impl PyFile {
/// Open or create an HDF5 file.
///
/// Parameters:
/// path: file path
/// path: file path, or a URL (`http://`, `https://`, `s3://`, `gs://`,
/// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write
#[new]
#[pyo3(signature = (path, mode="r"))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
let filename = path.to_string();
match mode {
"r" => {
let file = py.detach(|| {
crate::no_panic(|| clawhdf5_rs::File::open(path).map_err(to_py_err))
})?;
Ok(Self {
inner: Some(FileInner::Read(root_group(Arc::new(file)))),
filename,
})
if is_url(path) {
if mode != "r" {
return Err(PyValueError::new_err(format!(
"remote files are read-only: mode '{mode}' is not supported for a URL"
)));
}
let handle = Handle::open_url(py, path, &clawhdf5_remote::Options::default())?;
return Ok(Self::from_handle(handle, filename));
}
match mode {
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
"w" => Ok(Self {
filename,
inner: Some(FileInner::Write(WriteState {
@@ -74,12 +104,122 @@ impl PyFile {
groups: Vec::new(),
})),
}),
other => Err(PyErr::new::<pyo3::exceptions::PyValueError, _>(format!(
other => Err(PyValueError::new_err(format!(
"unsupported mode '{other}'; expected 'r' or 'w'"
))),
}
}
/// Open a remote file for reading, with options.
///
/// The file is read through a block cache with range requests: opening
/// costs one request (it also fetches the first block), and a read
/// fetches only the blocks it needs. The GIL is released while waiting
/// on the network.
///
/// Parameters (all optional):
/// block_size: bytes per cached block (default 1 MiB)
/// cache_size: byte budget of the block cache (default 64 MiB)
/// headers: dict of extra HTTP headers (e.g. Authorization), sent only
/// to the URL's own origin
/// retries: retries of a request that failed transiently (default 3)
/// timeout: seconds to connect and receive response headers (default 30)
/// allow_full_download: when the server ignores Range requests,
/// download the whole file once instead of failing (default False)
/// max_full_download: largest file such a download may fetch
/// (default 1 GiB)
/// require_validator: refuse a server that sends neither ETag nor
/// Last-Modified (default False)
/// max_redirects: redirects followed per request (default 5)
/// max_parallel: requests of one read in flight at once (default 8)
#[staticmethod]
#[allow(clippy::too_many_arguments)]
#[pyo3(signature = (url, *, block_size=None, cache_size=None, headers=None, retries=None,
timeout=None, allow_full_download=None, max_full_download=None,
require_validator=None, max_redirects=None, max_parallel=None))]
fn open_url(
py: Python<'_>,
url: &str,
block_size: Option<u64>,
cache_size: Option<u64>,
headers: Option<HashMap<String, String>>,
retries: Option<u32>,
timeout: Option<f64>,
allow_full_download: Option<bool>,
max_full_download: Option<u64>,
require_validator: Option<bool>,
max_redirects: Option<u32>,
max_parallel: Option<usize>,
) -> PyResult<Self> {
let mut options = clawhdf5_remote::Options::default();
if let Some(b) = block_size {
if b == 0 {
return Err(PyValueError::new_err("block_size must be positive"));
}
options.cache.block_size = b;
options.cache.coalesce_gap = b;
// The opening request fetches the first block, not 1 MiB.
options.http.first_request = b;
}
if let Some(c) = cache_size {
options.cache.capacity = c;
}
let http = &mut options.http;
if let Some(h) = headers {
http.headers = h.into_iter().collect();
}
if let Some(r) = retries {
http.retries = r;
}
if let Some(t) = timeout {
if !(t.is_finite() && t > 0.0) {
return Err(PyValueError::new_err("timeout must be a positive number"));
}
http.timeout = Duration::from_secs_f64(t);
}
if let Some(a) = allow_full_download {
http.allow_full_download = a;
}
if let Some(m) = max_full_download {
http.max_full_download = m;
}
if let Some(v) = require_validator {
http.require_validator = v;
}
if let Some(r) = max_redirects {
http.max_redirects = r;
}
if let Some(p) = max_parallel {
if p == 0 {
return Err(PyValueError::new_err("max_parallel must be positive"));
}
http.max_parallel = p;
}
let handle = Handle::open_url(py, url, &options)?;
Ok(Self::from_handle(handle, url.to_string()))
}
/// For a remote file, what its block cache has done so far (reads,
/// hits, misses, requests, bytes fetched, ...); `None` for a local file.
#[getter]
fn remote_stats<'py>(&self, py: Python<'py>) -> PyResult<Option<Bound<'py, PyDict>>> {
let Some(storage) = self.read_file()?.handle.remote_storage() else {
return Ok(None);
};
let s = storage.stats();
let d = PyDict::new(py);
d.set_item("reads", s.reads)?;
d.set_item("hits", s.hits)?;
d.set_item("misses", s.misses)?;
d.set_item("waits", s.waits)?;
d.set_item("requests", s.requests)?;
d.set_item("fetch_calls", s.fetch_calls)?;
d.set_item("bytes_fetched", s.bytes_fetched)?;
d.set_item("evictions", s.evictions)?;
d.set_item("cached_bytes", s.cached_bytes)?;
Ok(Some(d))
}
/// Close the file. In write mode, this finalizes and writes the file.
fn close(&mut self) -> PyResult<()> {
let inner = self.inner.take().ok_or_else(|| {
@@ -121,7 +261,7 @@ impl PyFile {
/// List the names of all children in the root group.
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let names = self.read_file()?.member_names()?;
let names = self.read_file()?.member_names(py)?;
Ok(PyList::new(py, names)?.into_any().unbind())
}
@@ -139,8 +279,8 @@ impl PyFile {
self.keys(py)?.call_method0(py, "__iter__")
}
fn __len__(&self) -> PyResult<usize> {
Ok(self.read_file()?.member_names()?.len())
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
Ok(self.read_file()?.member_names(py)?.len())
}
/// The root group's name, `/`.
@@ -149,7 +289,7 @@ impl PyFile {
"/"
}
/// The path the file was opened with.
/// The path (or URL) the file was opened with.
#[getter]
fn filename(&self) -> &str {
&self.filename
@@ -204,9 +344,9 @@ impl PyFile {
/// Attribute access. In read mode, returns attributes of the root group.
/// In write mode, returns a writable attrs handle.
#[getter]
fn attrs(&self) -> PyResult<PyAttrs> {
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match self.inner.as_ref() {
Some(FileInner::Read(root)) => root.attrs(),
Some(FileInner::Read(root)) => root.attrs(py),
Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))),
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"file is closed",
@@ -216,9 +356,10 @@ impl PyFile {
fn __repr__(&self) -> String {
match &self.inner {
Some(FileInner::Read(root)) => {
format!("<HDF5 File (read, {} bytes)>", root.file.as_bytes().len())
}
Some(FileInner::Read(root)) => match root.handle.redacted_url() {
Some(url) => format!("<HDF5 File (read, \"{url}\")>"),
None => format!("<HDF5 File (read, \"{}\")>", self.filename),
},
Some(FileInner::Write(s)) => {
format!("<HDF5 File (write, \"{}\")>", s.path.display())
}
@@ -226,8 +367,8 @@ impl PyFile {
}
}
fn __contains__(&self, key: &str) -> PyResult<bool> {
Ok(self.read_file()?.contains(key))
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
self.read_file()?.contains(py, key)
}
}
@@ -271,11 +412,6 @@ fn parse_compression(
}
}
fn root_group(file: Arc<clawhdf5_rs::File>) -> ReadGroup {
let root = file.superblock().root_group_address;
ReadGroup::new(file, String::new(), root)
}
/// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| {
@@ -309,6 +445,18 @@ fn finalize_write(state: WriteState) -> PyResult<()> {
mod tests {
use super::*;
#[test]
fn urls_and_paths() {
assert!(is_url("http://h/f.h5"));
assert!(is_url("s3://bucket/key.h5"));
assert!(is_url("git+https://x"));
assert!(!is_url("data.h5"));
assert!(!is_url("/tmp/a://b.h5"));
assert!(!is_url("dir/x://y"));
assert!(!is_url("1http://x"));
assert!(!is_url("://x"));
}
#[test]
fn parse_gzip_compression() {
assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6));
+59 -76
View File
@@ -3,11 +3,12 @@
use std::collections::HashMap;
use std::sync::{Arc, Mutex, OnceLock};
use pyo3::exceptions::{PyIOError, PyKeyError, PyValueError};
use pyo3::exceptions::{PyIOError, PyKeyError, PyOSError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::PyList;
use crate::attrs::PyAttrs;
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
/// Shared state for a group being written.
@@ -34,9 +35,9 @@ enum GroupInner {
}
impl PyGroup {
pub(crate) fn from_read(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
pub(crate) fn from_read(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self {
inner: GroupInner::Read(ReadGroup::new(file, path, addr)),
inner: GroupInner::Read(ReadGroup::new(handle, path, addr)),
}
}
@@ -60,9 +61,10 @@ impl PyGroup {
/// h5py). It keeps its own address and, once listed, its links, so looking
/// up a child neither resolves the path from the root nor scans the group's
/// links again: visiting every member of a large group is linear, not
/// quadratic.
/// quadratic. (Edits never add or remove links, so these stay valid in a
/// file open for editing.)
pub(crate) struct ReadGroup {
pub file: Arc<clawhdf5_rs::File>,
pub handle: Arc<Handle>,
pub path: String,
pub addr: u64,
/// Link name -> object address (soft links resolved), filled on first use.
@@ -72,9 +74,9 @@ pub(crate) struct ReadGroup {
}
impl ReadGroup {
pub(crate) fn new(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
pub(crate) fn new(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self {
file,
handle,
path,
addr,
links: OnceLock::new(),
@@ -82,17 +84,14 @@ impl ReadGroup {
}
}
fn links(&self) -> PyResult<&HashMap<String, u64>> {
fn links(&self, py: Python<'_>) -> PyResult<&HashMap<String, u64>> {
if let Some(links) = self.links.get() {
return Ok(links);
}
let entries = crate::no_panic(|| {
clawhdf5_format::group_v2::resolve_group_children(
self.file.as_bytes(),
self.file.superblock(),
self.addr,
)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", node::name(&self.path))))
let (addr, path) = (self.addr, &self.path);
let entries = self.handle.with(py, |f| {
clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr)
.map_err(|e| node::format_err(path, e, PyValueError::new_err))
})?;
let map = entries
.into_iter()
@@ -102,7 +101,7 @@ impl ReadGroup {
}
/// The path and address of `key` (a name, a relative or an absolute path).
fn locate(&self, key: &str) -> PyResult<(String, u64)> {
fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> {
let path = node::join(&self.path, key);
let rel = if self.path.is_empty() {
Some(path.as_str())
@@ -112,24 +111,24 @@ impl ReadGroup {
path.strip_prefix(self.path.as_str())
.and_then(|r| r.strip_prefix('/'))
};
let addr = match rel {
// A direct child: the link table, when it has the name.
Some(name) if !name.is_empty() && !name.contains('/') => {
match self.links()?.get(name) {
Some(&a) => a,
None => node::resolve_from(&self.file, self.addr, name, &path)?,
if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/'))
&& let Some(&a) = self.links(py)?.get(name)
{
return Ok((path, a));
}
}
Some(rel) => node::resolve_from(&self.file, self.addr, rel, &path)?,
None => node::address(&self.file, &path)?,
};
Ok((path, addr))
let addr = self.addr;
let found = self.handle.with(py, |f| match rel {
Some(rel) => node::resolve_from(f, addr, rel, &path),
None => node::address(f, &path),
})?;
Ok((path, found))
}
/// `group[key]`.
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
let (path, addr) = self.locate(key)?;
node::open(py, &self.file, path, addr)
let (path, addr) = self.locate(py, key)?;
node::open(py, &self.handle, path, addr)
}
/// `group.get(key, default)`.
@@ -148,48 +147,57 @@ impl ReadGroup {
}
/// Names of the group's datasets and subgroups, sorted (h5py's order).
pub(crate) fn member_names(&self) -> PyResult<&[String]> {
pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> {
if let Some(m) = self.members.get() {
return Ok(m);
}
let links = self.links(py)?;
let path = &self.path;
let mut names = self.handle.with(py, |f| {
let mut names = Vec::new();
for (name, &addr) in self.links()? {
let hdr = node::header_at(&self.file, addr, &node::join(&self.path, name))?;
for (name, &addr) in links {
if matches!(
node::kind(&hdr),
node::kind_at(f, addr, &node::join(path, name))?,
Some(node::Kind::Dataset | node::Kind::Group)
) {
names.push(name.clone());
}
}
Ok(names)
})?;
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
Ok(self.members.get_or_init(|| names))
}
pub(crate) fn contains(&self, key: &str) -> bool {
self.locate(key)
.and_then(|(path, addr)| node::header_at(&self.file, addr, &path))
.ok()
.and_then(|h| node::kind(&h))
.is_some_and(|k| k != node::Kind::Datatype)
/// `key in group`: whether `key` names a dataset or group. A failed
/// read of the file (a network error) is raised, not `False`.
pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
let found = self
.locate(py, key)
.and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path)));
match found {
Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)),
Err(e) if e.is_instance_of::<PyOSError>(py) => Err(e),
Err(_) => Ok(false),
}
}
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
self.member_names()?
self.member_names(py)?
.iter()
.map(|n| self.get_item(py, n))
.collect()
}
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
self.member_names()?
self.member_names(py)?
.iter()
.map(|n| Ok((n.clone(), self.get_item(py, n)?)))
.collect()
}
pub(crate) fn attrs(&self) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
}
}
@@ -210,7 +218,7 @@ impl PyGroup {
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
match &self.inner {
GroupInner::Read(g) => {
let list = PyList::new(py, g.member_names()?)?;
let list = PyList::new(py, g.member_names(py)?)?;
Ok(list.into_any().unbind())
}
GroupInner::Write(state) => {
@@ -236,9 +244,9 @@ impl PyGroup {
self.keys(py)?.call_method0(py, "__iter__")
}
fn __len__(&self) -> PyResult<usize> {
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match &self.inner {
GroupInner::Read(g) => Ok(g.member_names()?.len()),
GroupInner::Read(g) => Ok(g.member_names(py)?.len()),
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
}
}
@@ -301,9 +309,9 @@ impl PyGroup {
/// Attribute access.
#[getter]
fn attrs(&self) -> PyResult<PyAttrs> {
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match &self.inner {
GroupInner::Read(g) => g.attrs(),
GroupInner::Read(g) => g.attrs(py),
GroupInner::Write(state) => {
let store = Arc::clone(&state.lock().unwrap().attrs);
Ok(PyAttrs::from_write(store))
@@ -311,10 +319,10 @@ impl PyGroup {
}
}
fn __repr__(&self) -> String {
fn __repr__(&self, py: Python<'_>) -> String {
match &self.inner {
GroupInner::Read(g) => {
let n = g.member_names().map_or(0, |m| m.len());
let n = g.member_names(py).map_or(0, |m| m.len());
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
}
GroupInner::Write(state) => {
@@ -324,9 +332,9 @@ impl PyGroup {
}
}
fn __contains__(&self, key: &str) -> PyResult<bool> {
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
match &self.inner {
GroupInner::Read(g) => Ok(g.contains(key)),
GroupInner::Read(g) => g.contains(py, key),
GroupInner::Write(state) => {
let guard = state.lock().unwrap();
Ok(guard.datasets.iter().any(|d| d.name == key))
@@ -357,31 +365,6 @@ pub(crate) fn finalize_write_group(
mod tests {
use super::*;
#[test]
fn member_names_are_sorted() {
let mut b = clawhdf5_rs::FileBuilder::new();
b.create_dataset("zeta").with_f64_data(&[1.0]);
b.create_dataset("alpha").with_f64_data(&[1.0]);
let mut g = b.create_group("mid");
g.create_dataset("x").with_f64_data(&[1.0]);
let finished = g.finish();
b.add_group(finished);
let bytes = b.finish().unwrap();
let file = Arc::new(clawhdf5_rs::File::from_bytes(bytes).unwrap());
let root = file.superblock().root_group_address;
let top = ReadGroup::new(Arc::clone(&file), String::new(), root);
assert_eq!(top.member_names().unwrap(), ["alpha", "mid", "zeta"]);
let (path, addr) = top.locate("mid").unwrap();
assert_eq!(path, "mid");
let mid = ReadGroup::new(Arc::clone(&file), path, addr);
assert_eq!(mid.member_names().unwrap(), ["x"]);
assert!(top.contains("mid/x"));
assert!(mid.contains("/alpha"));
assert!(mid.contains("x") && mid.contains("./x"));
assert!(!top.contains("nope"));
assert!(!mid.contains("alpha"));
}
#[test]
fn finalize_group() {
let state = WriteGroupState {
+130
View File
@@ -0,0 +1,130 @@
//! The open file every object of a `File` shares.
//!
//! Every read goes through [`Handle::with`], which releases the GIL and
//! parses through `File::storage()`, so the same code serves a local file
//! (memory-mapped) and a remote one (`clawhdf5-remote`: range requests
//! through a block cache, so a network read never holds the GIL).
//!
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
//! with the GIL released, and code that holds it never touches Python.
use std::path::PathBuf;
use std::sync::{Arc, PoisonError, RwLock};
use clawhdf5_rs::File;
use pyo3::exceptions::PyOSError;
use pyo3::prelude::*;
use crate::to_py_err;
/// Where the file's bytes come from.
pub(crate) enum Source {
/// A local path (memory-mapped).
Local(#[allow(dead_code)] PathBuf),
/// A URL, read through `clawhdf5-remote`'s block cache.
Remote {
url: String,
storage: Arc<clawhdf5_remote::RemoteStorage>,
},
}
pub(crate) struct Handle {
file: RwLock<Option<File>>,
source: Source,
pub offset_size: u8,
pub length_size: u8,
pub root: u64,
}
fn closed_after_failed_reopen() -> PyErr {
PyOSError::new_err("the file could not be reopened after an edit; open it again")
}
impl Handle {
fn new(file: File, source: Source) -> Arc<Self> {
let sb = file.superblock();
let (offset_size, length_size, root) =
(sb.offset_size, sb.length_size, sb.root_group_address);
Arc::new(Self {
file: RwLock::new(Some(file)),
source,
offset_size,
length_size,
root,
})
}
/// A local file, read-only.
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
Ok(Self::new(file, Source::Local(PathBuf::from(path))))
}
/// A remote file (`http(s)://`, `s3://`, ...).
pub(crate) fn open_url(
py: Python<'_>,
url: &str,
options: &clawhdf5_remote::Options,
) -> PyResult<Arc<Self>> {
let (file, storage) = py.detach(|| {
crate::no_panic(|| {
let storage = clawhdf5_remote::storage_for_url(url, options).map_err(remote_err)?;
let file = File::open_storage(storage.clone()).map_err(to_py_err)?;
Ok((file, storage))
})
})?;
Ok(Self::new(
file,
Source::Remote {
url: url.to_string(),
storage,
},
))
}
/// Run `f` on the file with the GIL released (a remote read may wait
/// on the network; other Python threads run meanwhile). `f` must not
/// touch Python.
pub(crate) fn with<R: Send>(
&self,
py: Python<'_>,
f: impl FnOnce(&File) -> PyResult<R> + Send,
) -> PyResult<R> {
py.detach(|| self.with_detached(f))
}
/// [`with`](Self::with) for code that already runs without the GIL.
pub(crate) fn with_detached<R>(&self, f: impl FnOnce(&File) -> PyResult<R>) -> PyResult<R> {
crate::no_panic(|| {
let guard = self.file.read().unwrap_or_else(PoisonError::into_inner);
let file = guard.as_ref().ok_or_else(closed_after_failed_reopen)?;
f(file)
})
}
/// The remote file's block cache.
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
match &self.source {
Source::Remote { storage, .. } => Some(storage),
Source::Local(_) => None,
}
}
/// The URL of a remote file, credentials and query values redacted.
pub(crate) fn redacted_url(&self) -> Option<String> {
match &self.source {
Source::Remote { url, .. } => Some(clawhdf5_remote::redact_url(url)),
Source::Local(_) => None,
}
}
}
/// A `clawhdf5_remote::Error` as a Python exception: the network side
/// (unreachable, a status, no range support, a changed file) is `OSError`,
/// a file that is not HDF5 is what `to_py_err` makes of it.
pub(crate) fn remote_err(e: clawhdf5_remote::Error) -> PyErr {
match e {
clawhdf5_remote::Error::Hdf5(e) => to_py_err(e),
other => PyOSError::new_err(other.to_string()),
}
}
+6 -1
View File
@@ -14,6 +14,7 @@ mod convert;
mod dataset;
mod file;
mod group;
mod handle;
mod node;
mod select;
@@ -63,7 +64,7 @@ fn _panic_for_test() -> PyResult<()> {
/// Convert a `clawhdf5_rs::Error` into a `PyErr`.
///
/// Maps different error variants to more specific Python exception types:
/// - I/O errors -> `PyIOError`
/// - I/O errors, and failed reads of a remote file -> `PyIOError`/`PyOSError`
/// - Format/parsing errors -> `PyValueError`
/// - Missing dataset/path errors -> `PyKeyError`
/// - Invalid arguments -> `PyValueError`
@@ -73,6 +74,10 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr {
use clawhdf5_rs::Error;
match &e {
Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()),
// A failed read of the storage: a network error on a remote file.
Error::Format(clawhdf5_format::error::FormatError::Storage(_)) => {
PyErr::new::<pyo3::exceptions::PyOSError, _>(e.to_string())
}
Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()),
Error::NotADataset(_) | Error::MissingMessage(_) => {
PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string())
+72 -51
View File
@@ -1,16 +1,24 @@
//! Resolving paths to objects in a file opened for reading.
//!
//! Everything here parses through `File::storage()` (the `clawhdf5_format`
//! `*_in` functions), never `File::as_bytes()`, so it works the same on a
//! memory-mapped local file and on a remote one; and it runs inside
//! `Handle::with`, without the GIL.
use std::sync::Arc;
use clawhdf5_format::attribute::AttributeMessage;
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
use clawhdf5_format::error::FormatError;
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
use clawhdf5_rs::File;
use pyo3::exceptions::{PyKeyError, PyOSError, PyTypeError, PyValueError};
use pyo3::prelude::*;
use crate::dataset::PyDataset;
use crate::dataset::{DatasetMeta, PyDataset};
use crate::group::PyGroup;
use crate::handle::Handle;
/// Join `key` onto the group path `base` the way h5py does: an absolute key
/// starts from the root, a relative one from `base`. Paths are kept without
@@ -33,42 +41,46 @@ pub(crate) fn name(path: &str) -> String {
format!("/{path}")
}
/// A format error met at `path`: a failed read of the storage (a network
/// error on a remote file) is an `OSError`, anything else `other(message)`.
pub(crate) fn format_err(path: &str, e: FormatError, other: fn(String) -> PyErr) -> PyErr {
let msg = format!("{}: {e}", name(path));
match e {
FormatError::Storage(_) => PyOSError::new_err(msg),
_ => other(msg),
}
}
fn value_err(msg: String) -> PyErr {
PyValueError::new_err(msg)
}
/// The address of the object at `path`, resolved from the root group.
pub(crate) fn address(file: &clawhdf5_rs::File, path: &str) -> PyResult<u64> {
pub(crate) fn address(file: &File, path: &str) -> PyResult<u64> {
resolve_from(file, file.superblock().root_group_address, path, path)
}
/// The address of `rel` resolved from the group at `group` (`full` is the
/// resulting path, for the error message).
pub(crate) fn resolve_from(
file: &clawhdf5_rs::File,
group: u64,
rel: &str,
full: &str,
) -> PyResult<u64> {
pub(crate) fn resolve_from(file: &File, group: u64, rel: &str, full: &str) -> PyResult<u64> {
if rel.is_empty() {
return Ok(group);
}
crate::no_panic(|| {
clawhdf5_format::group_v2::resolve_path_from(file.as_bytes(), file.superblock(), group, rel)
.map_err(|e| {
PyKeyError::new_err(format!(
clawhdf5_format::group_v2::resolve_path_from_in(file.storage(), file.superblock(), group, rel)
.map_err(|e| match e {
FormatError::Storage(_) => format_err(full, e, value_err),
e => PyKeyError::new_err(format!(
"Unable to open object (object '{}' doesn't exist): {e}",
name(full)
))
})
)),
})
}
/// The object header at `addr` (the object at `path`).
pub(crate) fn header_at(file: &clawhdf5_rs::File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
crate::no_panic(|| {
pub(crate) fn header_at(file: &File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
let sb = file.superblock();
let at = usize::try_from(addr)
.map_err(|_| PyValueError::new_err(format!("{}: address out of range", name(path))))?;
ObjectHeader::parse(file.as_bytes(), at, sb.offset_size, sb.length_size)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))
})
ObjectHeader::parse_in(file.storage(), addr, sb.offset_size, sb.length_size)
.map_err(|e| format_err(path, e, value_err))
}
/// What an object header describes.
@@ -96,29 +108,50 @@ pub(crate) fn kind(hdr: &ObjectHeader) -> Option<Kind> {
}
}
/// The kind of the object at `addr`, from its header.
pub(crate) fn kind_at(file: &File, addr: u64, path: &str) -> PyResult<Option<Kind>> {
Ok(kind(&header_at(file, addr, path)?))
}
/// What opening an object found, read without the GIL.
enum Found {
Dataset(DatasetMeta),
Group,
Datatype,
Other,
}
/// Open the object at `addr` (whose path is `path`) as a `Dataset` or
/// `Group`. Both keep the address, so later reads resolve nothing.
pub(crate) fn open(
py: Python<'_>,
file: &Arc<clawhdf5_rs::File>,
handle: &Arc<Handle>,
path: String,
addr: u64,
) -> PyResult<Py<PyAny>> {
let hdr = header_at(file, addr, &path)?;
match kind(&hdr) {
Some(Kind::Dataset) => Ok(PyDataset::open(py, Arc::clone(file), path, addr, &hdr)?
let found = handle.with(py, |f| {
let hdr = header_at(f, addr, &path)?;
Ok(match kind(&hdr) {
Some(Kind::Dataset) => Found::Dataset(DatasetMeta::load(f, addr, &hdr, &path)?),
Some(Kind::Group) => Found::Group,
Some(Kind::Datatype) => Found::Datatype,
None => Found::Other,
})
})?;
match found {
Found::Dataset(meta) => Ok(PyDataset::new(py, Arc::clone(handle), path, addr, meta)
.into_pyobject(py)?
.into_any()
.unbind()),
Some(Kind::Group) => Ok(PyGroup::from_read(Arc::clone(file), path, addr)
Found::Group => Ok(PyGroup::from_read(Arc::clone(handle), path, addr)
.into_pyobject(py)?
.into_any()
.unbind()),
Some(Kind::Datatype) => Err(PyTypeError::new_err(format!(
Found::Datatype => Err(PyTypeError::new_err(format!(
"{}: committed (named) datatypes are not supported by clawhdf5",
name(&path)
))),
None => Err(PyValueError::new_err(format!(
Found::Other => Err(PyValueError::new_err(format!(
"{}: not a dataset, group or datatype",
name(&path)
))),
@@ -126,32 +159,26 @@ pub(crate) fn open(
}
/// The dataspace message of an object header.
pub(crate) fn dataspace(file: &clawhdf5_rs::File, hdr: &ObjectHeader) -> PyResult<Dataspace> {
crate::no_panic(|| {
pub(crate) fn dataspace(file: &File, hdr: &ObjectHeader, path: &str) -> PyResult<Dataspace> {
let sb = file.superblock();
let msg = hdr
.messages
.iter()
.find(|m| m.msg_type == MessageType::Dataspace)
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
let data = clawhdf5_format::shared_message::message_data(
file.as_bytes(),
let data = clawhdf5_format::shared_message::message_data_in(
file.storage(),
msg,
sb.offset_size,
sb.length_size,
)
.map_err(|e| PyValueError::new_err(e.to_string()))?;
Dataspace::parse(&data, sb.length_size).map_err(|e| PyValueError::new_err(e.to_string()))
})
.map_err(|e| format_err(path, e, value_err))?;
Dataspace::parse(&data, sb.length_size).map_err(|e| format_err(path, e, value_err))
}
/// The chunk shape of a chunked dataset (one entry per dataset dimension),
/// or `None` for other layouts or a layout message that does not parse.
pub(crate) fn chunk_shape(
file: &clawhdf5_rs::File,
hdr: &ObjectHeader,
rank: usize,
) -> Option<Vec<u64>> {
pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Option<Vec<u64>> {
let sb = file.superblock();
let msg = hdr
.messages
@@ -179,24 +206,18 @@ pub(crate) fn is_null(space: &Dataspace) -> bool {
/// The attributes of the object at `addr` (whose path is `path`), sorted by
/// name (h5py's order). Attributes whose messages cannot be parsed are left
/// out, as the facade's `attrs()` does.
pub(crate) fn attributes(
file: &clawhdf5_rs::File,
addr: u64,
path: &str,
) -> PyResult<Vec<AttributeMessage>> {
pub(crate) fn attributes(file: &File, addr: u64, path: &str) -> PyResult<Vec<AttributeMessage>> {
let hdr = header_at(file, addr, path)?;
crate::no_panic(|| {
let sb = file.superblock();
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant(
file.as_bytes(),
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
file.storage(),
&hdr,
sb.offset_size,
sb.length_size,
)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))?;
.map_err(|e| format_err(path, e, value_err))?;
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
Ok(attrs)
})
}
#[cfg(test)]
+135
View File
@@ -1,6 +1,10 @@
"""Shared fixtures for the clawhdf5 Python binding tests."""
import os
import re
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import pytest
@@ -16,3 +20,134 @@ def h5py():
pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable")
pytest.skip("h5py not installed")
return mod
# ---------------------------------------------------------------------------
# An HTTP server for remote reads
# ---------------------------------------------------------------------------
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
class RangeServer:
"""A static file server on 127.0.0.1, in a thread of this process, that
answers `Range: bytes=a-b` with 206 and `Content-Range` (the way S3 and
common web servers do), sends an ETag and honours `If-Match`.
- `ranges=False`: ignores `Range` and answers 200 with the whole file,
like a server without range support.
- `down` (set by `close()`): hang up on every request.
- `delay`: seconds to wait before answering each request after the
first `delay_after` ones (a slow network).
- `log`: every request as `(method, path, range header)`.
"""
def __init__(self, root, ranges=True):
self.root = str(root)
self.ranges = ranges
self.delay = 0.0
self.delay_after = 0
self.down = False
self.log = []
self._lock = threading.Lock()
server = self
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def log_message(self, *args): # quiet
pass
def do_HEAD(self):
self._serve(body=False)
def do_GET(self):
self._serve(body=True)
def _serve(self, body):
with server._lock:
server.log.append((self.command, self.path, self.headers.get("Range")))
n = len(server.log)
if server.delay and n > server.delay_after:
time.sleep(server.delay)
if server.down:
# Hang up without an answer (open keep-alive
# connections outlive shutdown(), so close() sets this).
self.close_connection = True
return
path = os.path.join(server.root, self.path.lstrip("/").split("?")[0])
if not os.path.isfile(path):
self.send_response(404)
self.send_header("Content-Length", "0")
self.end_headers()
return
with open(path, "rb") as fh:
data = fh.read()
st = os.stat(path)
etag = f'"{st.st_mtime_ns:x}-{st.st_size:x}"'
want = self.headers.get("If-Match")
if want is not None and want != etag and want != "*":
self.send_response(412)
self.send_header("Content-Length", "0")
self.end_headers()
return
rng = self.headers.get("Range") if server.ranges else None
m = _RANGE.match(rng.strip()) if rng else None
if m and (m.group(1) or m.group(2)):
size = len(data)
if m.group(1):
start = int(m.group(1))
end = int(m.group(2)) if m.group(2) else size - 1
else:
start = max(0, size - int(m.group(2)))
end = size - 1
if start >= size:
self.send_response(416)
self.send_header("Content-Range", f"bytes */{size}")
self.send_header("Content-Length", "0")
self.end_headers()
return
end = min(end, size - 1)
part = data[start : end + 1]
self.send_response(206)
self.send_header("Content-Range", f"bytes {start}-{end}/{size}")
else:
part = data
self.send_response(200)
if server.ranges:
self.send_header("Accept-Ranges", "bytes")
self.send_header("ETag", etag)
self.send_header("Content-Length", str(len(part)))
self.send_header("Content-Type", "application/x-hdf5")
self.end_headers()
if body:
try:
self.wfile.write(part)
except (BrokenPipeError, ConnectionResetError):
pass
self.httpd = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
self.httpd.daemon_threads = True
self.port = self.httpd.server_address[1]
self.thread = threading.Thread(target=self.httpd.serve_forever, daemon=True)
self.thread.start()
def url(self, name):
return f"http://127.0.0.1:{self.port}/{name}"
def requests(self):
with self._lock:
return len(self.log)
def close(self):
self.down = True
self.httpd.shutdown()
self.httpd.server_close()
@pytest.fixture
def range_server(tmp_path):
"""A range-capable server over `tmp_path`."""
server = RangeServer(tmp_path)
yield server
server.close()
+20 -3
View File
@@ -178,15 +178,32 @@ def _write_fixture(h5py, path):
g.attrs["depth"] = np.int8(3)
@pytest.fixture(scope="module")
def pair(h5py, tmp_path_factory):
path = str(tmp_path_factory.mktemp("h5") / "fixture.h5")
@pytest.fixture(scope="module", params=["local", "http", "http-1k-blocks"])
def pair(request, h5py, tmp_path_factory):
"""The fixture file through h5py and through clawhdf5: opened locally,
and over HTTP range requests (a local server in this process) with the
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
through many small ranges."""
from conftest import RangeServer
root = tmp_path_factory.mktemp("h5")
path = str(root / "fixture.h5")
_write_fixture(h5py, path)
theirs = h5py.File(path, "r")
server = None
if request.param == "local":
ours = clawhdf5.File(path, "r")
else:
server = RangeServer(root)
if request.param == "http":
ours = clawhdf5.File(server.url("fixture.h5"))
else:
ours = clawhdf5.File.open_url(server.url("fixture.h5"), block_size=1024)
yield ours, theirs, path
theirs.close()
ours.close()
if server is not None:
server.close()
def _all_datasets(h5py, f):
+244
View File
@@ -0,0 +1,244 @@
"""Remote files: `clawhdf5.File(url)` / `File.open_url(url, ...)` read over
HTTP range requests (clawhdf5-remote's block cache), against a server in this
process (conftest.RangeServer). Values are compared with h5py reading the
same file locally; the rest checks what the server saw (only the blocks a
read needs are fetched), the failure modes (no range support, a missing
file, a file that changes, a server that goes away: errors, never wrong
data), and that the GIL is released while a read waits on the network."""
import os
import sys
import threading
import time
import numpy as np
import pytest
import clawhdf5
from conftest import RangeServer
def _write(h5py, path):
rng = np.random.default_rng(7)
with h5py.File(path, "w") as f:
f.create_dataset("contig", data=rng.standard_normal((400, 300)))
f.create_dataset(
"chunked",
data=rng.integers(0, 1000, size=(512, 512), dtype="<i4"),
chunks=(64, 64),
compression="gzip",
)
f.create_dataset("strings", data=["alpha", "beta", "gamma"], dtype=h5py.string_dtype())
g = f.create_group("grp")
g.create_dataset("small", data=np.arange(10, dtype="<u2"))
g.attrs["units"] = "m/s"
f.attrs["title"] = "remote test"
@pytest.fixture
def remote_file(h5py, tmp_path, range_server):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
return path, range_server.url("remote.h5"), range_server
def test_remote_reads_match_h5py(h5py, remote_file):
path, url, server = remote_file
with h5py.File(path, "r") as theirs, clawhdf5.File(url) as ours:
assert ours.filename == url
assert list(ours.keys()) == list(theirs.keys())
assert "grp/small" in ours and "nope" not in ours
for name in ["contig", "chunked", "grp/small"]:
np.testing.assert_array_equal(ours[name][...], theirs[name][...])
np.testing.assert_array_equal(ours[name][3:7], theirs[name][3:7])
np.testing.assert_array_equal(ours["chunked"][100:130, 200:300:3], theirs["chunked"][100:130, 200:300:3])
np.testing.assert_array_equal(ours["chunked"][[1, 70, 300], 5], theirs["chunked"][[1, 70, 300], 5])
assert list(ours["strings"][...]) == list(theirs["strings"][...])
assert ours["grp"].attrs["units"] == theirs["grp"].attrs["units"]
assert ours.attrs["title"] == theirs.attrs["title"]
assert ours["chunked"].maxshape == theirs["chunked"].maxshape
assert server.requests() >= 2
def test_a_small_read_fetches_only_its_blocks(h5py, remote_file):
"""With 4 KiB blocks, opening and reading one chunk of a 1 MB chunked
dataset costs a handful of requests and a few blocks, not the file."""
path, url, server = remote_file
size = os.path.getsize(path)
f = clawhdf5.File.open_url(url, block_size=4096)
opened = server.requests()
assert opened == 1, server.log
ds = f["chunked"]
got = ds[0:10, 0:10]
with h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(got, theirs["chunked"][0:10, 0:10])
stats = f.remote_stats
assert stats["bytes_fetched"] < size / 4, (stats, size)
assert server.requests() - opened <= 12, server.log
# A second read of the same region is served by the cache.
before = server.requests()
ds[0:10, 0:10]
assert server.requests() == before
assert f.remote_stats["hits"] > stats["hits"]
assert clawhdf5.File(str(path), "r").remote_stats is None
def test_server_without_range_support(h5py, tmp_path):
"""A server that ignores Range answers 200 with the whole file: that is
an OSError by default, and a whole download when allowed."""
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path, ranges=False)
try:
url = server.url("remote.h5")
with pytest.raises(OSError, match="range"):
clawhdf5.File(url)
with clawhdf5.File.open_url(url, allow_full_download=True) as ours, h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(ours["chunked"][...], theirs["chunked"][...])
np.testing.assert_array_equal(ours["contig"][5], theirs["contig"][5])
with pytest.raises(OSError):
clawhdf5.File.open_url(url, allow_full_download=True, max_full_download=1000)
finally:
server.close()
def test_errors_are_oserrors(remote_file):
_, url, server = remote_file
with pytest.raises(OSError, match="404"):
clawhdf5.File(server.url("missing.h5"))
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "r+")
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "w")
with pytest.raises(OSError, match="unsupported URL"):
clawhdf5.File("nosuchscheme://x/y.h5")
with pytest.raises(ValueError):
clawhdf5.File.open_url(url, block_size=0)
with pytest.raises(TypeError):
clawhdf5.File.open_url(url, no_such_option=1)
def test_object_store_urls_need_their_features():
"""The default wheel has no S3/GCS/Azure clients (aws-lc-rs builds C):
such a URL is an OSError naming the build feature."""
for url, feature in [("s3://bucket/k.h5", "s3"), ("gs://b/k.h5", "gcs"), ("az://c/k.h5", "azure")]:
try:
clawhdf5.File(url)
except OSError as e:
if "feature" in str(e):
assert f"`{feature}`" in str(e), str(e)
else:
pytest.fail(f"{url} opened")
def test_https_needs_the_https_feature():
"""The default wheel has no TLS stack (rustls needs ring, which builds C):
an https URL is an OSError that names the build feature."""
with pytest.raises(OSError) as e:
clawhdf5.File("https://127.0.0.1:1/x.h5")
msg = str(e.value)
# Built with `--features https` the error is the refused connection.
assert "https" in msg or "connect" in msg.lower() or "refused" in msg.lower(), msg
def test_a_changed_file_is_an_error_not_mixed_data(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=1024)
first = f["grp/small"][...]
# Rewrite the file with other values: new ETag, same name.
time.sleep(0.01)
with h5py.File(path, "w") as g:
g.create_dataset("contig", data=np.zeros((400, 300)))
with pytest.raises(OSError, match="changed"):
f["contig"][...]
np.testing.assert_array_equal(first, np.arange(10, dtype="<u2"))
def test_a_server_that_goes_away_is_an_error(h5py, tmp_path):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path)
f = clawhdf5.File.open_url(server.url("remote.h5"), block_size=1024, retries=0, timeout=2)
ds = f["contig"]
server.close()
with pytest.raises(OSError):
ds[...]
def test_threads_read_one_remote_file(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=2048)
with h5py.File(path, "r") as theirs:
expected = theirs["chunked"][...]
errors = []
def work(i):
try:
rows = slice((i * 37) % 400, (i * 37) % 400 + 64)
np.testing.assert_array_equal(f["chunked"][rows], expected[rows])
except Exception as e: # noqa: BLE001
errors.append(e)
threads = [threading.Thread(target=work, args=(i,)) for i in range(16)]
for t in threads:
t.start()
for t in threads:
t.join()
assert not errors, errors[:3]
def test_remote_reads_release_the_gil(h5py, remote_file):
"""A read waiting on a slow server lets other Python threads run: a
thread counting in a loop keeps counting (and never stalls for long)
while the main thread reads through requests that each take 0.2 s."""
_, url, server = remote_file
f = clawhdf5.File.open_url(url, block_size=1024, max_parallel=1)
ds = f["contig"]
server.delay = 0.2
server.delay_after = server.requests()
old = sys.getswitchinterval()
sys.setswitchinterval(0.001)
stop = threading.Event()
progress = {"n": 0, "worst": 0.0}
def spin():
last = time.perf_counter()
while not stop.is_set():
now = time.perf_counter()
progress["worst"] = max(progress["worst"], now - last)
last = now
progress["n"] += 1
t = threading.Thread(target=spin)
try:
t.start()
time.sleep(0.02)
t0 = time.perf_counter()
before = server.requests()
ds[0:2]
took = time.perf_counter() - t0
stop.set()
t.join()
finally:
sys.setswitchinterval(old)
assert server.requests() > before
assert took >= 0.2, took
assert progress["n"] > 1000, progress
# Held across a 0.2 s request, the spinner would stall that long.
assert progress["worst"] < 0.1, (progress, took)
def test_a_clawhdf5_written_file_reads_the_same_remotely(tmp_path, range_server):
path = tmp_path / "ours.h5"
data = np.arange(3000, dtype="<f8").reshape(100, 30)
with clawhdf5.File(str(path), "w") as f:
f.create_dataset("d", data=data, chunks=[10, 30], compression="gzip")
g = f.create_group("g")
g.create_dataset("i", data=np.arange(5, dtype="<i4"))
f.attrs["k"] = 3
with clawhdf5.File(range_server.url("ours.h5")) as f, clawhdf5.File(str(path)) as local:
np.testing.assert_array_equal(f["d"][...], data)
np.testing.assert_array_equal(f["d"][5:9, ::4], local["d"][5:9, ::4])
np.testing.assert_array_equal(f["g/i"][...], np.arange(5))
assert f.attrs["k"] == local.attrs["k"]
assert "file (read" in repr(f).lower() and "127.0.0.1" in repr(f)
+15 -3
View File
@@ -7,8 +7,9 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
`File::open_storage` gives the facade's read API over any `Storage` (see
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is
next. Every count in §1–§2 was
object stores) and URLs in `h5rs` (see the M3 status below); the Python
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
@@ -475,7 +476,7 @@ fast path within benchmark noise.
a request counter exposed for tests and users.
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
Python bindings, with these choices:
Python bindings (done 2026-09-27, below), with these choices:
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
sits below the facade.
@@ -512,6 +513,17 @@ fast path within benchmark noise.
same work without a cache: 141 936 requests. Per file: A lists in 2
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
6.4 MB: 35 001 object headers spread over the file).
- *Status 2026-09-27, Python bindings:* done on branch
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
`File.open_url(url, **options)` (cache and HTTP options) go through
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
parsing (path lookups, headers, attributes, listings, the global heap)
moved from `File::as_bytes` to `File::storage()` and the `*_in`
functions, and every file access, metadata included, runs with the GIL
released. Checked by running the whole read-vs-h5py suite over an
in-process range server (1 MiB and 1 KiB blocks) and by request counts
in `crates/clawhdf5-py/tests/test_remote.py`.
**M4 — wasm lazy loading (1–2 weeks).**
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
+9 -6
View File
@@ -844,9 +844,9 @@ cache, but:
`read_*_zerocopy`) need the file in memory and answer
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
panics for such a file (`File::contiguous_bytes()` is the fallible form).
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
`remote` feature).
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
(`h5rs` and the Python bindings read through `File::storage`, and take
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -862,9 +862,12 @@ cache, but:
**Status:** open (added 2026-09-26, milestone M3 of
`docs/design/range-reads.md`).
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
parses through `File::as_bytes`, which a remote file does not have; the
wasm reader's `openUrl` is milestone M4.
- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
the default wheel reads plain `http://` only: `https://` needs a wheel
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.
+5 -2
View File
@@ -131,14 +131,17 @@ run_step "cargo clippy (h5rs remote)" cargo clippy \
# js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C.
# clawhdf5-remote is checked by default (plain HTTP) and with its
# object-store feature, and h5rs with URL support (remote); the https
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in.
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. The
# Python bindings (clawhdf5-py, remote reads over plain HTTP) are checked too:
# their https/s3/gcs/azure features are opt-in for the same reason.
no_c_in_default_build() {
local entry crate features found=0
for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \
clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \
clawhdf5-tools \
clawhdf5-wasm \
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote; do
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote \
clawhdf5-py; do
crate=${entry%%:*}
features=()
[ "$entry" != "$crate" ] && features=(--features "${entry#*:}")