Files
clawhdf5/crates/clawhdf5-remote/README.md
T
osobhandClaude Opus 5.5 4ff3e40fea clawhdf5-remote: object stores through object_store (S3, GCS, Azure)
ObjectStoreStorage (feature `object-store`, pure Rust) reads one object
of any object_store store by ranged get_opts, pinned at open by a head
request: If-Match with its ETag (and the ETag and size of every response
compared), else its version or modification time. A change is
RemoteError::FileChanged. object_store is async and Storage is not, so
the storage owns a small multi-threaded tokio runtime (two workers) and
blocks the calling thread on it; the ranges of one read_ranges call are
fetched concurrently (up to 8). From inside another tokio runtime it
refuses with RemoteError::Usage instead of blocking a worker, and it
shuts its runtime down in the background on drop so dropping it in async
code does not panic.

open_object(store, path, options) opens a file through a block cache
(first block prefetched); open_url accepts s3://, gs:// and az:// with
the `s3`, `gcs` and `azure` features, configured from the environment by
object_store's from_env builders. Those pull object_store's cloud clients
and aws-lc-rs (C), so they are opt-in; without them the URL is a clean
UnsupportedScheme error naming the feature.

Tests against object_store's in-memory and local-file stores (no cloud):
every fixture's transcript equals File::open's, a multi-block object is
fetched in coalesced block runs, an object replaced while open is an
error, and a missing object or a read from inside a runtime is a clean
error. ci-test.sh lints all backends, runs these tests (with s3 for its
URL parsing test) and checks object-store for C in the no-C step.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 17:15:06 -05:00

4.8 KiB

clawhdf5-remote

Read HDF5 files where they live — on an HTTP(S) server or in an object store (S3, GCS, Azure) — with clawhdf5, without downloading them first.

let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/data.h5")?;
let temps = file.dataset("/grid/temperature")?.read_f64()?;

The result is an ordinary clawhdf5::File: groups, datasets, attributes, selections, variable-length data. Only the bytes an operation needs are fetched, by Range requests, through a block cache.

How it reads

  • Block cache (BlockCache, mandatory for remote files; see docs/design/range-reads.md §2): aligned blocks of 1 MiB by default, LRU with a byte budget (64 MiB by default), the missing blocks of one read fetched as runs of consecutive blocks (a gap of one block is fetched to merge two runs), each request at most 8 MiB, the requests of one read in parallel. A read that misses more than half the budget is not kept, so a big dataset read does not evict the metadata. Readers on several threads share one cache; a block is never fetched twice at once — a second reader waits for the first one's request.
  • Opening costs one request: a GET of the first block, whose Content-Range gives the file's length. HDF5 files keep the superblock and usually the root group's metadata there, so listing a small file often needs nothing more.
  • Pinned to one version of the file: a strong ETag is sent back as If-Match (else Last-Modified as If-Unmodified-Since) and checked on every response, as is the length. A file replaced while open is an error (RemoteError::FileChanged), never a mix of old and new bytes. A server with neither validator can only be checked by length; HttpOptions::require_validator refuses it.
  • Servers that ignore Range (answer 200 with the whole file) are refused (RemoteError::RangeNotSupported) without reading the body, unless HttpOptions::allow_full_download is set; then the file is downloaded once and read from memory.
  • Retries: connection failures, timeouts, 408/429/5xx and bodies that end early are retried with exponential backoff (3 retries, from 200 ms). Bodies are requested with Accept-Encoding: identity; an encoded body is refused.

The zero-copy methods of clawhdf5 (read_raw_ref, read_*_zerocopy, File::as_bytes) borrow the whole file from memory, so they are errors (as_bytes a panic; use File::contiguous_bytes) on a remote file.

Features

Feature What C code
http (default) http:// through ureq, no TLS none
https https:// through rustls, ring provider, Mozilla roots ring (C and assembly)
object-store ObjectStoreStorage and open_object over any object_store store (in-memory, local files, or one you configure) none
s3, gcs, azure s3://bucket/key, gs://bucket/key, az://container/key in open_url, configured from the environment (AWS_*, GOOGLE_*, AZURE_*) as object_store's from_env builders read it aws-lc-rs (object_store's cloud clients)

The default build and object-store compile no C (scripts/ci-test.sh checks both).

Object stores

object_store is async; Storage is synchronous (parsing is CPU work). ObjectStoreStorage owns a small tokio runtime (two worker threads) and blocks the calling thread on it for each read, so it is read from ordinary threads, several at once. From inside an async runtime it refuses (RemoteError::Usage) rather than block a worker: read in tokio::task::spawn_blocking. The object is pinned by its ETag (If-Match, and compared on every response), else its version or modification time, and its size. The ranges of one read are fetched concurrently (up to 8).

use std::sync::Arc;
use clawhdf5_remote::object_store::{memory::InMemory, ObjectStore};
let store: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
// ... put a file at "data.h5" ...
let (file, cache) = clawhdf5_remote::open_object(store, "data.h5", &Default::default())?;

The tests use object_store's in-memory and local-file stores; no cloud account is needed. The cloud schemes are only built (and unit-tested for URL parsing) in CI, not run against a real bucket.

Counting requests

use clawhdf5_remote::{storage_for_url, Options};
let storage = storage_for_url(url, &Options::default())?;
let file = clawhdf5::File::open_storage(storage.clone())?;
// ... read ...
let s = storage.stats(); // requests, bytes_fetched, hits, misses, cached_bytes, ...

cargo run -p clawhdf5-remote --example range_server -- DIR serves a directory with range support (the server the tests use), and cargo run -p clawhdf5-remote --example read_url -- URL [DATASET] lists a file and prints what it cost.