Files
clawhdf5/crates/clawhdf5-remote/README.md
T
osobhandClaude Opus 5.5 b55b24b7ba docs: crate READMEs describe each crate as it is today
Every crate under crates/ now has a README (android, bench, cli, napi and
wasm had none), each saying what the crate is, its main types and
functions (names checked against the code), its cargo features with
defaults and which ones build C (checked with `cargo tree`), and links to
the top-level docs.

Corrections to the old stubs:
- clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs
  clawhdf5-format as a dependency.
- clawhdf5-filters: deflate backends only, and no library crate depends
  on it; the filter pipeline and every other codec are in -format.
- clawhdf5-gpu: vector distance compute, not I/O; not used by
  HDF5Memory::search.
- clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO.
- clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and
  search(q, k, ef).
- clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm
  backends are reported but run the scalar kernels.
- clawhdf5-gpu: the old example called l2_distances, which does not
  exist (l2_search).
- clawhdf5-agent: it described a "vector store" with "GPU acceleration";
  it now covers HDF5Memory, search options, WAL, signing, the graph.
- crates.io/docs.rs badges removed and `cargo install <crate>` replaced:
  nothing is published; depend on git.
- fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh.
- tools: the FileEditor interop tests that live in this crate.
- remote, py: license, other front ends, limits, File.mode/flush/chunks.

The Rust examples of the facade, format, filters, accel, ann, derive and
agent READMEs were compiled and run as tests (netcdf4, gpu and remote
compiled only) in a scratch crate; the CLI example was run.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:30 -05:00

7.1 KiB

clawhdf5-remote

Read HDF5 files where they live — on an HTTP(S) server or in an object store (S3, GCS, Azure) — with clawhdf5, without downloading them first.

let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/data.h5")?;
let temps = file.dataset("/grid/temperature")?.read_f64()?;

The result is an ordinary clawhdf5::File: groups, datasets, attributes, selections, variable-length data. Only the bytes an operation needs are fetched, by Range requests, through a block cache.

How it reads

  • Block cache (BlockCache, mandatory for remote files; see docs/design/range-reads.md §2): aligned blocks of 1 MiB by default, LRU with a byte budget (64 MiB by default), the missing blocks of one read fetched as runs of consecutive blocks (a gap of one block is fetched to merge two runs), each request at most 8 MiB, the requests of one read in parallel. A read that misses more than half the budget is not kept, so a big dataset read does not evict the metadata. Readers on several threads share one cache; a block is never fetched twice at once — a second reader waits for the first one's request.
  • Opening costs one request: a GET of the first block, whose Content-Range gives the file's length. HDF5 files keep the superblock and usually the root group's metadata there, so listing a small file often needs nothing more.
  • Pinned to one version of the file: a strong ETag is sent back as If-Match (else Last-Modified as If-Unmodified-Since) and checked on every response, as is the length. A file replaced while open is an error (RemoteError::FileChanged), never a mix of old and new bytes. A server with neither validator can only be checked by length; HttpOptions::require_validator refuses it.
  • Servers that ignore Range (answer 200 with the whole file) are refused (RemoteError::RangeNotSupported) without reading the body, unless HttpOptions::allow_full_download is set; then the file is downloaded once and read from memory. A 200 whose body is no longer than the range asked for is the whole (small) file, and is accepted.
  • Retries: connection failures, timeouts, 408/429/5xx and bodies that end early are retried with exponential backoff (3 retries, from 200 ms). Bodies are requested with Accept-Encoding: identity; an encoded body is refused.
  • Timeouts scale with the request: HttpOptions::timeout (30 s) to connect and to receive the headers, and for the body that plus its size at HttpOptions::min_speed (16 KiB/s) — a slow link is not cut off mid-block, a stalled connection still fails.
  • Redirects are followed up to HttpOptions::max_redirects (5; 0 refuses them), never from https to http. Once a redirect leaves the URL's origin (scheme, host, port), HttpOptions::headers (API keys, Authorization, cookies) are no longer sent.
  • Credentials stay out of messages: every error and Debug output shows URLs through redact_url — no user:password@, query values replaced by REDACTED (a presigned S3/GCS URL's signature lives there).
  • Claimed lengths are not trusted: nothing is allocated for the length a server reports; a read spanning more than the cache budget is fetched in pieces as data arrives, and download(&storage, max_bytes) reads a whole file only up to a limit (DEFAULT_MAX_DOWNLOAD, 1 GiB).

The zero-copy methods of clawhdf5 (read_raw_ref, read_*_zerocopy, File::as_bytes) borrow the whole file from memory, so they are errors (as_bytes a panic; use File::contiguous_bytes) on a remote file.

Features

Feature What C code
http (default) http:// through ureq, no TLS none
https https:// through rustls, ring provider, Mozilla roots ring (C and assembly)
object-store ObjectStoreStorage and open_object over any object_store store (in-memory, local files, or one you configure) none
s3, gcs, azure s3://bucket/key, gs://bucket/key, az://container/key in open_url, configured from the environment (AWS_*, GOOGLE_*, AZURE_*) as object_store's from_env builders read it aws-lc-rs (object_store's cloud clients)

The default build and object-store compile no C (scripts/ci-test.sh checks both).

Object stores

object_store is async; Storage is synchronous (parsing is CPU work). ObjectStoreStorage owns a small tokio runtime (two worker threads): each read runs there while the calling thread waits, so it works from any thread, several at once — including tokio::task::spawn_blocking and code inside another runtime (where spawn_blocking is still the better place, since a read blocks the thread it is called on). The object is pinned by its ETag (If-Match, and compared on every response), else its version or modification time, and its size. The ranges of one read are fetched concurrently (up to 8).

use std::sync::Arc;
use clawhdf5_remote::object_store::{memory::InMemory, ObjectStore};
let store: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
// ... put a file at "data.h5" ...
let (file, cache) = clawhdf5_remote::open_object(store, "data.h5", &Default::default())?;

The tests use object_store's in-memory and local-file stores; no cloud account is needed. The cloud schemes are only built (and unit-tested for URL parsing) in CI, not run against a real bucket.

Counting requests

use clawhdf5_remote::{storage_for_url, Options};
let storage = storage_for_url(url, &Options::default())?;
let file = clawhdf5::File::open_storage(storage.clone())?;
// ... read ...
let s = storage.stats(); // requests, bytes_fetched, hits, misses, cached_bytes, ...

cargo run -p clawhdf5-remote --example range_server -- DIR serves a directory with range support (the server the tests use), and cargo run -p clawhdf5-remote --example read_url -- URL [DATASET] lists a file and prints what it cost.

Other front ends

  • h5rs (built with --features remote, or remote-https) takes URLs as FILE arguments: clawhdf5-tools.
  • Python: clawhdf5.File("http://…") and File.open_url(url, ...) go through this crate: clawhdf5-py.
  • The browser does not use this crate (its cache fetches by blocking); clawhdf5-wasm's openUrl has its own restartable cache: examples/wasm-viewer.

Limits

Files a SWMR writer is still appending to cannot be followed remotely (the file is pinned at open, so growth is RemoteError::FileChanged); the block size is fixed rather than taken from a paged file's page size; the cloud backends are built and unit-tested but have not been run against a real bucket. The full list is under "Remote files (clawhdf5-remote) limits" in docs/known-issues.md; the design is milestone M3 of docs/design/range-reads.md.

Not on crates.io yet; depend on it from git:

[dependencies]
clawhdf5-remote = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }

License

MIT