clawhdf5-remote: block cache and HTTP range reads (open_url)

Range-read milestone M3, first half: a new crate with the block cache the
design makes mandatory for remote files and an HTTP backend, so
open_url("http://...") gives a clawhdf5::File over File::open_storage.

BlockCache wraps any Storage: aligned blocks (1 MiB by default, the size
docs/design/range-reads.md section 2 measured), LRU with a byte budget,
the missing blocks of one read_at/read_ranges fetched with one backend
read_ranges call as runs of consecutive blocks (a one-block gap filled to
merge runs, each request at most 8 MiB), and reads that miss more than
half the budget not kept. Thread-safe without holding the lock across a
fetch: a block being fetched is in flight, a second reader waits for it
instead of fetching it again, and a failed fetch fails its waiters and is
not cached. A backend holding the file in memory passes through.

HttpStorage (ureq, no TLS by default; `https` adds rustls with ring):
opening is one ranged GET of the first block, whose Content-Range gives
the length (the cache keeps the bytes). The file is pinned by a strong
ETag (If-Match), else Last-Modified (If-Unmodified-Since), and its length,
checked on every response: a change is RemoteError::FileChanged, never
mixed data. A server that ignores Range is refused without reading the
body unless a full download is allowed. Connection errors, timeouts,
408/429/5xx and short bodies are retried with exponential backoff;
Accept-Encoding: identity, and an encoded body is refused. read_ranges
fetches its ranges in parallel.

Tests (a std-only HTTP/1.1 server in tests/common/server.rs, also the
range_server example): every fixture read over HTTP gives File::open's
transcript (CLAWHDF5_REMOTE_CORPUS adds the conformance corpus), with
request counts per file with and without the cache; an h5py-written file
against libhdf5's values; a multi-block file fetched in whole blocks, each
once; a server ignoring Range; a file replaced mid-read (ETag,
Last-Modified, length only); truncated bodies and 503s (retried, then an
error, never cached); a slow server with 8 concurrent readers (no block
fetched twice); bad URLs, 404, encoded bodies, non-HDF5 data. The cache
has unit tests for coalescing, splitting, LRU order, large reads,
failures and concurrent in-flight dedup.

ci-test.sh: clawhdf5-remote joins the no-C default-build check, and its
https feature is linted.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-26 17:13:28 -05:00
co-authored by Claude Opus 5.5
parent f191dc09d5
commit db2554dd81
14 changed files with 3024 additions and 6 deletions
+71
View File
@@ -0,0 +1,71 @@
# clawhdf5-remote
Read HDF5 files where they live — on an HTTP(S) server — with
[clawhdf5](../../README.md), without downloading them first.
```rust
let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/data.h5")?;
let temps = file.dataset("/grid/temperature")?.read_f64()?;
```
The result is an ordinary `clawhdf5::File`: groups, datasets, attributes,
selections, variable-length data. Only the bytes an operation needs are
fetched, by `Range` requests, through a block cache.
## How it reads
- **Block cache** (`BlockCache`, mandatory for remote files; see
`docs/design/range-reads.md` §2): aligned blocks of 1 MiB by default, LRU
with a byte budget (64 MiB by default), the missing blocks of one read
fetched as runs of consecutive blocks (a gap of one block is fetched to
merge two runs), each request at most 8 MiB, the requests of one read in
parallel. A read that misses more than half the budget is not kept, so a
big dataset read does not evict the metadata. Readers on several threads
share one cache; a block is never fetched twice at once — a second reader
waits for the first one's request.
- **Opening costs one request**: a `GET` of the first block, whose
`Content-Range` gives the file's length. HDF5 files keep the superblock
and usually the root group's metadata there, so listing a small file
often needs nothing more.
- **Pinned to one version of the file**: a strong `ETag` is sent back as
`If-Match` (else `Last-Modified` as `If-Unmodified-Since`) and checked on
every response, as is the length. A file replaced while open is an error
(`RemoteError::FileChanged`), never a mix of old and new bytes. A server
with neither validator can only be checked by length;
`HttpOptions::require_validator` refuses it.
- **Servers that ignore `Range`** (answer `200` with the whole file) are
refused (`RemoteError::RangeNotSupported`) without reading the body,
unless `HttpOptions::allow_full_download` is set; then the file is
downloaded once and read from memory.
- **Retries**: connection failures, timeouts, `408`/`429`/`5xx` and bodies
that end early are retried with exponential backoff (3 retries, from
200 ms). Bodies are requested with `Accept-Encoding: identity`; an encoded
body is refused.
The zero-copy methods of `clawhdf5` (`read_raw_ref`, `read_*_zerocopy`,
`File::as_bytes`) borrow the whole file from memory, so they are errors
(`as_bytes` a panic; use `File::contiguous_bytes`) on a remote file.
## Features
| Feature | What | C code |
|---|---|---|
| `http` (default) | `http://` through `ureq`, no TLS | none |
| `https` | `https://` through rustls, ring provider, Mozilla roots | ring (C and assembly) |
The default build compiles no C (`scripts/ci-test.sh` checks it).
## Counting requests
```rust
use clawhdf5_remote::{storage_for_url, Options};
let storage = storage_for_url(url, &Options::default())?;
let file = clawhdf5::File::open_storage(storage.clone())?;
// ... read ...
let s = storage.stats(); // requests, bytes_fetched, hits, misses, cached_bytes, ...
```
`cargo run -p clawhdf5-remote --example range_server -- DIR` serves a
directory with range support (the server the tests use), and
`cargo run -p clawhdf5-remote --example read_url -- URL [DATASET]` lists a
file and prints what it cost.