Range-read milestone M3, first half: a new crate with the block cache the
design makes mandatory for remote files and an HTTP backend, so
open_url("http://...") gives a clawhdf5::File over File::open_storage.
BlockCache wraps any Storage: aligned blocks (1 MiB by default, the size
docs/design/range-reads.md section 2 measured), LRU with a byte budget,
the missing blocks of one read_at/read_ranges fetched with one backend
read_ranges call as runs of consecutive blocks (a one-block gap filled to
merge runs, each request at most 8 MiB), and reads that miss more than
half the budget not kept. Thread-safe without holding the lock across a
fetch: a block being fetched is in flight, a second reader waits for it
instead of fetching it again, and a failed fetch fails its waiters and is
not cached. A backend holding the file in memory passes through.
HttpStorage (ureq, no TLS by default; `https` adds rustls with ring):
opening is one ranged GET of the first block, whose Content-Range gives
the length (the cache keeps the bytes). The file is pinned by a strong
ETag (If-Match), else Last-Modified (If-Unmodified-Since), and its length,
checked on every response: a change is RemoteError::FileChanged, never
mixed data. A server that ignores Range is refused without reading the
body unless a full download is allowed. Connection errors, timeouts,
408/429/5xx and short bodies are retried with exponential backoff;
Accept-Encoding: identity, and an encoded body is refused. read_ranges
fetches its ranges in parallel.
Tests (a std-only HTTP/1.1 server in tests/common/server.rs, also the
range_server example): every fixture read over HTTP gives File::open's
transcript (CLAWHDF5_REMOTE_CORPUS adds the conformance corpus), with
request counts per file with and without the cache; an h5py-written file
against libhdf5's values; a multi-block file fetched in whole blocks, each
once; a server ignoring Range; a file replaced mid-read (ETag,
Last-Modified, length only); truncated bodies and 503s (retried, then an
error, never cached); a slow server with 8 concurrent readers (no block
fetched twice); bad URLs, 404, encoded bodies, non-HDF5 data. The cache
has unit tests for coalescing, splitting, LRU order, large reads,
failures and concurrent in-flight dedup.
ci-test.sh: clawhdf5-remote joins the no-C default-build check, and its
https feature is linted.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
72 lines
3.2 KiB
Markdown
72 lines
3.2 KiB
Markdown
# clawhdf5-remote
|
|
|
|
Read HDF5 files where they live — on an HTTP(S) server — with
|
|
[clawhdf5](../../README.md), without downloading them first.
|
|
|
|
```rust
|
|
let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/data.h5")?;
|
|
let temps = file.dataset("/grid/temperature")?.read_f64()?;
|
|
```
|
|
|
|
The result is an ordinary `clawhdf5::File`: groups, datasets, attributes,
|
|
selections, variable-length data. Only the bytes an operation needs are
|
|
fetched, by `Range` requests, through a block cache.
|
|
|
|
## How it reads
|
|
|
|
- **Block cache** (`BlockCache`, mandatory for remote files; see
|
|
`docs/design/range-reads.md` §2): aligned blocks of 1 MiB by default, LRU
|
|
with a byte budget (64 MiB by default), the missing blocks of one read
|
|
fetched as runs of consecutive blocks (a gap of one block is fetched to
|
|
merge two runs), each request at most 8 MiB, the requests of one read in
|
|
parallel. A read that misses more than half the budget is not kept, so a
|
|
big dataset read does not evict the metadata. Readers on several threads
|
|
share one cache; a block is never fetched twice at once — a second reader
|
|
waits for the first one's request.
|
|
- **Opening costs one request**: a `GET` of the first block, whose
|
|
`Content-Range` gives the file's length. HDF5 files keep the superblock
|
|
and usually the root group's metadata there, so listing a small file
|
|
often needs nothing more.
|
|
- **Pinned to one version of the file**: a strong `ETag` is sent back as
|
|
`If-Match` (else `Last-Modified` as `If-Unmodified-Since`) and checked on
|
|
every response, as is the length. A file replaced while open is an error
|
|
(`RemoteError::FileChanged`), never a mix of old and new bytes. A server
|
|
with neither validator can only be checked by length;
|
|
`HttpOptions::require_validator` refuses it.
|
|
- **Servers that ignore `Range`** (answer `200` with the whole file) are
|
|
refused (`RemoteError::RangeNotSupported`) without reading the body,
|
|
unless `HttpOptions::allow_full_download` is set; then the file is
|
|
downloaded once and read from memory.
|
|
- **Retries**: connection failures, timeouts, `408`/`429`/`5xx` and bodies
|
|
that end early are retried with exponential backoff (3 retries, from
|
|
200 ms). Bodies are requested with `Accept-Encoding: identity`; an encoded
|
|
body is refused.
|
|
|
|
The zero-copy methods of `clawhdf5` (`read_raw_ref`, `read_*_zerocopy`,
|
|
`File::as_bytes`) borrow the whole file from memory, so they are errors
|
|
(`as_bytes` a panic; use `File::contiguous_bytes`) on a remote file.
|
|
|
|
## Features
|
|
|
|
| Feature | What | C code |
|
|
|---|---|---|
|
|
| `http` (default) | `http://` through `ureq`, no TLS | none |
|
|
| `https` | `https://` through rustls, ring provider, Mozilla roots | ring (C and assembly) |
|
|
|
|
The default build compiles no C (`scripts/ci-test.sh` checks it).
|
|
|
|
## Counting requests
|
|
|
|
```rust
|
|
use clawhdf5_remote::{storage_for_url, Options};
|
|
let storage = storage_for_url(url, &Options::default())?;
|
|
let file = clawhdf5::File::open_storage(storage.clone())?;
|
|
// ... read ...
|
|
let s = storage.stats(); // requests, bytes_fetched, hits, misses, cached_bytes, ...
|
|
```
|
|
|
|
`cargo run -p clawhdf5-remote --example range_server -- DIR` serves a
|
|
directory with range support (the server the tests use), and
|
|
`cargo run -p clawhdf5-remote --example read_url -- URL [DATASET]` lists a
|
|
file and prints what it cost.
|