Lazy remote files in the browser (M4), SWMR reader (M5), Python remote reads and editing #19

Merged
osobh merged 33 commits from feat/p3-wasm-swmr-python into main 2026-09-27 14:47:26 +00:00
64 changed files with 10170 additions and 959 deletions
+302
View File
@@ -2,6 +2,308 @@
## Unreleased
### Deterministic errors on damaged chunked datasets (2026-09-27)
- A read through the file's chunk cache listed a damaged dataset's chunks in
hash-map order, seeded per `File`, so two opens of the same file could
report different failing chunks (`cve-2025-2310.h5`, and an intermittent
failure of the storage and wasm corpus comparisons). The cache now keeps
the chunk index's own order, as the uncached readers use: every path names
the same first damaged chunk. Values of readable datasets were never
affected.
### Range reads, milestone M5: reading files a SWMR writer is appending to (2026-09-27)
Design: `docs/design/swmr.md`.
- **Fix: files with the SWMR-write flag were bounded by a stale end of
file** (since the end-of-file check of 2026-09-26, unreleased). A
libhdf5 SWMR writer (h5py `f.swmr_mode = True`) does not keep
the superblock's end-of-file address up to date; a copy h5py made of its
own file mid-write records 715 in a 6 030-byte file.
`Superblock::data_end` only ignored the recorded end when it lay past the
end of the file, so such a file listed, but every chunked read failed
("unexpected EOF: need 787 bytes, have 715") and `h5rs check` reported
its chunk indexes past the end of the file. For a v3 superblock with the
SWMR-write flag the data now ends at the end of the file, as libhdf5's
SWMR reader reads it (it skips the end-of-allocation check). Every open
path (`File::open`, `open_buffered`, `from_bytes`, `open_storage`,
`MmapFile`, `LazyFile`, `h5rs`) reads such a copy now.
- **`File::open_swmr(path)` / `File::open_storage_swmr(storage)`**: open a
file a SWMR writer may still be appending to, as libhdf5's SWMR reader
(h5py `File(path, "r", swmr=True)`) does. The file is read with
positioned reads (the new `clawhdf5::FileStorage`: `pread`/`seek_read`,
never mapped, `len()` the file's current length), reads are bounded by
the file's length at the time of each read, and the chunk cache is not
used (a cached chunk index would hide new chunks; a cached edge chunk
would read as fill where the writer has since written). A file open for
writing without SWMR is refused with `Error::Locked`, as libhdf5 refuses
it.
Only a file whose superblock has the SWMR-write flag when it is opened is
read this way; any other file (its writer has closed it) reads exactly as
`File::open` reads it, bounded by its recorded end of file, and
`is_swmr_read()` is `false`.
- **`Dataset::refresh()`** reads the dataset's object header again
(`H5Drefresh`, h5py `Dataset.refresh()`), so `shape()` and later reads
see the writer's appends; a handle keeps its extent until refreshed.
- **Bounded retries.** On a SWMR-read file, an operation (open, lookups
and listings, refresh, every dataset read including strings and
variable-length data, attributes, `File::decode_*`) that fails with an error a concurrent write can cause —
the failures libhdf5's SWMR reader retries: a checksum mismatch, a read
past the file's current end, an object header prefix that does not
decode — is run again from the start, up to
`File::swmr_read_attempts()` times (`SWMR_READ_ATTEMPTS` = 100,
libhdf5's default metadata read attempts for SWMR readers,
`H5Pset_metadata_read_attempts`; `set_swmr_read_attempts` changes it),
pausing 1 µs doubling to 10 ms between attempts. Results come only from
an attempt in which every structure verified, so a torn read is at worst
an error, never data. Any other error is returned at once.
`File::swmr_retries()` counts the retries.
`File::swmr_writer_active()` reads the superblock flags again to tell
when the writer has closed the file.
- Tests (`crates/clawhdf5/tests/swmr_interop.rs`): the mid-write copy
(fixture `tests/fixtures/swmr_mid_write.h5`) through every open path and
against h5py's SWMR reader; garbled reads (a storage that corrupts the
next reads) retried, never returned, and given up after the attempts;
every read path (listings, attributes compact and dense, strings,
variable-length data; fixture `tests/fixtures/swmr_strings_attrs.h5`)
with each of its reads failing once in turn returns the same result;
and a live test: an h5py writer appends to a 1-D and a 2-D dataset with
one unlimited dimension (Extensible Array index, one of them gzip) and a
2-D dataset with two (v2 B-tree index) for 2 500 steps
(`CLAWHDF5_SWMR_STEPS`), flushing after each, while two Rust reader
threads refresh and read all three in a loop — every value must be the
one the writer wrote at its position, extents never shrink — beside
h5py's own SWMR reader doing the same checks; after the writer closes,
the live handle and a new `File::open` read exactly what h5py reads.
Run on tank, 2026-09-27, with h5py 3.16 / HDF5 2.0:
`CLAWHDF5_REQUIRE_INTEROP=1 CLAWHDF5_PYTHON=.venv/bin/python cargo test
-p clawhdf5 --test swmr_interop`; also at 20 000 steps in a release
build. A test with the chunk cache left on in live mode fails it.
- Not covered: SWMR writing, remote SWMR (`BlockCache` caches blocks and
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
groups or attributes (a SWMR writer cannot add objects or attributes).
### Correctness: edits planned from another file after a rename or `chdir` (2026-09-27)
- **`FileEditor` planned each edit by re-opening its path but wrote
through the file it held open** (fixed 2026-09-27; on main since PR #18,
no release). When the path came to name another file between edits — a
rename or replacement, or, for a relative path, a change of working
directory — an edit was laid out from the other file's metadata and
written into the held one, corrupting it (h5py: "invalid dataset size,
likely file corruption"). The Python `'r+'` handle had the same flaw in
its reads: it reopened the path after every edit, so reads came from the
other file. The editor now plans every edit from the file it holds, and
its path is canonicalised at open. New `FileEditor::reader()` opens the
held file anew for reading (on Linux through `/proc/self/fd`, so it
follows a renamed file; elsewhere by the path, refused when the path no
longer names the held file), without sharing the editor's lock; the
Python handle reads through it, and a `'w'` file is written at the
absolute path it was opened with. Tests: `edit_tests.rs`'s
`edits_go_to_the_file_held_not_the_path`; `test_edit.py`'s
`test_relative_path_and_chdir` and `test_path_replaced_between_edits`
(the review's repro).
### Correctness: zero extents in Fixed/Extensible Array chunk indexes (2026-09-27)
- **A chunked dataset whose maximum (or, with none recorded, current)
extent is 0 along a dimension made the reader divide by zero** (fixed
2026-09-27): `h5rs check` panicked ("attempt to divide by zero",
`chunk_grid.rs`) and the next `FileEditor::resize` failed with an
internal error. The unfixed editor produced such files by resizing a
clawhdf5-written dataset to a zero extent (12 of 30 extra random-edit
seeds on clawhdf5-written files). Such an index has no slot for any
chunk of the dataset, and `ChunkGrid::offsets` now says so instead of
dividing by the zero stride. Tests: `chunk_grid`'s
`zero_extent_has_no_chunks`, `edit_interop.rs`'s
`zero_extent_resizes_without_a_recorded_maximum` (a file the unfixed
editor left checks clean and resizes on; a 2.7.0-written file through
zero extents checks clean at each step), and `test_edit.py`'s random
edits on seeds 10 to 39 of a clawhdf5-written file.
### Correctness: resizing chunked datasets with no recorded maximum (2026-09-27)
- **`FileEditor::resize` scrambled the values of a chunked dataset whose
dataspace records no maximum dimensions when it shrank it** (fixed
2026-09-27). The editor shipped on main in PR #18 (a4c2ace) and reached
Python as `Dataset.resize` in `'r+'` files; no release has it. clawhdf5's
writer stores such a dataspace for every chunked dataset created without
a `maxshape`, with a Fixed Array (or Single Chunk) chunk index. The
editor patched only the current dimensions, and with no maximum
recorded the maximum is the current dimensions — which is also what the
Fixed Array linearises chunks by — so a shrink moved every chunk after
the first row: h5py, h5dump and our reader all read wrong values
without complaint (20x20, chunks 6x6, resized to 15x15: row 6 read
`0 0 0 0 0 0 120 ...`). After a shrink the dataset could not grow back
either (`3 exceeds the maximum 0`). libhdf5 itself never writes such a
dataspace (`H5S_set_extent_simple` records the maximum, equal to the
dimensions when none is given); reading one, `H5S_extent_get_dims`
reports the current dimensions as the maximum and `H5S_set_extent`
checks against no maximum at all, so libhdf5's own `H5Dset_extent`
scrambles such a file the same way (and lets it grow past its Fixed
Array). The editor now records the maximum libhdf5 would have written
— the dimensions before the first resize, the ones the index was built
with — then changes the current ones (the dataspace message grows by one
length per dimension and moves in the header when it has to). The
dataset then shrinks, grows back to that extent and refuses more, as
one libhdf5 wrote would. The writer (`FileBuilder`) now records the
maximum of every chunked dataset too, as libhdf5 does, so h5py can
resize what it writes (8 more bytes per dimension). Tests:
`crates/clawhdf5/tests/edit_resize_interop.rs` (a 2.7.0-written fixture,
new `FileBuilder` files and h5py files through shrinks, zero extents
and growth, checked against a model with our reader and h5py; h5py
resizing a `FileBuilder` file) and `test_edit.py`'s numpy-model checks.
### Python bindings: in-place editing (2026-09-27)
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
file for editing through `clawhdf5::FileEditor`, holding its exclusive
lock until `close()`:
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
one increasing index list) and broadcasting (numpy's rules for slices
and integers, allowing extra leading length-1 axes; the exact shape for
an index list, a scalar only where h5py expands it). A numpy array is
converted to the dataset's dtype as libhdf5 converts it in native byte
order (integers saturate; floats are truncated toward zero and clipped;
integers go into h5py's bool enum by value, as libhdf5 stores them);
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
does. NaN into an integer dataset is a `ValueError`.
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
and errors (`TypeError` for a dataset that is not chunked).
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
enum, complex as the `r`/`i` compound), except that `str` becomes
fixed-length UTF-8.
- Every edit is written and synced before it returns, then the file is
reopened: datasets and attrs objects taken earlier see new shapes and
attributes, and reads on other threads wait while an edit is written.
- What the editor cannot do raises `NotImplementedError` and writes
nothing (deleting attributes or objects, creating datasets or groups,
compound fields by name, variable-length data, ...;
`docs/known-issues.md`).
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
- Boolean-mask keys (`ds[mask]`, `ds[mask] = v`), which h5py supports,
raise `NotImplementedError` (they raised `TypeError`).
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
by h5py and by clawhdf5, and both read back through h5py after every
edit, on files h5py writes with `libver` earliest, v114 and latest and on
a clawhdf5-written one: a fixed sequence over every chunk index kind,
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
and compound types, and 16 random sequences of 40 edits (writes,
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
and leave its file unchanged. A matrix of every numeric source dtype into
every numeric dataset dtype at the edge values, dense attribute storage,
locking, objects seeing each other's edits, readers racing a writer
(never a partly written dataset). Every edited file must pass `h5dump`
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
a file opened `'r+'`.
### Python bindings: remote files (2026-09-27)
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
through the block cache, the whole read API (groups, attributes, every
dataset type and index the local reader handles). A URL is any
`scheme://…`; a remote file is read-only (another mode is a
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
`allow_full_download`, `max_full_download`, `require_validator`,
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
cache's counters. The default wheel builds plain HTTP only (no C: rustls
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
check now covers `clawhdf5-py`.
- **Every read parses through `File::storage()`** instead of
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
group listings, variable-length data through the global heap), inside one
shared file handle that releases the GIL for all file access, not only
dataset reads: a read waiting on the network lets other Python threads
run. A failed read of the storage (a network error, a file changed on
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
data; `key in group` raises it instead of answering `False`.
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
blocks and 1 KiB blocks) against a range-capable `http.server` in the
test process, plus `tests/test_remote.py`: request counts of a small
read, cache hits, a server without `Range` support (refused, or a
whole download when allowed), a file changed on the server, a server
that hangs up, 16 threads on one remote file, and a thread that keeps
running while a read waits on 0.2 s requests.
### Range reads, milestone M4: remote files in the browser (2026-09-27)
- **`openUrl(url, opts)`** in `clawhdf5-wasm` opens an HDF5/NetCDF-4 file
on a web server without downloading it: the returned `RemoteFile` has
`H5File`'s methods (`kind`, `list`, `info`, `attrs`, `attrErrors`,
`read`, `readHyperslab`), each returning a promise, and `stats()`
(requests, bytes fetched, file size). Only the byte ranges a call needs
are fetched, with `fetch` and `Range` headers, in 1 MiB blocks
(`blockSize`) kept in a 64 MiB cache (`cacheSize`). `open(bytes)` is
unchanged.
- **How, on the browser's main thread:** the design's restartable
"NeedBytes" mode (`clawhdf5_wasm::lazy::LazyStorage`). A call runs as a
pass over the blocks fetched so far; a read that misses records the
missing blocks and fails, the pass's result is dropped (even if a parser
caught the error and carried on), the blocks are fetched and the pass is
re-run. No Web Worker and no synchronous XHR (what h5wasm's lazy files
need). No block is evicted while a call is in flight, so every call ends;
the cache is trimmed between calls, raw-data blocks before metadata.
`clawhdf5-remote`'s `BlockCache` is not reused: it fetches by blocking,
and it evicts, and does not keep large reads, during a read, where a
restartable pass needs every block it read to stay until it finishes.
- **Every answer is checked** (`js/remote.js`): a `206` with exactly the
bytes asked for (`Content-Range`, when the page can see it, and the body
length), and the same `ETag`/`Last-Modified` and length as at open, or
the call fails: never data from another file or offset. A server that
answers `200` to a `Range` request is downloaded whole (up to
`maxDownload`, 512 MiB) unless `fallback: "error"`. Other options:
`headers`, `credentials`, `parallel` (6 requests at a time), `fetch`.
- **The viewer** (`examples/wasm-viewer`) has a URL box, opens
`?file=<url>` lazily, and shows the requests and bytes fetched.
- Counted on tank, 2026-09-27 (`WASM_BIG_MB=200 bash
examples/wasm-viewer/test/run.sh`, requests as `test/serve.py` counted
them): in a 200 MB h5py file, listing the root, reading two small
datasets, a group's attributes, the large dataset's shape and a
10-value window of it (25 million values) took 5 requests and 6 MiB.
With
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, the 622 corpus files
up to 16 MiB that open read over HTTP as from bytes: the same listings,
attributes and values (datasets up to 2^20 values), and an error
wherever bytes give one. Natively, `cargo test -p clawhdf5-wasm --test
lazy` with the same variable compares 656 files (up to 64 MiB, error
messages included, against the facade's range-storage path).
- `Reader::open_storage` (the wasm crate's core over any `Storage`), and
variable-length strings resolve through the file's storage rather than
`File::as_bytes`.
- The package is larger: the facade's `Storage` read path is now
reachable from JavaScript (it was compiled out before), and the promise
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
build machine was shared); the viewer README's size table predates M4.
- **Hardened after review (2026-09-27):**
- Sizes a server or a dataset names are errors, never an abort of the
wasm module (which took every open file on the page with it): a read
past 2 GiB aborted in the lazy cache, reachable by a hostile server
claiming a large file and a 2 GiB heap collection, and `read()` of a
256 MiB `u8` dataset aborted widening it to 64 bits. New option
`maxFetch` (512 MiB, at most 1 GiB): what one call may fetch, and the
longest single read, refused before fetching. `read()` refuses a
dataset that would take more than 1 GiB to decode, naming
`readHyperslab`. A file of 4 GiB or more is refused at open (wasm32
reads offsets as 32-bit); `maxDownload` is at most 1 GiB.
- Response bodies are read as they arrive and cut off at the length
asked for (`maxDownload` for a `200`): a `206` with a gigabyte body
was buffered whole before its length was checked.
- Listing a group reads every child's header, and every node of each
level of the group's index, in one pass: 3000 datasets (h5py, 198 MB)
listed in 6 passes and 73 requests at 1 MiB blocks instead of 185
passes and 184 serial requests (`libver="latest"`: 9 passes instead of
189). In `clawhdf5-format`, the B-tree v1/v2 collectors, the symbol
table node loop and the dense-link loop read (without using) the
siblings after the first that fails, then return that error: same
results and errors, more reads only on failure (free in memory). The
lazy cache no longer re-fetches a cached block to merge two requests.
- `headers` may be a `Headers` instance or `[name, value]` pairs (a
`Headers` was silently dropped); when one range request fails the
others in flight are aborted; `parallel` must be an integer from 1 to
1024.
- Tests: `test/serve.py` serves ranges without exposed `Content-Range`/
`ETag` (`/noexpose/`, and `/unexposed/` for a real cross-origin page
in Chromium), so the HEAD-length path runs end to end; hostile and
oversized files (`make_fixture.py`'s `write_limits`), flooding
bodies, aborted siblings.
### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
+13 -8
View File
@@ -181,14 +181,19 @@ Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal F
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
file over HTTP with `File::open`.
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only, file held in memory;
no Zstd/SZIP since they link C) and the `examples/wasm-viewer/` page.
`examples/wasm-viewer/test/run.sh` builds the package (needs the
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node
and headless Chromium (a Playwright download in `~/.cache/ms-playwright`
on tank); the CI container has neither, so CI runs the native
`clawhdf5-wasm` `h5py_interop` test on the same fixture. Size numbers are
in the example's README.
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
call is re-run after each wave of misses; no block evicted while a call
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
version) and tests it under Node and headless Chromium (a Playwright
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
(range server with request counts, 200 MB budget file); the CI container
has neither, so CI runs the native `h5py_interop` and `lazy` tests
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
numbers in the example's README predate `openUrl`.
- Python and Node.js bindings for cross-language use
- NetCDF-4 compatibility for scientific data interop
+72 -2
View File
@@ -101,6 +101,10 @@ breaking change, are in [CHANGELOG.md](CHANGELOG.md).
HTTP server (or in S3/GCS/Azure, opt-in) by range requests through a
block cache, without downloading it; `h5rs` takes URLs with its `remote`
feature. See [Reading remote files](#reading-remote-files).
- `File::open_swmr` follows a file an h5py/libhdf5 SWMR writer is still
appending to (`Dataset::refresh`, bounded retries); copies of such files
taken mid-write read with every open path. See
[Following a file a SWMR writer is appending to](#following-a-file-a-swmr-writer-is-appending-to).
**Tooling**
- CI now runs the h5py/netCDF4 interop suites for real (they had been skipping
@@ -504,6 +508,43 @@ built with `--features remote` takes the same URLs:
`https://` is the `https` feature (rustls with ring, which compiles C).
Limits are in [known issues](docs/known-issues.md).
### Following a file a SWMR writer is appending to
`File::open_swmr` reads a file that a libhdf5 writer in SWMR mode (h5py
`f.swmr_mode = True`) is still appending to, as h5py's
`File(path, "r", swmr=True)` does: `Dataset::refresh()` picks up the new
extent, every read reads the chunk index as it is now, and a read that
races the writer (a checksum that fails mid-flush) is retried, up to 100
attempts as in libhdf5, and never returned torn.
```rust
use std::time::{Duration, Instant};
let file = clawhdf5::File::open_swmr("live.h5")?;
let mut ds = file.dataset("samples")?;
let (mut seen, mut last_growth) = (0, Instant::now());
// Stop when the writer closes the file, or when the dataset has not grown
// for a minute: a writer that crashed or was killed never clears the
// SWMR-write flag, so `swmr_writer_active()` alone can stay true forever.
while file.swmr_writer_active()? && last_growth.elapsed() < Duration::from_secs(60) {
ds.refresh()?;
let n = ds.shape()?[0];
if n > seen {
// read rows seen..n ...
(seen, last_growth) = (n, Instant::now());
}
std::thread::sleep(Duration::from_millis(100));
}
ds.refresh()?; // the final extent
```
`swmr_writer_active()` reads the superblock's SWMR-write flag, which libhdf5
clears only when the writer closes the file; a file whose writer died keeps
it set (as the mid-write copy in `tests/fixtures/swmr_mid_write.h5` does),
so a follower needs its own stop condition, like the idle timeout above.
Design and limits: [docs/design/swmr.md](docs/design/swmr.md).
### Python
`crates/clawhdf5-py` is a Python package (PyO3 + numpy) that reads HDF5 with
@@ -533,8 +574,36 @@ with clawhdf5.File("data.h5", "r") as f:
records = f["table"] # compound -> numpy structured array
ids = records["id"] # one field
# A file on a web server: range requests through a block cache, nothing
# downloaded up front; the same read API. The GIL is released while waiting.
with clawhdf5.File("http://data.example.org/run42.h5") as f:
first = f["group/temperatures"][0]
f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024,
headers={"Authorization": "Bearer ..."})
```
An existing file opened with `"r+"` is edited in place (through
`clawhdf5::FileEditor`), with h5py's indexing, broadcasting and numeric
conversion; each edit is on disk when the statement returns:
```python
with clawhdf5.File("data.h5", "r+") as f:
f["group/temperatures"][100:200, ::4] = 0.0
f["series"].resize(5000, axis=0) # chunked datasets, within maxshape
f["series"][4000:] = new_values
f["group"].attrs["calibrated"] = True
```
Creating or deleting datasets, groups and attributes in an existing file is
not supported (`NotImplementedError`); limits are in
[known issues](docs/known-issues.md).
The default build reads `http://` URLs only; build with
`maturin develop --release --features https` (rustls with ring, which
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
object-store URLs.
Reads cover integers and IEEE floats of every width in either byte order,
`bool`, enums, complex, fixed and variable-length strings, variable-length
sequences, opaque, HDF5 array types and compounds; other types (references,
@@ -549,7 +618,8 @@ non-default fill value (`docs/known-issues.md`). An index list is read one
group of neighbouring chunks at a time.
Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`)
covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests
in `crates/clawhdf5-py/tests` compare every read with h5py; run them with
in `crates/clawhdf5-py/tests` compare every read and every in-place edit
with h5py; run them with
`pip install pytest h5py && pytest crates/clawhdf5-py/tests`.
### Agent Memory
@@ -757,7 +827,7 @@ clawhdf5 workspace (19 crates, ~86K lines of Rust in src/, ~104K with tests
├── Bindings
│ ├── clawhdf5-py — Python (PyO3)
│ ├── clawhdf5-napi — Node.js (napi-rs)
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only)
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only; remote files by HTTP range requests)
│
└── Tooling
├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check
+53 -18
View File
@@ -576,7 +576,14 @@ pub fn find_attribute_in_file(
offset_size: u8,
length_size: u8,
) -> Result<Option<AttributeMessage>, FormatError> {
find_attribute_core(file_data, header, name, offset_size, length_size)
find_attribute_core(
file_data,
header,
name,
offset_size,
length_size,
&mut Vec::new(),
)
}
/// [`find_attribute_in_file`] over any [`Storage`] (see
@@ -594,16 +601,48 @@ pub fn find_attribute_in<S: Storage + ?Sized>(
) -> Result<Option<AttributeMessage>, FormatError> {
match file_data.as_contiguous() {
Some(all) => find_attribute_in_file(all, header, name, offset_size, length_size),
None => find_attribute_core(file_data, header, name, offset_size, length_size),
None => find_attribute_core(
file_data,
header,
name,
offset_size,
length_size,
&mut Vec::new(),
),
}
}
/// [`find_attribute_in`], also returning the errors of the attributes it
/// could not read on the way (which it leaves out rather than failing
/// the call): the attribute asked for may be one of them. A reader of a
/// file that is being written uses them to tell a read that raced the
/// writer from an absent attribute.
pub fn find_attribute_reporting_in<S: Storage + ?Sized>(
file_data: &S,
header: &ObjectHeader,
name: &str,
offset_size: u8,
length_size: u8,
) -> Result<(Option<AttributeMessage>, Vec<FormatError>), FormatError> {
let mut errors = Vec::new();
let found = find_attribute_core(
file_data,
header,
name,
offset_size,
length_size,
&mut errors,
)?;
Ok((found, errors))
}
fn find_attribute_core<S: Storage + ?Sized>(
file_data: &S,
header: &ObjectHeader,
name: &str,
offset_size: u8,
length_size: u8,
errors: &mut Vec<FormatError>,
) -> Result<Option<AttributeMessage>, FormatError> {
let attr_info = find_attribute_info(header, offset_size)?;
let dense = attr_info
@@ -612,12 +651,10 @@ fn find_attribute_core<S: Storage + ?Sized>(
let Some((fh_addr, btree_addr)) = dense else {
// Compact only (or dense storage without a name index, which a
// listing reports): as a listing finds it.
return Ok(
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?
.0
.into_iter()
.find(|a| a.name == name),
);
let (attrs, errs) =
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?;
errors.extend(errs);
return Ok(attrs.into_iter().find(|a| a.name == name));
};
let btree_hdr = BTreeV2Header::parse_in(
file_data,
@@ -627,12 +664,10 @@ fn find_attribute_core<S: Storage + ?Sized>(
)?;
let fh = FractalHeapHeader::parse_in(file_data, fh_addr, offset_size, length_size)?;
if btree_hdr.tree_type != ATTRIBUTE_NAME_INDEX || btree_hdr.record_size < 4 {
return Ok(
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?
.0
.into_iter()
.find(|a| a.name == name),
);
let (attrs, errs) =
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?;
errors.extend(errs);
return Ok(attrs.into_iter().find(|a| a.name == name));
}
// A listing has the compact attributes first.
@@ -671,10 +706,10 @@ fn find_attribute_core<S: Storage + ?Sized>(
AttributeMessage::parse_in_storage(&d, file_data, offset_size, length_size)
});
// One that cannot be read is left out, as from a listing.
if let Ok(attr) = attr
&& attr.name == name
{
return Ok(Some(attr));
match attr {
Ok(attr) if attr.name == name => return Ok(Some(attr)),
Ok(_) => {}
Err(e) => errors.push(e),
}
}
Ok(None)
+18 -5
View File
@@ -199,19 +199,32 @@ fn collect_symbol_table_nodes_inner<S: Storage + ?Sized>(
// Leaf: children are SNOD addresses
Ok(node.children)
} else {
// Internal: recurse into children
// Internal: recurse into children. After the first child that
// fails, the others are only read (as `storage::touch` does), not
// descended into; that error is returned.
let mut result = Vec::new();
let mut failed = None;
for &child_addr in &node.children {
let child_snods = collect_symbol_table_nodes_inner(
if failed.is_some() {
// Parsing reads the node's header, then its body.
let _ = BTreeV1Node::parse_in(file, child_addr, offset_size, length_size);
continue;
}
match collect_symbol_table_nodes_inner(
file,
child_addr,
offset_size,
length_size,
depth + 1,
)?;
result.extend(child_snods);
) {
Ok(child_snods) => result.extend(child_snods),
Err(e) => failed = Some(e),
}
}
match failed {
Some(e) => Err(e),
None => Ok(result),
}
Ok(result)
}
}
+19 -1
View File
@@ -504,7 +504,18 @@ fn collect_internal_records<S: Storage + ?Sized>(
// Interleave: child[0], record[0], child[1], record[1], ..., child[nr]
// We collect child[0] records, then record[0], then child[1], etc.
// After the first child that fails, the others are only touched (see
// `storage::touch`); that error is returned.
let mut failed = None;
for (i, &(child_addr, child_nrec)) in node.children.iter().enumerate() {
if failed.is_some() {
let len = usize::try_from(node_size)
.unwrap_or(usize::MAX)
.min(1 << 16);
crate::storage::touch(file, child_addr, len);
continue;
}
if let Err(e) = (|| -> Result<(), FormatError> {
if child_depth == 0 {
// Before parsing, so a refused tree is not also a large allocation.
spend(budget, usize::from(child_nrec))?;
@@ -540,9 +551,16 @@ fn collect_internal_records<S: Storage + ?Sized>(
data: data.to_vec(),
});
}
Ok(())
})() {
failed = Some(e);
}
}
Ok(())
match failed {
Some(e) => Err(e),
None => Ok(()),
}
}
/// The records of a B-tree v2 that fall in one key range, found by
+27 -7
View File
@@ -262,6 +262,11 @@ struct CachedChunk {
struct DatasetEntry {
/// Chunk coordinate -> ChunkInfo (offset + size in file).
index: Option<Arc<HashMap<ChunkCoord, ChunkInfo>>>,
/// The same chunks in the order the chunk index lists them: what
/// [`ChunkCache::chunks_for`] returns, so a cached read walks (and, on a
/// damaged file, fails at) the chunks in the same order as an uncached
/// one, rather than in hash-map order.
ordered: Option<Arc<Vec<ChunkInfo>>>,
/// Pre-built chunk index for O(1) coordinate lookups.
chunk_index: Option<Arc<ChunkIndex>>,
/// Pre-computed chunk layout for fast assembly.
@@ -274,6 +279,7 @@ struct DatasetEntry {
impl DatasetEntry {
fn weight(&self) -> usize {
self.index.as_ref().map_or(0, |m| m.len())
+ self.ordered.as_ref().map_or(0, |o| o.len())
+ self.chunk_index.as_ref().map_or(0, |c| c.num_chunks())
}
}
@@ -562,11 +568,22 @@ impl ChunkCache {
rank: usize,
build: impl FnOnce() -> Result<Vec<ChunkInfo>, E>,
) -> Result<Vec<ChunkInfo>, E> {
Ok(self
.index_for(addr, rank, build)?
.values()
.cloned()
.collect())
if let Some(ordered) = self.lock().touch(addr).ordered.clone() {
return Ok(ordered.as_ref().clone());
}
let chunks = build()?;
let map: HashMap<ChunkCoord, ChunkInfo> = chunks
.iter()
.map(|ci| (ci.offsets.iter().take(rank).copied().collect(), ci.clone()))
.collect();
let mut inner = self.lock();
let entry = inner.touch(addr);
// Another thread may have built this dataset's index meanwhile: keep
// the first one, so every reader sees the same order.
let ordered = Arc::clone(entry.ordered.get_or_insert_with(|| Arc::new(chunks)));
entry.index.get_or_insert_with(|| Arc::new(map));
inner.trim_datasets(addr);
Ok(ordered.as_ref().clone())
}
fn index_for<E>(
@@ -580,11 +597,14 @@ impl ChunkCache {
}
let chunks = build()?;
let map: HashMap<ChunkCoord, ChunkInfo> = chunks
.into_iter()
.map(|ci| (ci.offsets.iter().take(rank).copied().collect(), ci))
.iter()
.map(|ci| (ci.offsets.iter().take(rank).copied().collect(), ci.clone()))
.collect();
let mut inner = self.lock();
let entry = inner.touch(addr);
if entry.index.is_none() {
entry.ordered = Some(Arc::new(chunks));
}
let index = Arc::clone(entry.index.get_or_insert_with(|| Arc::new(map)));
inner.trim_datasets(addr);
Ok(index)
+28
View File
@@ -135,6 +135,12 @@ impl ChunkGrid {
let mut rem = index;
for p in 0..rank {
let d = self.order[p];
// A zero stride: a later dimension has no chunks (its maximum,
// or with none recorded its current extent, is 0), so no slot of
// the index is a chunk of the dataset.
if self.down[p] == 0 {
return None;
}
let scaled = rem / self.down[p];
rem %= self.down[p];
if scaled >= self.cur_chunks[d] {
@@ -193,6 +199,28 @@ mod tests {
assert_eq!(g.offsets(11), Some(vec![2, 3]));
}
#[test]
fn zero_extent_has_no_chunks() {
// No maximum recorded and a zero current dimension: every stride
// before it is 0 (this divided by zero).
let g = ChunkGrid::fixed_array(&[1, 0], None, &[6, 6]).unwrap();
for i in 0..16 {
assert_eq!(g.offsets(i), None);
}
let g = ChunkGrid::fixed_array(&[0, 0, 3], Some(&[4, 0, 3]), &[2, 2, 3]).unwrap();
for i in 0..16 {
assert_eq!(g.offsets(i), None);
}
let g = ChunkGrid::extensible_array(&[0, 5], Some(&[u64::MAX, 0]), &[2, 2]).unwrap();
for i in 0..16 {
assert_eq!(g.offsets(i), None);
}
// A zero last dimension leaves the other strides alone.
let g = ChunkGrid::fixed_array(&[4, 0], Some(&[4, 6]), &[2, 3]).unwrap();
assert_eq!(g.offsets(0), None);
assert_eq!(g.linear_index(&[1, 1]), 3);
}
#[test]
fn rejects_two_unlimited_dims_after_the_first() {
assert!(ChunkGrid::fixed_array(&[4, 6], Some(&[u64::MAX, u64::MAX]), &[2, 3]).is_err());
+16
View File
@@ -130,6 +130,22 @@ pub(crate) fn build_chunked_dataset_oh(
fill_message: &[u8],
refcount: u32,
) -> Result<Vec<u8>, FormatError> {
// libhdf5 records every simple dataspace's maximum dimensions (the
// dimensions themselves when none are given, `H5S_set_extent_simple`).
// Without them libhdf5 takes the maximum to be the current dimensions,
// so a resize by libhdf5 (h5py's `Dataset.resize`) would also change
// the maximum a Fixed Array chunk index is laid out by, and move every
// chunk already written.
let recorded;
let ds = if ds.space_type == DataspaceType::Simple && ds.max_dimensions.is_none() {
recorded = Dataspace {
max_dimensions: Some(ds.dimensions.clone()),
..ds.clone()
};
&recorded
} else {
ds
};
let mut w = ObjectHeaderWriter::new();
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
+27 -3
View File
@@ -79,13 +79,29 @@ pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
length_size,
)?;
// The names are read one by one from the heap's data segment; read
// (up to 1 MiB of) it first, so a storage that records what it lacks
// asks for it at once (see `storage::touch`).
if !snod_addrs.is_empty() {
let len = usize::try_from(heap.data_segment_size).map_or(1 << 20, |n| n.min(1 << 20));
crate::storage::touch(file_data, heap.data_segment_address, len);
}
let mut entries = Vec::new();
let mut heap_checked = false;
// After the first node that fails, the others are only read (as
// `storage::touch` does); that error is returned.
let mut failed = None;
for snod_addr in snod_addrs {
if failed.is_some() {
let _ = SymbolTableNode::parse_in(file_data, snod_addr, offset_size);
continue;
}
let mut node = || -> Result<(), FormatError> {
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
for entry in &snod.entries {
// Like libhdf5, look at the heap's free list only once a name is
// needed: an empty group with a damaged heap still lists.
// Like libhdf5, look at the heap's free list only once a name
// is needed: an empty group with a damaged heap still lists.
if !heap_checked {
heap.validate_free_list_in(file_data, length_size)?;
heap_checked = true;
@@ -97,9 +113,17 @@ pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
cache_type: entry.cache_type,
});
}
Ok(())
};
if let Err(e) = node() {
failed = Some(e);
}
}
Ok(entries)
match failed {
Some(e) => Err(e),
None => Ok(entries),
}
}
/// Symbol table cache type for a soft link: the scratch pad's first four bytes
+15 -4
View File
@@ -126,6 +126,9 @@ fn for_each_dense_link<S: Storage + ?Sized>(
)?;
let records = collect_btree_v2_records_in(file_data, &btree_hdr, offset_size, length_size)?;
// After the first link that fails, the others are only read, not
// visited (a touch, see `storage::touch`); that error is returned.
let mut failed = None;
for record in &records {
// For type 5 (name index): hash(4) + heap_id(heap_id_length)
// For type 6 (creation order): creation_order(8) + heap_id(heap_id_length)
@@ -141,12 +144,20 @@ fn for_each_dense_link<S: Storage + ?Sized>(
let id_bytes = &record.data[id_offset..id_offset + fh.heap_id_length as usize];
// Read managed object from fractal heap
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size)?;
if let Some(link) = parse_link(&link_data, offset_size)? {
visit(link);
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size);
if failed.is_some() {
continue;
}
match link_data.and_then(|d| parse_link(&d, offset_size)) {
Ok(Some(link)) => visit(link),
Ok(None) => {}
Err(e) => failed = Some(e),
}
}
Ok(())
match failed {
Some(e) => Err(e),
None => Ok(()),
}
}
/// Resolve entries from dense storage (fractal heap + B-tree v2).
+13
View File
@@ -202,6 +202,19 @@ pub(crate) fn len_usize<S: Storage + ?Sized>(file: &S) -> usize {
usize::try_from(file.len()).unwrap_or(usize::MAX)
}
/// Read `len` bytes at `offset` and drop them, ignoring any error.
///
/// For a traversal that has failed on one sibling (a B-tree child, a
/// symbol table node, a heap object) and would stop there: it first
/// touches the siblings it did not get to, so a storage that records what
/// it lacks — the browser's restartable reader, which fetches over the
/// network between attempts — learns about all of them in one attempt
/// instead of one per attempt. Results and errors are unchanged (the first
/// error is still the one returned); an in-memory read is free.
pub fn touch<S: Storage + ?Sized>(file: &S, offset: u64, len: usize) {
let _ = file.read_at(offset, len);
}
/// Bytes `[offset, offset + len)`, all of them.
///
/// A range that runs past the end of the storage is
+19 -7
View File
@@ -116,22 +116,27 @@ impl Superblock {
/// [`FormatError::TruncatedFile`]. Bytes past that address are not part
/// of the file: libhdf5 fails any read of them ("addr overflow" /
/// "address plus size exceeds file eoa"), so a reader should parse only
/// the data up to the returned end. As libhdf5 does for a SWMR reader,
/// the check is skipped for a version-3 superblock whose writer is still
/// writing it in SWMR mode (it extends the file as it goes); the data
/// then ends at the end of the file.
/// the data up to the returned end.
///
/// A version-3 superblock with the SWMR-write flag set belongs to a file
/// a SWMR writer has open (or had, and did not close). That writer does
/// not keep the recorded end of file up to date — a copy taken mid-write
/// can record an end of a few hundred bytes in a file of tens of
/// kilobytes — and libhdf5's SWMR reader skips its end-of-allocation
/// check for every read (`H5FD_read`). For such a superblock the data
/// ends at the end of the file, whatever end it records.
///
/// When the superblock's recorded base address differs from where the
/// superblock actually is (a user block added or removed after the file
/// was written), libhdf5 moves the recorded end of file by the same
/// amount, and so does this.
pub fn data_end(&self, user_block: u64, file_len: u64) -> Result<u64, FormatError> {
let eof =
i128::from(self.eof_address) - i128::from(self.base_address) + i128::from(user_block);
if eof < 0 || eof > i128::from(file_len) {
if self.version >= 3 && self.is_swmr_write() {
return Ok(file_len.saturating_sub(user_block));
}
let eof =
i128::from(self.eof_address) - i128::from(self.base_address) + i128::from(user_block);
if eof < 0 || eof > i128::from(file_len) {
return Err(FormatError::TruncatedFile {
stored_eof: u64::try_from(eof).unwrap_or(self.eof_address),
actual_len: file_len,
@@ -617,6 +622,13 @@ mod tests {
let mut swmr = Superblock::parse(&build_v2_bytes(8, 3), 0).unwrap();
swmr.consistency_flags = swmr_flags::WRITE_ACCESS | swmr_flags::SWMR_WRITE;
assert_eq!(swmr.data_end(0, 1000), Ok(1000));
// ... nor bounded by its recorded end, which the writer does not
// keep up to date (2048 here).
assert_eq!(swmr.data_end(0, 17_857), Ok(17_857));
assert_eq!(swmr.data_end(512, 17_857), Ok(17_345));
// Without the SWMR-write flag the recorded end bounds the data.
swmr.consistency_flags = swmr_flags::WRITE_ACCESS;
assert_eq!(swmr.data_end(0, 17_857), Ok(2048));
}
#[test]
+17
View File
@@ -39,6 +39,23 @@ impl MmapReader {
Ok(Self { _file: file, mmap })
}
/// Memory-map a file that is already open (for reading).
///
/// The mapping references `file`'s open file description for as long as
/// it lives, so a `flock` taken through that description (or a
/// `try_clone` of it) is held until the reader is dropped.
///
/// # Safety
///
/// The same contract as [`open`](Self::open): the file must not be
/// modified while the mapping is active.
pub fn from_file(file: fs::File) -> io::Result<Self> {
// SAFETY: a read-only mapping; the caller keeps the file unmodified
// while it is alive.
let mmap = unsafe { Mmap::map(&file)? };
Ok(Self { _file: file, mmap })
}
/// Zero-copy access to the entire file contents.
pub fn as_bytes(&self) -> &[u8] {
&self.mmap
+9
View File
@@ -17,11 +17,20 @@ crate-type = ["cdylib", "rlib"]
[dependencies]
clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" }
clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
# Remote files (`clawhdf5.File(url)`): plain HTTP by default, which builds
# no C. HTTPS and the object stores are opt-in features below.
clawhdf5-remote = { path = "../clawhdf5-remote", version = "2.7.0" }
pyo3 = "0.29"
numpy = "0.29"
[features]
extension-module = ["pyo3/extension-module"]
# https:// URLs (rustls with ring, which compiles C and assembly).
https = ["clawhdf5-remote/https"]
# s3://, gs://, az:// URLs (object_store; its cloud clients build aws-lc-rs, C).
s3 = ["clawhdf5-remote/s3"]
gcs = ["clawhdf5-remote/gcs"]
azure = ["clawhdf5-remote/azure"]
[package.metadata.docs.rs]
features = []
+74 -3
View File
@@ -43,8 +43,9 @@ with clawhdf5.File("data.h5", "r") as f:
Other types raise `TypeError`.
- Keys are h5py's: integers, slices with a positive step, `...`, one
increasing list of integers, compound field names. Each maps onto a
hyperslab selection. `None`, negative steps and boolean masks are refused
with h5py's errors.
hyperslab selection. `None` and negative steps are refused
with h5py's errors; boolean masks (which h5py supports) raise
`NotImplementedError`, for reads and writes.
- What is read from the file: a selection whose bounding box covers at
most half the dataset decodes only the chunks (or contiguous rows) the box
overlaps. The library decodes the whole dataset for a larger box
@@ -61,6 +62,36 @@ with clawhdf5.File("data.h5", "r") as f:
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
dataspace (h5py's `Empty`).
## Remote files
A URL instead of a path reads the file where it is, through
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
```python
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
```
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise `OSError` instead of mixing versions.
Network failures are `OSError` too.
- Remote files are read-only.
- Schemes: the default build (no C) reads `http://`. `https://` needs
`maturin develop --release --features https` (rustls with ring, which
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
`azure` features (credentials from the environment; aws-lc-rs, C).
## Writing
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
@@ -68,6 +99,41 @@ chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
written on `close()`.
## Editing a file in place
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
where it is, through clawhdf5's `FileEditor`; the file is locked until
`close()`, and every edit is written and synced before the statement
returns.
```python
with clawhdf5.File("data.h5", "r+") as f:
ds = f["grid"]
ds[10:20, ::2] = 0 # h5py keys and broadcasting
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
f["series"].attrs["units"] = "K"
f.attrs.create("version", 2, dtype="u1")
```
- Values: a numpy array is converted to the dataset's dtype as libhdf5
converts it (integers saturate at the target's limits; floats are
truncated toward zero and clipped); anything else goes through
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
value). A few libhdf5 edge cases differ on purpose; see
`docs/known-issues.md`.
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
see its result.
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
`str` is stored as a fixed-length UTF-8 string (h5py stores a
variable-length one), so h5py reads it back as `bytes`.
- Not supported (`NotImplementedError`, nothing written): creating or
deleting datasets, groups and attributes, writing compound fields by
name, variable-length data, HDF5 array types, and whatever
`FileEditor` refuses (listed in `docs/known-issues.md`).
## Tests
```bash
@@ -76,7 +142,12 @@ pytest crates/clawhdf5-py/tests
```
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
writes. `scripts/ci-test.sh` builds the wheel and runs these in CI.
writes, opened locally and over HTTP (an in-process range server,
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
the wheel and runs these in CI.
## License
+166 -65
View File
@@ -1,22 +1,62 @@
//! PyAttrs — dict-like access to HDF5 attributes.
use std::sync::{Arc, Mutex};
use std::sync::{Arc, Mutex, PoisonError};
use clawhdf5_format::attribute::AttributeMessage;
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
use pyo3::exceptions::{PyKeyError, PyNotImplementedError, PyTypeError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple};
use crate::convert::{Converter, Elements, resolve_vl};
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value};
use crate::handle::Handle;
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, edit, node, py_to_attr_value};
/// The attributes of an object in a file opened for reading (or editing).
struct ReadAttrs {
handle: Arc<Handle>,
addr: u64,
path: String,
/// Sorted by name, with the file generation they were read at: an edit
/// (`attrs[name] = value`, here or through another handle on the same
/// object) makes them re-read.
cache: Mutex<(u64, Arc<Vec<AttributeMessage>>)>,
}
impl ReadAttrs {
fn current(&self, py: Python<'_>) -> PyResult<Arc<Vec<AttributeMessage>>> {
let generation = self.handle.generation();
{
let cached = self.cache.lock().unwrap_or_else(PoisonError::into_inner);
if cached.0 == generation {
return Ok(Arc::clone(&cached.1));
}
}
let (addr, path) = (self.addr, &self.path);
let attrs = Arc::new(self.handle.with(py, |f| node::attributes(f, addr, path))?);
*self.cache.lock().unwrap_or_else(PoisonError::into_inner) =
(generation, Arc::clone(&attrs));
Ok(attrs)
}
fn check_writable(&self) -> PyResult<()> {
if self.handle.is_writable() {
return Ok(());
}
Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"cannot set attributes on a read-only file (open it with mode 'r+')",
))
}
fn set(&self, py: Python<'_>, name: &str, value: clawhdf5_rs::AttrValue) -> PyResult<()> {
self.check_writable()?;
let path = node::name(&self.path);
self.handle.edit(py, |ed| ed.set_attr(&path, name, &value))
}
}
/// Backing storage for attributes.
enum AttrsInner {
/// Attributes of an object in a file opened for reading, sorted by name.
Read {
file: Arc<clawhdf5_rs::File>,
attrs: Vec<AttributeMessage>,
},
Read(ReadAttrs),
/// Writable attribute list shared with a parent (PyFile or PyGroup).
Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>),
}
@@ -26,8 +66,11 @@ enum AttrsInner {
/// In read mode, values are what h5py returns: numpy scalars for scalar
/// attributes, numpy arrays otherwise, `str` for variable-length strings,
/// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace.
/// In write mode, attributes set here are accumulated and written when
/// the parent file is closed.
/// In a file opened with `'r+'`, `attrs[name] = value` adds or replaces an
/// attribute in the file at once (as h5py stores it, except that `str`
/// values become fixed-length UTF-8 strings). In write mode (`'w'`),
/// attributes set here are accumulated and written when the parent file is
/// closed.
#[pyclass(name = "Attrs")]
pub struct PyAttrs {
inner: AttrsInner,
@@ -36,10 +79,21 @@ pub struct PyAttrs {
impl PyAttrs {
/// The attributes of the object at `addr` (whose path is `path`) in a
/// file opened for reading.
pub(crate) fn read(file: Arc<clawhdf5_rs::File>, addr: u64, path: &str) -> PyResult<Self> {
let attrs = node::attributes(&file, addr, path)?;
pub(crate) fn read(
py: Python<'_>,
handle: Arc<Handle>,
addr: u64,
path: &str,
) -> PyResult<Self> {
let generation = handle.generation();
let attrs = Arc::new(handle.with(py, |f| node::attributes(f, addr, path))?);
Ok(Self {
inner: AttrsInner::Read { file, attrs },
inner: AttrsInner::Read(ReadAttrs {
handle,
addr,
path: path.to_string(),
cache: Mutex::new((generation, attrs)),
}),
})
}
@@ -49,14 +103,48 @@ impl PyAttrs {
inner: AttrsInner::Write(store),
}
}
fn set_value(
&self,
py: Python<'_>,
key: &str,
value: &Bound<'_, PyAny>,
dtype: Option<&Bound<'_, PyAny>>,
shape: Option<&Bound<'_, PyAny>>,
) -> PyResult<()> {
match &self.inner {
AttrsInner::Read(r) => {
r.check_writable()?;
let value = edit::attr_value(py, value, dtype, shape)?;
r.set(py, key, value)
}
AttrsInner::Write(store) => {
if dtype.is_some() || shape.is_some() {
return Err(PyNotImplementedError::new_err(
"attrs.create with a dtype or shape is only supported in a file opened \
with 'r+'",
));
}
let owned = py_to_attr_value(value)?;
let mut guard = store.lock().unwrap();
// Replace existing key if present.
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
entry.1 = owned;
} else {
guard.push((key.to_string(), owned));
}
Ok(())
}
}
}
}
#[pymethods]
impl PyAttrs {
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
match &self.inner {
AttrsInner::Read { file, attrs } => match attrs.iter().find(|a| a.name == key) {
Some(attr) => Ok(attr_to_py(py, file, attr)?.unbind()),
AttrsInner::Read(r) => match r.current(py)?.iter().find(|a| a.name == key) {
Some(attr) => Ok(attr_to_py(py, &r.handle, attr)?.unbind()),
None => Err(PyKeyError::new_err(format!(
"Can't open attribute (can't locate attribute: '{key}')"
))),
@@ -74,36 +162,63 @@ impl PyAttrs {
}
}
fn __setitem__(&self, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
/// `attrs[name] = value`. In a file opened with `'r+'` this writes the
/// attribute (numeric, bool, complex, bytes and str data, any shape)
/// into the file before returning; see the class docs.
fn __setitem__(&self, py: Python<'_>, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
self.set_value(py, key, value, None, None)
}
/// Deleting attributes is not supported in a file (the in-place editor
/// cannot remove them); in write mode it removes a pending attribute.
fn __delitem__(&self, key: &str) -> PyResult<()> {
match &self.inner {
AttrsInner::Read { .. } => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"cannot set attributes on a read-only file",
)),
AttrsInner::Read(_) => Err(PyNotImplementedError::new_err(format!(
"cannot delete attribute '{key}': deleting attributes is not supported by \
clawhdf5's in-place editor"
))),
AttrsInner::Write(store) => {
let owned = py_to_attr_value(value)?;
let mut guard = store.lock().unwrap();
// Replace existing key if present.
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
entry.1 = owned;
} else {
guard.push((key.to_string(), owned));
let before = guard.len();
guard.retain(|(k, _)| k != key);
if guard.len() == before {
return Err(PyKeyError::new_err(key.to_string()));
}
Ok(())
}
}
}
fn __len__(&self) -> usize {
/// h5py's `attrs.create(name, data, shape=None, dtype=None)`: `data`
/// converted to `dtype` and reshaped to `shape` first.
#[pyo3(signature = (name, data, shape=None, dtype=None))]
fn create(
&self,
py: Python<'_>,
name: &str,
data: &Bound<'_, PyAny>,
shape: Option<&Bound<'_, PyAny>>,
dtype: Option<&Bound<'_, PyAny>>,
) -> PyResult<()> {
self.set_value(py, name, data, dtype, shape)
}
/// h5py's `attrs.modify(name, value)`: same as `attrs[name] = value`.
fn modify(&self, py: Python<'_>, name: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
self.set_value(py, name, value, None, None)
}
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match &self.inner {
AttrsInner::Read { attrs, .. } => attrs.len(),
AttrsInner::Write(store) => store.lock().unwrap().len(),
AttrsInner::Read(r) => Ok(r.current(py)?.len()),
AttrsInner::Write(store) => Ok(store.lock().unwrap().len()),
}
}
fn __contains__(&self, key: &str) -> bool {
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
match &self.inner {
AttrsInner::Read { attrs, .. } => attrs.iter().any(|a| a.name == key),
AttrsInner::Write(store) => store.lock().unwrap().iter().any(|(k, _)| k == key),
AttrsInner::Read(r) => Ok(r.current(py)?.iter().any(|a| a.name == key)),
AttrsInner::Write(store) => Ok(store.lock().unwrap().iter().any(|(k, _)| k == key)),
}
}
@@ -113,15 +228,17 @@ impl PyAttrs {
Ok(iter)
}
fn __repr__(&self) -> String {
let n = self.__len__();
format!("<HDF5 Attrs ({n} members)>")
fn __repr__(&self, py: Python<'_>) -> String {
match self.__len__(py) {
Ok(n) => format!("<HDF5 Attrs ({n} members)>"),
Err(_) => "<HDF5 Attrs>".to_string(),
}
}
/// The value of `key`, or `default` if there is no such attribute.
#[pyo3(signature = (key, default=None))]
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
if self.__contains__(key) {
if self.__contains__(py, key)? {
self.__getitem__(py, key)
} else {
Ok(default.unwrap_or_else(|| py.None()))
@@ -131,7 +248,7 @@ impl PyAttrs {
/// Return attribute names as a list.
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let names: Vec<String> = match &self.inner {
AttrsInner::Read { attrs, .. } => attrs.iter().map(|a| a.name.clone()).collect(),
AttrsInner::Read(r) => r.current(py)?.iter().map(|a| a.name.clone()).collect(),
AttrsInner::Write(store) => store
.lock()
.unwrap()
@@ -146,9 +263,10 @@ impl PyAttrs {
/// Return attribute values as a list.
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let vals: Vec<Py<PyAny>> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs
AttrsInner::Read(r) => r
.current(py)?
.iter()
.map(|a| attr_to_py(py, file, a).map(Bound::unbind))
.map(|a| attr_to_py(py, &r.handle, a).map(Bound::unbind))
.collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store
.lock()
@@ -167,9 +285,10 @@ impl PyAttrs {
/// Return attribute (key, value) pairs as a list of tuples.
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
AttrsInner::Read { file, attrs } => attrs
AttrsInner::Read(r) => r
.current(py)?
.iter()
.map(|a| Ok((a.name.clone(), attr_to_py(py, file, a)?.unbind())))
.map(|a| Ok((a.name.clone(), attr_to_py(py, &r.handle, a)?.unbind())))
.collect::<PyResult<_>>()?,
AttrsInner::Write(store) => store
.lock()
@@ -189,12 +308,11 @@ impl PyAttrs {
/// An attribute's value as h5py returns it.
fn attr_to_py<'py>(
py: Python<'py>,
file: &clawhdf5_rs::File,
handle: &Handle,
attr: &AttributeMessage,
) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| {
let sb = file.superblock();
let conv = Converter::new(py, &attr.datatype, sb.offset_size)
let conv = Converter::new(py, &attr.datatype, handle.offset_size)
.map_err(|e| prefix_err(py, &attr.name, e))?;
if node::is_null(&attr.dataspace) {
return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any());
@@ -216,12 +334,11 @@ fn attr_to_py<'py>(
)));
}
let raw = &attr.raw_data[..want];
let file_data = file.as_bytes();
let (osz, lsz, unit) = (sb.offset_size, sb.length_size, conv.vl_unit);
Elements::Vl(
py.detach(|| resolve_vl(file_data, raw, n, osz, lsz, unit))
.map_err(|e| PyValueError::new_err(format!("attribute {}: {e}", attr.name)))?,
)
let (osz, lsz, unit) = (handle.offset_size, handle.length_size, conv.vl_unit);
let what = format!("attribute {}", attr.name);
Elements::Vl(handle.with(py, |f| {
resolve_vl(f.storage(), raw, n, osz, lsz, unit).map_err(|e| e.into_py(&what))
})?)
} else {
Elements::Bytes(attr.raw_data.clone())
};
@@ -244,19 +361,3 @@ fn prefix_err(py: Python<'_>, name: &str, e: PyErr) -> PyErr {
PyValueError::new_err(msg)
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn write_attrs_len() {
let store = Arc::new(Mutex::new(Vec::new()));
store
.lock()
.unwrap()
.push(("key".into(), OwnedAttrValue::I64(99)));
let attrs = PyAttrs::from_write(store);
assert_eq!(attrs.__len__(), 1);
}
}
+49 -21
View File
@@ -17,6 +17,7 @@ use std::collections::HashMap;
use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder};
use clawhdf5_format::global_heap::GlobalHeapCollection;
use clawhdf5_format::storage::Storage;
use numpy::PyArray1;
use pyo3::exceptions::{PyTypeError, PyValueError};
use pyo3::prelude::*;
@@ -538,16 +539,16 @@ fn object_array<'py>(
/// their bytes: each element's stored length times `unit` (1 for strings,
/// the base type's size for sequences). Pure Rust, so it runs without the
/// GIL.
pub(crate) fn resolve_vl(
file_data: &[u8],
pub(crate) fn resolve_vl<S: Storage + ?Sized>(
file: &S,
raw: &[u8],
count: usize,
offset_size: u8,
length_size: u8,
unit: usize,
) -> Result<Vec<Vec<u8>>, String> {
) -> Result<Vec<Vec<u8>>, VlError> {
let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size)
.map_err(|e| e.to_string())?;
.map_err(|e| VlError::Invalid(e.to_string()))?;
let undefined = match offset_size {
2 => 0xFFFF,
4 => 0xFFFF_FFFF,
@@ -558,43 +559,70 @@ pub(crate) fn resolve_vl(
for vl in &refs {
if vl.collection_address == 0 || vl.collection_address == undefined {
if vl.length != 0 {
return Err(format!(
return Err(VlError::Invalid(format!(
"variable-length element of length {} has no heap address",
vl.length
));
)));
}
out.push(Vec::new());
continue;
}
let coll = match collections.entry(vl.collection_address) {
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
std::collections::hash_map::Entry::Vacant(e) => {
let addr = usize::try_from(vl.collection_address)
.map_err(|_| "global heap address out of range".to_string())?;
e.insert(
GlobalHeapCollection::parse(file_data, addr, length_size)
.map_err(|e| e.to_string())?,
)
}
std::collections::hash_map::Entry::Vacant(e) => e.insert(
GlobalHeapCollection::parse_in(file, vl.collection_address, length_size)
.map_err(VlError::from_format)?,
),
};
let index = u16::try_from(vl.object_index)
.map_err(|_| format!("global heap object index {} out of range", vl.object_index))?;
let index = u16::try_from(vl.object_index).map_err(|_| {
VlError::Invalid(format!(
"global heap object index {} out of range",
vl.object_index
))
})?;
let obj = coll.get_object(index).ok_or_else(|| {
format!(
VlError::Invalid(format!(
"global heap object {index} not found in the collection at {}",
vl.collection_address
)
))
})?;
let need = (vl.length as usize)
.checked_mul(unit)
.ok_or("variable-length element too long")?;
.ok_or_else(|| VlError::Invalid("variable-length element too long".into()))?;
if need > obj.data.len() {
return Err(format!(
return Err(VlError::Invalid(format!(
"variable-length element of {need} bytes in a {}-byte heap object",
obj.data.len()
));
)));
}
out.push(obj.data[..need].to_vec());
}
Ok(out)
}
/// Why variable-length elements could not be resolved.
#[derive(Debug)]
pub(crate) enum VlError {
/// Reading the file failed (a network error on a remote file).
Storage(String),
/// The references or the heap are not valid.
Invalid(String),
}
impl VlError {
fn from_format(e: clawhdf5_format::error::FormatError) -> Self {
match e {
clawhdf5_format::error::FormatError::Storage(_) => VlError::Storage(e.to_string()),
e => VlError::Invalid(e.to_string()),
}
}
/// As a Python exception, the message prefixed with `what`: a storage
/// failure is an `OSError`, anything else a `ValueError`.
pub(crate) fn into_py(self, what: &str) -> PyErr {
match self {
VlError::Storage(m) => pyo3::exceptions::PyOSError::new_err(format!("{what}: {m}")),
VlError::Invalid(m) => PyValueError::new_err(format!("{what}: {m}")),
}
}
}
+256 -68
View File
@@ -6,20 +6,52 @@
//! whole dataset instead); the
//! bytes it returns become the numpy array's buffer without a copy (see
//! `convert`). All file access and decoding runs with the GIL released, so
//! Python threads reading the same or different datasets run in parallel.
//! Python threads reading the same or different datasets run in parallel,
//! and a remote file's network reads never hold the GIL.
use std::sync::Arc;
use std::sync::{Arc, Mutex, PoisonError};
use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::object_header::ObjectHeader;
use pyo3::exceptions::{PyTypeError, PyValueError};
use clawhdf5_rs::File;
use pyo3::exceptions::{PyNotImplementedError, PyOSError, PyTypeError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::{PyList, PyTuple};
use crate::attrs::PyAttrs;
use crate::convert::{Converter, Elements, resolve_vl};
use crate::convert::{Converter, Elements, VlError, resolve_vl};
use crate::handle::Handle;
use crate::select::{self, Plan};
use crate::{PyEmpty, node, to_py_err};
use crate::{PyEmpty, edit, node, to_py_err};
/// What opening a dataset reads from the file (without the GIL).
pub(crate) struct DatasetMeta {
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
shape: Option<Vec<u64>>,
chunks: Option<Vec<u64>>,
datatype: Datatype,
}
impl DatasetMeta {
pub(crate) fn load(f: &File, addr: u64, hdr: &ObjectHeader, path: &str) -> PyResult<Self> {
let null = node::is_null(&node::dataspace(f, hdr, path)?);
let ds = f.dataset_at(addr).map_err(to_py_err)?;
let shape = if null {
None
} else {
Some(ds.shape().map_err(to_py_err)?)
};
let datatype = ds.raw_datatype().map_err(to_py_err)?;
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(f, hdr, s.len()));
Ok(Self {
shape,
chunks,
datatype,
})
}
}
/// A dataset in a file opened for reading.
///
@@ -30,13 +62,14 @@ use crate::{PyEmpty, node, to_py_err};
/// ```
#[pyclass(name = "Dataset")]
pub struct PyDataset {
file: Arc<clawhdf5_rs::File>,
handle: Arc<Handle>,
path: String,
/// Where the dataset's object header is: reads open it from here rather
/// than resolve `path` again.
addr: u64,
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
shape: Option<Vec<u64>>,
/// The shape (`None` for a null dataspace, h5py's `Empty`), with the
/// file generation it was read at: an edit (a resize) may change it.
shape: Mutex<(u64, Option<Vec<u64>>)>,
/// The chunk shape, for a chunked dataset.
chunks: Option<Vec<u64>>,
datatype: Datatype,
@@ -45,39 +78,25 @@ pub struct PyDataset {
}
impl PyDataset {
pub(crate) fn open(
pub(crate) fn new(
py: Python<'_>,
file: Arc<clawhdf5_rs::File>,
handle: Arc<Handle>,
path: String,
addr: u64,
hdr: &ObjectHeader,
) -> PyResult<Self> {
crate::no_panic(|| {
let null = node::is_null(&node::dataspace(&file, hdr)?);
let (shape, datatype) = {
let ds = file.dataset_at(addr).map_err(to_py_err)?;
let shape = if null {
None
} else {
Some(ds.shape().map_err(to_py_err)?)
};
(shape, ds.raw_datatype().map_err(to_py_err)?)
};
let conv = Converter::new(py, &datatype, file.superblock().offset_size)
meta: DatasetMeta,
) -> Self {
let generation = handle.generation();
let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
.map_err(|e| e.value(py).to_string());
let chunks = shape
.as_ref()
.and_then(|s| node::chunk_shape(&file, hdr, s.len()));
Ok(Self {
file,
Self {
handle,
path,
addr,
shape,
chunks,
datatype,
shape: Mutex::new((generation, meta.shape)),
chunks: meta.chunks,
datatype: meta.datatype,
conv,
})
})
}
}
fn converter(&self) -> PyResult<&Converter> {
@@ -86,10 +105,54 @@ impl PyDataset {
.map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path))))
}
/// The current shape: the one read at open, or re-read after an edit.
fn dims(&self, py: Python<'_>) -> PyResult<Option<Vec<u64>>> {
let generation = self.handle.generation();
{
let cached = self.shape.lock().unwrap_or_else(PoisonError::into_inner);
if cached.0 == generation {
return Ok(cached.1.clone());
}
}
let addr = self.addr;
let null = self
.shape
.lock()
.unwrap_or_else(PoisonError::into_inner)
.1
.is_none();
let shape = if null {
None
} else {
Some(self.handle.with(py, |f| {
f.dataset_at(addr)
.and_then(|ds| ds.shape())
.map_err(to_py_err)
})?)
};
*self.shape.lock().unwrap_or_else(PoisonError::into_inner) = (generation, shape.clone());
Ok(shape)
}
fn check_writable(&self) -> PyResult<()> {
if self.handle.is_writable() {
Ok(())
} else {
Err(PyOSError::new_err(format!(
"{}: the file is open read-only; open it with mode 'r+' to change it",
node::name(&self.path)
)))
}
}
/// Read the selection described by `plan` into a numpy array.
fn read_plan<'py>(&self, py: Python<'py>, plan: &Plan) -> PyResult<Bound<'py, PyAny>> {
fn read_plan<'py>(
&self,
py: Python<'py>,
plan: &Plan,
dims: &[u64],
) -> PyResult<Bound<'py, PyAny>> {
let conv = self.converter()?;
let dims = self.shape.as_deref().unwrap_or(&[]);
let out_shape = plan.out_shape();
let arr = if plan.is_empty() {
@@ -103,10 +166,10 @@ impl PyDataset {
};
let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size);
let read_shape = plan.read_shape();
let file = &*self.file;
let handle = &*self.handle;
let addr = self.addr;
// Everything below touches only Rust data: release the GIL.
let read = || -> Result<Elements, ReadError> {
let read = |file: &File| -> Result<Elements, ReadError> {
let ds = file.dataset_at(addr)?;
let mut blocks = Vec::with_capacity(reads.len());
for read in reads {
@@ -147,7 +210,7 @@ impl PyDataset {
let sb = file.superblock();
let n = read_shape.iter().product();
resolve_vl(
file.as_bytes(),
file.storage(),
&raw,
n,
sb.offset_size,
@@ -155,12 +218,20 @@ impl PyDataset {
unit,
)
.map(Elements::Vl)
.map_err(ReadError::Other)
.map_err(ReadError::Vl)
};
let data = py
.detach(|| {
std::panic::catch_unwind(std::panic::AssertUnwindSafe(read))
.unwrap_or_else(|p| Err(ReadError::Panic(crate::panic_text(&*p))))
handle
.with_detached(|f| {
Ok(
std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| read(f)))
.unwrap_or_else(|p| {
Err(ReadError::Panic(crate::panic_text(&*p)))
}),
)
})
.unwrap_or_else(|e| Err(ReadError::Py(e)))
})
.map_err(|e| e.into_py(&self.path))?;
let joined = conv.to_array(py, data, &read_shape, false)?;
@@ -183,6 +254,8 @@ impl PyDataset {
/// An error from the read closure, turned into a Python error with the GIL.
enum ReadError {
Lib(clawhdf5_rs::Error),
Vl(VlError),
Py(PyErr),
Other(String),
Panic(String),
}
@@ -197,6 +270,8 @@ impl ReadError {
fn into_py(self, path: &str) -> PyErr {
match self {
ReadError::Lib(e) => to_py_err(e),
ReadError::Vl(e) => e.into_py(&node::name(path)),
ReadError::Py(e) => e,
ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))),
ReadError::Panic(msg) => crate::InternalError::new_err(format!(
"{}: clawhdf5 internal error (please report it): {msg}",
@@ -243,7 +318,7 @@ impl PyDataset {
/// The shape of the dataset (`None` for an empty/null dataspace).
#[getter]
fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
match &self.shape {
match self.dims(py)? {
Some(s) => Ok(PyTuple::new(py, s)?.into_any()),
None => Ok(py.None().into_bound(py)),
}
@@ -252,22 +327,32 @@ impl PyDataset {
/// The maximum shape (`None` per unlimited dimension), like h5py.
#[getter]
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
crate::no_panic(|| {
let Some(shape) = &self.shape else {
let Some(shape) = self.dims(py)? else {
return Ok(py.None().into_bound(py));
};
let addr = self.addr;
let max = self
.file
.dataset_at(self.addr)
.handle
.with(py, |f| {
f.dataset_at(addr)
.and_then(|ds| ds.max_dimensions())
.map_err(to_py_err)?
.unwrap_or_else(|| shape.clone());
.map_err(to_py_err)
})?
.unwrap_or(shape);
let items: Vec<Option<u64>> = max
.into_iter()
.map(|d| (d != u64::MAX).then_some(d))
.collect();
Ok(PyTuple::new(py, items)?.into_any())
})
}
/// The chunk shape, or `None` for a dataset that is not chunked.
#[getter]
fn chunks<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
match &self.chunks {
Some(c) => Ok(PyTuple::new(py, c)?.into_any()),
None => Ok(py.None().into_bound(py)),
}
}
/// The dataset's numpy dtype, as h5py reports it.
@@ -277,14 +362,14 @@ impl PyDataset {
}
#[getter]
fn ndim(&self) -> usize {
self.shape.as_ref().map_or(0, Vec::len)
fn ndim(&self, py: Python<'_>) -> PyResult<usize> {
Ok(self.dims(py)?.map_or(0, |s| s.len()))
}
/// Number of elements (`None` for an empty/null dataspace, as h5py).
#[getter]
fn size(&self) -> Option<u64> {
self.shape.as_ref().map(|s| s.iter().product())
fn size(&self, py: Python<'_>) -> PyResult<Option<u64>> {
Ok(self.dims(py)?.map(|s| s.iter().product()))
}
/// The dataset's full name, e.g. `/group/data`.
@@ -293,10 +378,11 @@ impl PyDataset {
node::name(&self.path)
}
/// The dataset's attributes (read-only, dict-like).
/// The dataset's attributes (dict-like; writable in a file opened with
/// `'r+'`).
#[getter]
fn attrs(&self) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
}
/// Read with h5py indexing: integers, slices with positive steps,
@@ -308,7 +394,7 @@ impl PyDataset {
py: Python<'py>,
key: &Bound<'py, PyAny>,
) -> PyResult<Bound<'py, PyAny>> {
let Some(dims) = &self.shape else {
let Some(dims) = self.dims(py)? else {
let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty());
let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>();
if is_empty_tuple || is_ellipsis {
@@ -317,8 +403,109 @@ impl PyDataset {
}
return Err(PyValueError::new_err("Empty datasets cannot be sliced"));
};
let plan = select::parse(key, dims)?;
self.read_plan(py, &plan)
let plan = select::parse(key, &dims)?;
self.read_plan(py, &plan, &dims)
}
/// Write with h5py indexing (file opened with `'r+'`): `ds[key] = value`.
///
/// The key is what `ds[key]` reads (without compound field names). The
/// value is converted to the dataset's dtype as h5py converts it (a
/// numpy array as libhdf5 does, clipping out-of-range numbers; anything
/// else through `numpy.asarray(value, dtype=ds.dtype)`), and broadcast
/// to the selection as h5py broadcasts. The edit is written and synced
/// before this returns; what the in-place editor cannot write raises
/// `NotImplementedError` and leaves the file as it was.
fn __setitem__(
&self,
py: Python<'_>,
key: &Bound<'_, PyAny>,
value: &Bound<'_, PyAny>,
) -> PyResult<()> {
self.check_writable()?;
let Some(dims) = self.dims(py)? else {
return Err(PyNotImplementedError::new_err(
"writing to an empty (null dataspace) dataset is not supported",
));
};
let plan = select::parse(key, &dims)?;
if !plan.fields.is_empty() {
return Err(PyNotImplementedError::new_err(
"writing compound fields by name is not supported by clawhdf5's in-place editor; \
write whole elements",
));
}
let category = edit::category(&self.datatype)?;
let conv = self.converter()?;
let bytes = edit::dataset_bytes(
py,
value,
conv.dtype.bind(py),
category,
&plan,
self.chunks.as_deref(),
)?;
if plan.is_empty() {
return Ok(());
}
let sel = edit::selection(&plan, &dims)?;
let path = node::name(&self.path);
self.handle
.edit(py, |ed| ed.write_selection(&path, &sel, &bytes))
}
/// Change the dataset's shape (file opened with `'r+'`), as h5py's
/// `Dataset.resize`: `ds.resize((100, 20))`, or `ds.resize(100, axis=0)`.
/// Only chunked datasets, within their maximum shape; new elements read
/// as the fill value.
#[pyo3(signature = (size, axis=None))]
fn resize(&self, py: Python<'_>, size: &Bound<'_, PyAny>, axis: Option<isize>) -> PyResult<()> {
self.check_writable()?;
let Some(dims) = self.dims(py)? else {
return Err(PyTypeError::new_err("Empty datasets cannot be resized"));
};
if self.chunks.is_none() {
return Err(PyTypeError::new_err("Only chunked datasets can be resized"));
}
let shape: Vec<u64> = match axis {
Some(axis) => {
let rank = dims.len();
let a = usize::try_from(axis)
.ok()
.filter(|&a| a < rank)
.ok_or_else(|| {
PyValueError::new_err(format!(
"Invalid axis (0 to {} allowed)",
rank.saturating_sub(1)
))
})?;
let n: u64 = size.extract().map_err(|_| {
PyTypeError::new_err("Argument must be a single int if axis is specified")
})?;
let mut s = dims.clone();
s[a] = n;
s
}
// As h5py: without `axis` the size is a sequence (`tuple(size)`).
None => size.extract().map_err(|_| {
PyTypeError::new_err(format!(
"'{}' object is not iterable",
size.get_type()
.name()
.map(|n| n.to_string())
.unwrap_or_default()
))
})?,
};
if shape.len() != dims.len() {
return Err(PyValueError::new_err(format!(
"new shape {shape:?} has {} dimensions, the dataset {}",
shape.len(),
dims.len()
)));
}
let path = node::name(&self.path);
self.handle.edit(py, |ed| ed.resize(&path, &shape))
}
/// `numpy.asarray(ds)` reads the whole dataset.
@@ -330,20 +517,20 @@ impl PyDataset {
copy: Option<bool>,
) -> PyResult<Bound<'py, PyAny>> {
let _ = copy; // every read is a fresh array
let Some(dims) = &self.shape else {
let Some(dims) = self.dims(py)? else {
return Err(PyValueError::new_err("an empty dataset has no array value"));
};
let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any();
let plan = select::parse(&ellipsis, dims)?;
let arr = self.read_plan(py, &plan)?;
let plan = select::parse(&ellipsis, &dims)?;
let arr = self.read_plan(py, &plan, &dims)?;
match dtype {
Some(dt) => arr.call_method1("astype", (dt,)),
None => Ok(arr),
}
}
fn __len__(&self) -> PyResult<usize> {
match self.shape.as_deref() {
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match self.dims(py)?.as_deref() {
Some([first, ..]) => Ok(*first as usize),
_ => Err(PyTypeError::new_err(
"Attempt to take len() of scalar dataset",
@@ -361,9 +548,10 @@ impl PyDataset {
.unwrap_or_default(),
Err(_) => format!("{:?}", self.datatype),
};
let shape = match &self.shape {
Some(s) => format!("{s:?}"),
None => "None".to_string(),
let shape = match self.dims(py) {
Ok(Some(s)) => format!("{s:?}"),
Ok(None) => "None".to_string(),
Err(_) => "?".to_string(),
};
format!(
"<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">",
+387
View File
@@ -0,0 +1,387 @@
//! In-place editing (`clawhdf5.File(path, 'r+')`) through `FileEditor`:
//! turning what Python assigns into the bytes, selections and attribute
//! values the editor takes.
//!
//! Value conversion follows h5py (see `edit_helpers.py`, run inside the
//! extension module); what `FileEditor` cannot do is `NotImplementedError`
//! before anything is written.
use std::ffi::CString;
use clawhdf5_format::datatype::{
CharacterSet, CompoundMember, Datatype, DatatypeByteOrder, EnumMember, StringPadding,
};
use clawhdf5_format::selection::Selection;
use clawhdf5_rs::AttrValue;
use pyo3::exceptions::{PyNotImplementedError, PyTypeError};
use pyo3::prelude::*;
use pyo3::sync::PyOnceLock;
use pyo3::types::{PyBytes, PyModule, PyTuple};
use crate::select::{Axis, Plan};
/// The helper module, compiled once.
pub(crate) fn helpers(py: Python<'_>) -> PyResult<&Bound<'_, PyModule>> {
static HELPERS: PyOnceLock<Py<PyModule>> = PyOnceLock::new();
let module = HELPERS.get_or_try_init(py, || -> PyResult<Py<PyModule>> {
let code = CString::new(include_str!("edit_helpers.py"))
.map_err(|e| PyTypeError::new_err(e.to_string()))?;
Ok(PyModule::from_code(
py,
&code,
c"clawhdf5/edit_helpers.py",
c"clawhdf5._edit_helpers",
)?
.unbind())
})?;
Ok(module.bind(py))
}
fn not_implemented(what: impl std::fmt::Display) -> PyErr {
PyNotImplementedError::new_err(format!(
"{what} is not supported by clawhdf5's in-place editor"
))
}
/// How values for a dataset of type `dt` are converted (a category of
/// `edit_helpers._convert_array`), or why they cannot be written.
pub(crate) fn category(dt: &Datatype) -> PyResult<&'static str> {
match dt {
Datatype::FixedPoint { .. } => Ok("int"),
Datatype::FloatingPoint { .. } => Ok("float"),
Datatype::Enumeration {
base_type, members, ..
} => {
let is_bool = base_type.type_size() == 1
&& members.len() == 2
&& members
.iter()
.any(|m| m.name == "FALSE" && m.value.first() == Some(&0))
&& members
.iter()
.any(|m| m.name == "TRUE" && m.value.first() == Some(&1));
Ok(if is_bool { "bool" } else { "enum" })
}
Datatype::String {
padding: StringPadding::NullPad,
..
} => Ok("string"),
Datatype::String { padding, .. } => Err(not_implemented(format!(
"writing fixed-length strings padded {padding:?} (libhdf5 converts them \
differently from numpy)"
))),
Datatype::Compound { size, members } => {
if is_complex(*size, members) {
return Ok("complex");
}
check_exact(dt)?;
Ok("exact")
}
Datatype::Opaque { .. } => Ok("exact"),
Datatype::Array { .. } => Err(not_implemented("writing HDF5 array-type elements")),
Datatype::VariableLength { .. } => Err(not_implemented("writing variable-length data")),
Datatype::Reference { .. } => Err(not_implemented("writing references")),
Datatype::BitField { .. } => Err(not_implemented("writing bitfields")),
Datatype::Time { .. } => Err(not_implemented("writing time values")),
}
}
/// h5py's complex numbers: a compound of two identical floats `r`, `i`.
fn is_complex(size: u32, members: &[CompoundMember]) -> bool {
matches!(members, [r, i] if r.name == "r" && i.name == "i"
&& r.datatype == i.datatype
&& matches!(r.datatype, Datatype::FloatingPoint { size: fs, .. }
if r.byte_offset == 0 && i.byte_offset == u64::from(fs) && size == 2 * fs))
}
/// Compound members written byte for byte from the same numpy dtype: fine
/// unless libhdf5 would convert them on the way (strings padded other than
/// with NULs), or the editor cannot write them at all.
fn check_exact(dt: &Datatype) -> PyResult<()> {
match dt {
Datatype::Compound { members, .. } => {
members.iter().try_for_each(|m| check_exact(&m.datatype))
}
Datatype::Array { base_type, .. } => check_exact(base_type),
Datatype::String {
padding: StringPadding::NullPad,
..
}
| Datatype::FixedPoint { .. }
| Datatype::FloatingPoint { .. }
| Datatype::Enumeration { .. }
| Datatype::Opaque { .. }
| Datatype::BitField { .. } => Ok(()),
Datatype::String { .. } => Err(not_implemented(
"writing compounds with strings not padded with NULs",
)),
Datatype::VariableLength { .. } => Err(not_implemented(
"writing compounds with variable-length members",
)),
Datatype::Reference { .. } => Err(not_implemented("writing references")),
Datatype::Time { .. } => Err(not_implemented("writing time values")),
}
}
/// Largest point selection an index-list write builds (one coordinate
/// vector per element).
const MAX_POINTS: usize = 1 << 22;
/// The selection `plan` writes, whose elements are numbered as the value's
/// (row-major over the selection's shape).
pub(crate) fn selection(plan: &Plan, dims: &[u64]) -> PyResult<Selection> {
if plan.axes.is_empty() {
return Ok(Selection::All);
}
if plan.list_axis().is_none() {
let (reads, _) = plan.reads(dims, None, 1);
return match <[_; 1]>::try_from(reads) {
Ok([read]) => Ok(read.sel),
Err(_) => Err(PyTypeError::new_err("internal error: several hyperslabs")),
};
}
// An index list: the points, in the value's order.
let per_axis: Vec<Vec<u64>> = plan
.axes
.iter()
.map(|a| match a {
Axis::Index(i) => vec![*i],
Axis::Slice { start, step, count } => (0..*count).map(|k| start + k * step).collect(),
Axis::List(v) => v.clone(),
})
.collect();
let n = per_axis
.iter()
.try_fold(1usize, |acc, v| acc.checked_mul(v.len()))
.filter(|&n| n <= MAX_POINTS)
.ok_or_else(|| {
not_implemented(format!(
"an index-list write of more than {MAX_POINTS} elements (write it in slices)"
))
})?;
let mut points = Vec::with_capacity(n);
let mut at = vec![0usize; per_axis.len()];
for _ in 0..n {
points.push(at.iter().zip(&per_axis).map(|(&i, v)| v[i]).collect());
for d in (0..at.len()).rev() {
at[d] += 1;
if at[d] < per_axis[d].len() {
break;
}
at[d] = 0;
}
}
Ok(Selection::Points(points))
}
/// The bytes to write for `value` under `plan`, in the dataset's dtype.
pub(crate) fn dataset_bytes(
py: Python<'_>,
value: &Bound<'_, PyAny>,
dtype: &Bound<'_, PyAny>,
category: &str,
plan: &Plan,
chunks: Option<&[u64]>,
) -> PyResult<Vec<u8>> {
let shape = PyTuple::new(py, plan.out_shape())?;
let fancy = plan.list_axis().is_some();
let chunk_elems = chunks.map_or(0, |c| c.iter().fold(1u64, |a, &d| a.saturating_mul(d)));
let bytes = helpers(py)?.call_method1(
"dataset_values",
(value, dtype, category, shape, fancy, chunk_elems),
)?;
Ok(bytes.cast::<PyBytes>()?.as_bytes().to_vec())
}
fn ieee_float(size: u32, byte_order: DatatypeByteOrder) -> Option<Datatype> {
let (exponent_location, exponent_size, mantissa_size, exponent_bias) = match size {
2 => (10, 5, 10, 15),
4 => (23, 8, 23, 127),
8 => (52, 11, 52, 1023),
_ => return None,
};
Some(Datatype::FloatingPoint {
size,
byte_order,
bit_offset: 0,
bit_precision: (size * 8) as u16,
exponent_location,
exponent_size,
mantissa_location: 0,
mantissa_size,
exponent_bias,
})
}
/// The HDF5 datatype h5py writes for a numpy dtype string (`'<i4'`,
/// `'|b1'`, `'>f8'`, `'<c16'`, `'|S5'`).
fn datatype_of(dtype: &str) -> Option<Datatype> {
let order = match dtype.as_bytes().first()? {
b'<' | b'|' | b'=' => DatatypeByteOrder::LittleEndian,
b'>' => DatatypeByteOrder::BigEndian,
_ => return None,
};
let kind = dtype.as_bytes().get(1)?;
let size: u32 = dtype.get(2..)?.parse().ok()?;
match kind {
b'b' if size == 1 => Some(Datatype::Enumeration {
size: 1,
base_type: Box::new(Datatype::FixedPoint {
size: 1,
byte_order: DatatypeByteOrder::LittleEndian,
signed: true,
bit_offset: 0,
bit_precision: 8,
}),
members: vec![
EnumMember {
name: "FALSE".into(),
value: vec![0],
},
EnumMember {
name: "TRUE".into(),
value: vec![1],
},
],
}),
b'i' | b'u' if matches!(size, 1 | 2 | 4 | 8) => Some(Datatype::FixedPoint {
size,
byte_order: order,
signed: *kind == b'i',
bit_offset: 0,
bit_precision: (size * 8) as u16,
}),
b'f' => ieee_float(size, order),
b'c' => {
let part = ieee_float(size / 2, order)?;
Some(Datatype::Compound {
size,
members: vec![
CompoundMember {
name: "r".into(),
byte_offset: 0,
datatype: part.clone(),
},
CompoundMember {
name: "i".into(),
byte_offset: u64::from(size / 2),
datatype: part,
},
],
})
}
b'S' if size > 0 => Some(Datatype::String {
size,
padding: StringPadding::NullPad,
charset: CharacterSet::Ascii,
}),
_ => None,
}
}
/// An attribute value as h5py would store it (`attrs[name] = value`, or
/// `attrs.create(name, data, shape, dtype)`), except that `str` data is
/// stored as fixed-length UTF-8 strings (h5py stores variable-length ones,
/// which the editor cannot write).
pub(crate) fn attr_value(
py: Python<'_>,
value: &Bound<'_, PyAny>,
dtype: Option<&Bound<'_, PyAny>>,
shape: Option<&Bound<'_, PyAny>>,
) -> PyResult<AttrValue> {
if value.is_instance_of::<crate::PyEmpty>() {
return Err(not_implemented(
"writing an empty (null dataspace) attribute",
));
}
let (kind, dt, dims, data): (String, String, Vec<u64>, Vec<u8>) = helpers(py)?
.call_method1("attr_value", (value, dtype, shape))?
.extract()?;
let datatype = if kind == "str" {
let size: u32 = dt
.parse()
.map_err(|_| PyTypeError::new_err("bad string size"))?;
Datatype::String {
size,
padding: StringPadding::NullPad,
charset: CharacterSet::Utf8,
}
} else {
datatype_of(&dt).ok_or_else(|| not_implemented(format!("an attribute of dtype {dt}")))?
};
Ok(AttrValue::Raw {
datatype,
shape: dims,
data,
})
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn numpy_dtypes_map_to_h5py_types() {
assert!(matches!(
datatype_of("<i4"),
Some(Datatype::FixedPoint {
size: 4,
signed: true,
byte_order: DatatypeByteOrder::LittleEndian,
..
})
));
assert!(matches!(
datatype_of(">u2"),
Some(Datatype::FixedPoint {
size: 2,
signed: false,
byte_order: DatatypeByteOrder::BigEndian,
..
})
));
assert!(matches!(
datatype_of("<f2"),
Some(Datatype::FloatingPoint { size: 2, .. })
));
let c = datatype_of("<c16").unwrap();
assert!(matches!(&c, Datatype::Compound { size: 16, members } if is_complex(16, members)));
assert_eq!(category(&c).unwrap(), "complex");
let b = datatype_of("|b1").unwrap();
assert_eq!(category(&b).unwrap(), "bool");
assert!(matches!(
datatype_of("|S5"),
Some(Datatype::String { size: 5, .. })
));
assert!(datatype_of("<f16").is_none());
assert!(datatype_of("<M8").is_none());
assert!(datatype_of("|S0").is_none());
}
#[test]
fn list_writes_become_points_in_value_order() {
let plan = Plan {
axes: vec![
Axis::List(vec![1, 4]),
Axis::Slice {
start: 0,
step: 2,
count: 2,
},
Axis::Index(3),
],
fields: vec![],
scalar: false,
};
let sel = selection(&plan, &[5, 4, 4]).unwrap();
assert_eq!(
sel,
Selection::Points(vec![
vec![1, 0, 3],
vec![1, 2, 3],
vec![4, 0, 3],
vec![4, 2, 3]
])
);
}
}
+175
View File
@@ -0,0 +1,175 @@
"""Values for in-place writes (clawhdf5.File(path, 'r+')), prepared the way
h5py prepares them, so `ds[key] = value` stores what h5py would store.
Loaded by the extension module (src/edit.rs); not a public API.
h5py converts in two ways, and so does this module:
- a value that is not a numpy array (a list, a Python or numpy scalar) is
converted by numpy straight to the dataset's dtype
(`numpy.asarray(value, dtype=ds.dtype)`), with numpy's rules and errors;
- a numpy array is converted by libhdf5, whose numeric conversions clip to
the target's range instead of wrapping: integers saturate, floats are
truncated toward zero and clipped, a double too large for a float becomes
infinity. That is what `_convert_array` reproduces. Where libhdf5 has no
meaningful answer — NaN into an integer, for which it writes a different
arbitrary value per type — this raises ValueError instead of guessing.
"""
import numpy as np
def _no_path(src, dst):
return TypeError(f"No conversion path for dtype: {src!r} -> {dst!r}")
def _to_int(arr, dtype):
"""Integer target: libhdf5's saturating conversion."""
info = np.iinfo(dtype)
kind = arr.dtype.kind
if kind == "b":
return arr.astype(dtype)
if kind in "iu":
src = np.iinfo(arr.dtype)
lo = max(info.min, src.min)
hi = min(info.max, src.max)
clipped = np.clip(arr, np.array(lo, arr.dtype), np.array(hi, arr.dtype))
return clipped.astype(dtype)
if kind == "f":
if np.isnan(arr).any():
raise ValueError(
"cannot write NaN to an integer dataset (libhdf5 would store an arbitrary value)"
)
t = np.trunc(arr.astype(np.float64))
# info.max + 1 and info.min are powers of two: exact as floats.
over = t >= float(info.max + 1)
under = t < float(info.min)
out = np.where(over | under, 0.0, t).astype(dtype)
out[over] = info.max
out[under] = info.min
return out
raise _no_path(arr.dtype, dtype)
def _convert_array(arr, dtype, category):
kind = arr.dtype.kind
if category in ("int", "enum"):
if arr.dtype == dtype and kind in "iu":
return arr
if category == "enum" and kind not in "iu":
raise _no_path(arr.dtype, dtype)
return _to_int(arr, dtype)
if category == "bool":
if kind == "b":
return arr.astype(dtype)
if kind in "iu":
# h5py's bool is an enum over int8; libhdf5 converts integers
# into it by value (saturating), not to FALSE/TRUE, so 3 is
# stored as 3. Keep those bytes: a view, not a cast.
return _to_int(arr, np.dtype("i1")).view(dtype)
raise _no_path(arr.dtype, dtype)
if category == "float":
if kind not in "biuf":
raise _no_path(arr.dtype, dtype)
with np.errstate(over="ignore", invalid="ignore"):
return arr.astype(dtype)
if category == "complex":
if kind != "c":
raise _no_path(arr.dtype, dtype)
with np.errstate(over="ignore", invalid="ignore"):
return arr.astype(dtype)
if category == "string":
if kind != "S":
raise _no_path(arr.dtype, dtype)
return arr.astype(dtype)
# "exact": compound and opaque types, written only from the same dtype.
if arr.dtype == dtype:
return arr
raise _no_path(arr.dtype, dtype)
def _convert_other(value, dtype, category):
if category == "string":
items = np.asarray(value, dtype=object)
if any(isinstance(x, str) for x in items.flat):
meta = dtype.metadata or {}
if meta.get("h5py_encoding") == "utf-8":
enc = [x.encode("utf-8") if isinstance(x, str) else x for x in items.flat]
return np.array(enc, dtype=dtype).reshape(items.shape)
return np.asarray(value, dtype=dtype)
def _broadcast(arr, shape, fancy, chunk_elems):
"""h5py's broadcasting: numpy's rules against the selection's shape
(extra leading length-1 axes allowed) for slices and integers. For an
index list, the exact shape; a scalar only where h5py expands it to the
whole selection (a chunked dataset whose chunk holds at least as many
elements as the selection)."""
if arr.shape == shape:
return arr
if fancy:
size = int(np.prod(shape))
if arr.ndim == 0 and ((chunk_elems > 0 and size <= chunk_elems) or len(shape) == 1):
return np.broadcast_to(arr, shape)
raise TypeError("Broadcasting is not supported for complex selections")
if arr.ndim == 0:
return np.broadcast_to(arr, shape)
err = TypeError(f"Can't broadcast {arr.shape} -> {shape}")
src = arr.shape
while len(src) > len(shape) and src[0] == 1:
src = src[1:]
if len(src) > len(shape):
raise err
try:
return np.broadcast_to(arr.reshape(src), shape)
except ValueError:
raise err from None
def dataset_values(value, dtype, category, shape, fancy, chunk_elems):
"""The bytes to write for `value` under a selection of `shape`, as a
C-ordered array of the dataset's dtype."""
if isinstance(value, np.ndarray):
arr = _convert_array(value, dtype, category)
else:
arr = _convert_other(value, dtype, category)
arr = _broadcast(arr, tuple(shape), fancy, chunk_elems)
return np.ascontiguousarray(arr, dtype=dtype).tobytes()
def attr_value(value, dtype=None, shape=None):
"""(kind, dtype string, shape, bytes) for an attribute value, h5py's
`attrs[name] = value` / `attrs.create(name, data, shape, dtype)`:
- "str": `str` data (h5py would store a variable-length string; this
stores a fixed-length UTF-8 string, which clawhdf5 can write);
the dtype string is the byte length of the longest element;
- "raw": a numeric, bool or bytes array, as numpy lays it out.
"""
if dtype is not None:
arr = np.asarray(value, dtype=dtype, order="C")
else:
arr = np.asarray(value, order="C")
if shape is not None:
arr = arr.reshape(shape)
kind = arr.dtype.kind
if kind == "O":
if arr.size and all(isinstance(x, str) for x in arr.flat):
kind = "U"
elif arr.size and all(isinstance(x, bytes) for x in arr.flat):
arr = arr.astype(bytes)
kind = "S"
else:
raise TypeError(
f"clawhdf5 cannot write an attribute of Python objects ({value!r:.60})"
)
if kind == "U":
enc = [str(x).encode("utf-8") for x in arr.flat]
size = max([len(b) for b in enc] + [1])
data = np.array(enc, dtype=f"S{size}").reshape(arr.shape)
return ("str", str(size), arr.shape, data.tobytes())
if kind in "biufcS":
return ("raw", arr.dtype.str, arr.shape, np.ascontiguousarray(arr).tobytes())
raise NotImplementedError(
f"clawhdf5 cannot write an attribute of dtype {arr.dtype} in place"
)
+228 -32
View File
@@ -1,13 +1,17 @@
//! PyFile — the main entry point for opening and creating HDF5 files.
use std::collections::HashMap;
use std::path::PathBuf;
use std::sync::{Arc, Mutex};
use std::time::Duration;
use pyo3::exceptions::{PyNotImplementedError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::PyList;
use pyo3::types::{PyDict, PyList};
use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
/// Internal state for write mode.
@@ -23,8 +27,10 @@ struct WriteState {
/// Mirrors the h5py.File interface:
///
/// ```python
/// # Reading
/// # Reading, a local file or a URL (range requests, nothing downloaded
/// # up front)
/// f = clawhdf5.File('data.h5', 'r')
/// f = clawhdf5.File('https://example.org/data.h5')
/// ds = f['dataset']
/// f.close()
///
@@ -44,49 +50,200 @@ enum FileInner {
Write(WriteState),
}
/// Whether `s` is a URL (`scheme://…`) rather than a path: the scheme is a
/// letter followed by letters, digits, `+`, `-` or `.` (RFC 3986).
fn is_url(s: &str) -> bool {
let Some((scheme, _)) = s.split_once("://") else {
return false;
};
let mut chars = scheme.chars();
chars.next().is_some_and(|c| c.is_ascii_alphabetic())
&& chars.all(|c| c.is_ascii_alphanumeric() || matches!(c, '+' | '-' | '.'))
}
impl PyFile {
fn from_handle(handle: Arc<Handle>, filename: String) -> Self {
let root = handle.root;
Self {
inner: Some(FileInner::Read(ReadGroup::new(handle, String::new(), root))),
filename,
}
}
}
#[pymethods]
impl PyFile {
/// Open or create an HDF5 file.
///
/// Parameters:
/// path: file path
/// path: file path, or a URL (`http://`, `https://`, `s3://`, `gs://`,
/// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write
#[new]
#[pyo3(signature = (path, mode="r"))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
let filename = path.to_string();
match mode {
"r" => {
let file = py.detach(|| {
crate::no_panic(|| clawhdf5_rs::File::open(path).map_err(to_py_err))
})?;
Ok(Self {
inner: Some(FileInner::Read(root_group(Arc::new(file)))),
filename,
})
if is_url(path) {
if mode != "r" {
return Err(PyValueError::new_err(format!(
"remote files are read-only: mode '{mode}' is not supported for a URL"
)));
}
let handle = Handle::open_url(py, path, &clawhdf5_remote::Options::default())?;
return Ok(Self::from_handle(handle, filename));
}
match mode {
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
"r+" => Ok(Self::from_handle(
Handle::open_editable(py, path)?,
filename,
)),
"a" if std::path::Path::new(path).exists() => Ok(Self::from_handle(
Handle::open_editable(py, path)?,
filename,
)),
"a" => Err(PyNotImplementedError::new_err(format!(
"mode 'a' on {path}, which does not exist: clawhdf5 can only edit an existing \
file in place; create a new one with mode 'w'"
))),
"w" => Ok(Self {
filename,
inner: Some(FileInner::Write(WriteState {
path: PathBuf::from(path),
// Absolute now: the file is written at close, possibly
// after the working directory changed.
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
root_datasets: Vec::new(),
root_attrs: Arc::new(Mutex::new(Vec::new())),
groups: Vec::new(),
})),
}),
other => Err(PyErr::new::<pyo3::exceptions::PyValueError, _>(format!(
"unsupported mode '{other}'; expected 'r' or 'w'"
other => Err(PyValueError::new_err(format!(
"unsupported mode '{other}'; expected 'r', 'r+', 'a' or 'w'"
))),
}
}
/// Open a remote file for reading, with options.
///
/// The file is read through a block cache with range requests: opening
/// costs one request (it also fetches the first block), and a read
/// fetches only the blocks it needs. The GIL is released while waiting
/// on the network.
///
/// Parameters (all optional):
/// block_size: bytes per cached block (default 1 MiB)
/// cache_size: byte budget of the block cache (default 64 MiB)
/// headers: dict of extra HTTP headers (e.g. Authorization), sent only
/// to the URL's own origin
/// retries: retries of a request that failed transiently (default 3)
/// timeout: seconds to connect and receive response headers (default 30)
/// allow_full_download: when the server ignores Range requests,
/// download the whole file once instead of failing (default False)
/// max_full_download: largest file such a download may fetch
/// (default 1 GiB)
/// require_validator: refuse a server that sends neither ETag nor
/// Last-Modified (default False)
/// max_redirects: redirects followed per request (default 5)
/// max_parallel: requests of one read in flight at once (default 8)
#[staticmethod]
#[allow(clippy::too_many_arguments)]
#[pyo3(signature = (url, *, block_size=None, cache_size=None, headers=None, retries=None,
timeout=None, allow_full_download=None, max_full_download=None,
require_validator=None, max_redirects=None, max_parallel=None))]
fn open_url(
py: Python<'_>,
url: &str,
block_size: Option<u64>,
cache_size: Option<u64>,
headers: Option<HashMap<String, String>>,
retries: Option<u32>,
timeout: Option<f64>,
allow_full_download: Option<bool>,
max_full_download: Option<u64>,
require_validator: Option<bool>,
max_redirects: Option<u32>,
max_parallel: Option<usize>,
) -> PyResult<Self> {
let mut options = clawhdf5_remote::Options::default();
if let Some(b) = block_size {
if b == 0 {
return Err(PyValueError::new_err("block_size must be positive"));
}
options.cache.block_size = b;
options.cache.coalesce_gap = b;
// The opening request fetches the first block, not 1 MiB.
options.http.first_request = b;
}
if let Some(c) = cache_size {
options.cache.capacity = c;
}
let http = &mut options.http;
if let Some(h) = headers {
http.headers = h.into_iter().collect();
}
if let Some(r) = retries {
http.retries = r;
}
if let Some(t) = timeout {
if !(t.is_finite() && t > 0.0) {
return Err(PyValueError::new_err("timeout must be a positive number"));
}
http.timeout = Duration::from_secs_f64(t);
}
if let Some(a) = allow_full_download {
http.allow_full_download = a;
}
if let Some(m) = max_full_download {
http.max_full_download = m;
}
if let Some(v) = require_validator {
http.require_validator = v;
}
if let Some(r) = max_redirects {
http.max_redirects = r;
}
if let Some(p) = max_parallel {
if p == 0 {
return Err(PyValueError::new_err("max_parallel must be positive"));
}
http.max_parallel = p;
}
let handle = Handle::open_url(py, url, &options)?;
Ok(Self::from_handle(handle, url.to_string()))
}
/// For a remote file, what its block cache has done so far (reads,
/// hits, misses, requests, bytes fetched, ...); `None` for a local file.
#[getter]
fn remote_stats<'py>(&self, py: Python<'py>) -> PyResult<Option<Bound<'py, PyDict>>> {
let Some(storage) = self.read_file()?.handle.remote_storage() else {
return Ok(None);
};
let s = storage.stats();
let d = PyDict::new(py);
d.set_item("reads", s.reads)?;
d.set_item("hits", s.hits)?;
d.set_item("misses", s.misses)?;
d.set_item("waits", s.waits)?;
d.set_item("requests", s.requests)?;
d.set_item("fetch_calls", s.fetch_calls)?;
d.set_item("bytes_fetched", s.bytes_fetched)?;
d.set_item("evictions", s.evictions)?;
d.set_item("cached_bytes", s.cached_bytes)?;
Ok(Some(d))
}
/// Close the file. In write mode, this finalizes and writes the file.
fn close(&mut self) -> PyResult<()> {
let inner = self.inner.take().ok_or_else(|| {
PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed")
})?;
match inner {
FileInner::Read(_) => Ok(()),
FileInner::Read(root) => {
root.handle.close();
Ok(())
}
FileInner::Write(state) => finalize_write(state),
}
}
@@ -121,7 +278,7 @@ impl PyFile {
/// List the names of all children in the root group.
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
let names = self.read_file()?.member_names()?;
let names = self.read_file()?.member_names(py)?;
Ok(PyList::new(py, names)?.into_any().unbind())
}
@@ -139,8 +296,8 @@ impl PyFile {
self.keys(py)?.call_method0(py, "__iter__")
}
fn __len__(&self) -> PyResult<usize> {
Ok(self.read_file()?.member_names()?.len())
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
Ok(self.read_file()?.member_names(py)?.len())
}
/// The root group's name, `/`.
@@ -149,7 +306,31 @@ impl PyFile {
"/"
}
/// The path the file was opened with.
/// `'r'` for a file opened read-only (a local file or a URL), `'r+'`
/// for one open for editing or writing, as h5py reports it.
#[getter]
fn mode(&self) -> PyResult<&'static str> {
match &self.inner {
Some(FileInner::Read(root)) if !root.handle.is_writable() => Ok("r"),
Some(_) => Ok("r+"),
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"file is closed",
)),
}
}
/// Nothing to do: every edit is written and synced when it is made, and
/// a file opened with 'w' is written on `close()`.
fn flush(&self) {}
/// Deleting objects is not supported (h5py's `del f[name]`).
fn __delitem__(&self, key: &str) -> PyResult<()> {
Err(PyNotImplementedError::new_err(format!(
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
)))
}
/// The path (or URL) the file was opened with.
#[getter]
fn filename(&self) -> &str {
&self.filename
@@ -204,9 +385,9 @@ impl PyFile {
/// Attribute access. In read mode, returns attributes of the root group.
/// In write mode, returns a writable attrs handle.
#[getter]
fn attrs(&self) -> PyResult<PyAttrs> {
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match self.inner.as_ref() {
Some(FileInner::Read(root)) => root.attrs(),
Some(FileInner::Read(root)) => root.attrs(py),
Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))),
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"file is closed",
@@ -216,9 +397,10 @@ impl PyFile {
fn __repr__(&self) -> String {
match &self.inner {
Some(FileInner::Read(root)) => {
format!("<HDF5 File (read, {} bytes)>", root.file.as_bytes().len())
}
Some(FileInner::Read(root)) => match root.handle.redacted_url() {
Some(url) => format!("<HDF5 File (read, \"{url}\")>"),
None => format!("<HDF5 File (read, \"{}\")>", self.filename),
},
Some(FileInner::Write(s)) => {
format!("<HDF5 File (write, \"{}\")>", s.path.display())
}
@@ -226,8 +408,8 @@ impl PyFile {
}
}
fn __contains__(&self, key: &str) -> PyResult<bool> {
Ok(self.read_file()?.contains(key))
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
self.read_file()?.contains(py, key)
}
}
@@ -248,6 +430,13 @@ impl PyFile {
fn write_state_mut(&mut self) -> PyResult<&mut WriteState> {
match &mut self.inner {
Some(FileInner::Write(s)) => Ok(s),
Some(FileInner::Read(root)) if root.handle.is_writable() => {
Err(PyNotImplementedError::new_err(
"creating datasets or groups in an existing file is not supported by \
clawhdf5's in-place editor (mode 'r+' changes values, shapes and \
attributes)",
))
}
Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
"cannot write to a file opened for reading",
)),
@@ -271,11 +460,6 @@ fn parse_compression(
}
}
fn root_group(file: Arc<clawhdf5_rs::File>) -> ReadGroup {
let root = file.superblock().root_group_address;
ReadGroup::new(file, String::new(), root)
}
/// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| {
@@ -309,6 +493,18 @@ fn finalize_write(state: WriteState) -> PyResult<()> {
mod tests {
use super::*;
#[test]
fn urls_and_paths() {
assert!(is_url("http://h/f.h5"));
assert!(is_url("s3://bucket/key.h5"));
assert!(is_url("git+https://x"));
assert!(!is_url("data.h5"));
assert!(!is_url("/tmp/a://b.h5"));
assert!(!is_url("dir/x://y"));
assert!(!is_url("1http://x"));
assert!(!is_url("://x"));
}
#[test]
fn parse_gzip_compression() {
assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6));
+71 -77
View File
@@ -3,11 +3,12 @@
use std::collections::HashMap;
use std::sync::{Arc, Mutex, OnceLock};
use pyo3::exceptions::{PyIOError, PyKeyError, PyValueError};
use pyo3::exceptions::{PyIOError, PyKeyError, PyNotImplementedError, PyOSError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::PyList;
use crate::attrs::PyAttrs;
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
/// Shared state for a group being written.
@@ -34,9 +35,9 @@ enum GroupInner {
}
impl PyGroup {
pub(crate) fn from_read(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
pub(crate) fn from_read(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self {
inner: GroupInner::Read(ReadGroup::new(file, path, addr)),
inner: GroupInner::Read(ReadGroup::new(handle, path, addr)),
}
}
@@ -60,9 +61,10 @@ impl PyGroup {
/// h5py). It keeps its own address and, once listed, its links, so looking
/// up a child neither resolves the path from the root nor scans the group's
/// links again: visiting every member of a large group is linear, not
/// quadratic.
/// quadratic. (Edits never add or remove links, so these stay valid in a
/// file open for editing.)
pub(crate) struct ReadGroup {
pub file: Arc<clawhdf5_rs::File>,
pub handle: Arc<Handle>,
pub path: String,
pub addr: u64,
/// Link name -> object address (soft links resolved), filled on first use.
@@ -72,9 +74,9 @@ pub(crate) struct ReadGroup {
}
impl ReadGroup {
pub(crate) fn new(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
pub(crate) fn new(handle: Arc<Handle>, path: String, addr: u64) -> Self {
Self {
file,
handle,
path,
addr,
links: OnceLock::new(),
@@ -82,17 +84,14 @@ impl ReadGroup {
}
}
fn links(&self) -> PyResult<&HashMap<String, u64>> {
fn links(&self, py: Python<'_>) -> PyResult<&HashMap<String, u64>> {
if let Some(links) = self.links.get() {
return Ok(links);
}
let entries = crate::no_panic(|| {
clawhdf5_format::group_v2::resolve_group_children(
self.file.as_bytes(),
self.file.superblock(),
self.addr,
)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", node::name(&self.path))))
let (addr, path) = (self.addr, &self.path);
let entries = self.handle.with(py, |f| {
clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr)
.map_err(|e| node::format_err(path, e, PyValueError::new_err))
})?;
let map = entries
.into_iter()
@@ -102,7 +101,7 @@ impl ReadGroup {
}
/// The path and address of `key` (a name, a relative or an absolute path).
fn locate(&self, key: &str) -> PyResult<(String, u64)> {
fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> {
let path = node::join(&self.path, key);
let rel = if self.path.is_empty() {
Some(path.as_str())
@@ -112,24 +111,24 @@ impl ReadGroup {
path.strip_prefix(self.path.as_str())
.and_then(|r| r.strip_prefix('/'))
};
let addr = match rel {
// A direct child: the link table, when it has the name.
Some(name) if !name.is_empty() && !name.contains('/') => {
match self.links()?.get(name) {
Some(&a) => a,
None => node::resolve_from(&self.file, self.addr, name, &path)?,
if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/'))
&& let Some(&a) = self.links(py)?.get(name)
{
return Ok((path, a));
}
}
Some(rel) => node::resolve_from(&self.file, self.addr, rel, &path)?,
None => node::address(&self.file, &path)?,
};
Ok((path, addr))
let addr = self.addr;
let found = self.handle.with(py, |f| match rel {
Some(rel) => node::resolve_from(f, addr, rel, &path),
None => node::address(f, &path),
})?;
Ok((path, found))
}
/// `group[key]`.
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
let (path, addr) = self.locate(key)?;
node::open(py, &self.file, path, addr)
let (path, addr) = self.locate(py, key)?;
node::open(py, &self.handle, path, addr)
}
/// `group.get(key, default)`.
@@ -148,48 +147,57 @@ impl ReadGroup {
}
/// Names of the group's datasets and subgroups, sorted (h5py's order).
pub(crate) fn member_names(&self) -> PyResult<&[String]> {
pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> {
if let Some(m) = self.members.get() {
return Ok(m);
}
let links = self.links(py)?;
let path = &self.path;
let mut names = self.handle.with(py, |f| {
let mut names = Vec::new();
for (name, &addr) in self.links()? {
let hdr = node::header_at(&self.file, addr, &node::join(&self.path, name))?;
for (name, &addr) in links {
if matches!(
node::kind(&hdr),
node::kind_at(f, addr, &node::join(path, name))?,
Some(node::Kind::Dataset | node::Kind::Group)
) {
names.push(name.clone());
}
}
Ok(names)
})?;
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
Ok(self.members.get_or_init(|| names))
}
pub(crate) fn contains(&self, key: &str) -> bool {
self.locate(key)
.and_then(|(path, addr)| node::header_at(&self.file, addr, &path))
.ok()
.and_then(|h| node::kind(&h))
.is_some_and(|k| k != node::Kind::Datatype)
/// `key in group`: whether `key` names a dataset or group. A failed
/// read of the file (a network error) is raised, not `False`.
pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
let found = self
.locate(py, key)
.and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path)));
match found {
Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)),
Err(e) if e.is_instance_of::<PyOSError>(py) => Err(e),
Err(_) => Ok(false),
}
}
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
self.member_names()?
self.member_names(py)?
.iter()
.map(|n| self.get_item(py, n))
.collect()
}
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
self.member_names()?
self.member_names(py)?
.iter()
.map(|n| Ok((n.clone(), self.get_item(py, n)?)))
.collect()
}
pub(crate) fn attrs(&self) -> PyResult<PyAttrs> {
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
}
}
@@ -210,7 +218,7 @@ impl PyGroup {
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
match &self.inner {
GroupInner::Read(g) => {
let list = PyList::new(py, g.member_names()?)?;
let list = PyList::new(py, g.member_names(py)?)?;
Ok(list.into_any().unbind())
}
GroupInner::Write(state) => {
@@ -236,9 +244,9 @@ impl PyGroup {
self.keys(py)?.call_method0(py, "__iter__")
}
fn __len__(&self) -> PyResult<usize> {
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
match &self.inner {
GroupInner::Read(g) => Ok(g.member_names()?.len()),
GroupInner::Read(g) => Ok(g.member_names(py)?.len()),
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
}
}
@@ -293,17 +301,28 @@ impl PyGroup {
state.lock().unwrap().datasets.push(spec);
Ok(())
}
GroupInner::Read { .. } => Err(PyIOError::new_err(
GroupInner::Read(g) if g.handle.is_writable() => Err(PyNotImplementedError::new_err(
"creating datasets or groups in an existing file is not supported by \
clawhdf5's in-place editor (mode 'r+' changes values, shapes and attributes)",
)),
GroupInner::Read(_) => Err(PyIOError::new_err(
"cannot create datasets on a read-only group",
)),
}
}
/// Deleting objects is not supported (h5py's `del group[name]`).
fn __delitem__(&self, key: &str) -> PyResult<()> {
Err(PyNotImplementedError::new_err(format!(
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
)))
}
/// Attribute access.
#[getter]
fn attrs(&self) -> PyResult<PyAttrs> {
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
match &self.inner {
GroupInner::Read(g) => g.attrs(),
GroupInner::Read(g) => g.attrs(py),
GroupInner::Write(state) => {
let store = Arc::clone(&state.lock().unwrap().attrs);
Ok(PyAttrs::from_write(store))
@@ -311,10 +330,10 @@ impl PyGroup {
}
}
fn __repr__(&self) -> String {
fn __repr__(&self, py: Python<'_>) -> String {
match &self.inner {
GroupInner::Read(g) => {
let n = g.member_names().map_or(0, |m| m.len());
let n = g.member_names(py).map_or(0, |m| m.len());
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
}
GroupInner::Write(state) => {
@@ -324,9 +343,9 @@ impl PyGroup {
}
}
fn __contains__(&self, key: &str) -> PyResult<bool> {
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
match &self.inner {
GroupInner::Read(g) => Ok(g.contains(key)),
GroupInner::Read(g) => g.contains(py, key),
GroupInner::Write(state) => {
let guard = state.lock().unwrap();
Ok(guard.datasets.iter().any(|d| d.name == key))
@@ -357,31 +376,6 @@ pub(crate) fn finalize_write_group(
mod tests {
use super::*;
#[test]
fn member_names_are_sorted() {
let mut b = clawhdf5_rs::FileBuilder::new();
b.create_dataset("zeta").with_f64_data(&[1.0]);
b.create_dataset("alpha").with_f64_data(&[1.0]);
let mut g = b.create_group("mid");
g.create_dataset("x").with_f64_data(&[1.0]);
let finished = g.finish();
b.add_group(finished);
let bytes = b.finish().unwrap();
let file = Arc::new(clawhdf5_rs::File::from_bytes(bytes).unwrap());
let root = file.superblock().root_group_address;
let top = ReadGroup::new(Arc::clone(&file), String::new(), root);
assert_eq!(top.member_names().unwrap(), ["alpha", "mid", "zeta"]);
let (path, addr) = top.locate("mid").unwrap();
assert_eq!(path, "mid");
let mid = ReadGroup::new(Arc::clone(&file), path, addr);
assert_eq!(mid.member_names().unwrap(), ["x"]);
assert!(top.contains("mid/x"));
assert!(mid.contains("/alpha"));
assert!(mid.contains("x") && mid.contains("./x"));
assert!(!top.contains("nope"));
assert!(!mid.contains("alpha"));
}
#[test]
fn finalize_group() {
let state = WriteGroupState {
+230
View File
@@ -0,0 +1,230 @@
//! The open file every object of a `File` shares.
//!
//! Every read goes through [`Handle::with`], which releases the GIL and
//! parses through `File::storage()`, so the same code serves a local file
//! (memory-mapped), a remote one (`clawhdf5-remote`: range requests through
//! a block cache, so a network read never holds the GIL) and a file open
//! for editing.
//!
//! A file opened with `'r+'` also holds a [`FileEditor`]. An edit takes the
//! file's write lock, so no read runs while the file changes underneath it,
//! and reopens the file afterwards, through the editor's own open file
//! rather than its path: reads after an edit see the new bytes
//! (a grown file, a new dataspace), never a stale mapping or chunk cache.
//! Objects that cache something an edit can change compare
//! [`Handle::generation`] with the value they cached it at.
//!
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
//! with the GIL released, and code that holds it never touches Python.
use std::sync::atomic::{AtomicU64, Ordering};
use std::sync::{Arc, Mutex, PoisonError, RwLock};
use clawhdf5_rs::{File, FileEditor};
use pyo3::exceptions::PyOSError;
use pyo3::prelude::*;
use crate::{panic_text, to_py_err};
/// Where the file's bytes come from.
pub(crate) enum Source {
/// A local file (memory-mapped). Its path is not kept: nothing reopens
/// it by path (see `open_editable`).
Local,
/// A URL, read through `clawhdf5-remote`'s block cache.
Remote {
url: String,
storage: Arc<clawhdf5_remote::RemoteStorage>,
},
}
pub(crate) struct Handle {
/// The file as last opened; `None` if reopening it after an edit failed
/// (every read is then an error rather than a read of stale bytes).
file: RwLock<Option<File>>,
/// For `'r+'`: the editor, until the file is closed.
editor: Option<Mutex<Option<FileEditor>>>,
source: Source,
/// Bumped by every edit.
generation: AtomicU64,
pub offset_size: u8,
pub length_size: u8,
pub root: u64,
}
fn closed_after_failed_reopen() -> PyErr {
PyOSError::new_err("the file could not be reopened after an edit; open it again")
}
impl Handle {
fn new(file: File, source: Source, editor: Option<FileEditor>) -> Arc<Self> {
let sb = file.superblock();
let (offset_size, length_size, root) =
(sb.offset_size, sb.length_size, sb.root_group_address);
Arc::new(Self {
file: RwLock::new(Some(file)),
editor: editor.map(|e| Mutex::new(Some(e))),
source,
generation: AtomicU64::new(0),
offset_size,
length_size,
root,
})
}
/// A local file, read-only.
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
Ok(Self::new(file, Source::Local, None))
}
/// A local file, open for in-place editing (`'r+'`): the editor takes
/// the file's exclusive lock and checks that it can edit the file, then
/// the file is read through the editor's own file — never by path
/// again, so a later `os.chdir` or a rename or replacement of the path
/// cannot make reads (or the editor's plans) come from another file.
pub(crate) fn open_editable(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
let (file, editor) = py.detach(|| {
crate::no_panic(|| {
let editor = FileEditor::open(path).map_err(to_py_err)?;
let file = editor.reader().map_err(to_py_err)?;
Ok((file, editor))
})
})?;
Ok(Self::new(file, Source::Local, Some(editor)))
}
/// A remote file (`http(s)://`, `s3://`, ...).
pub(crate) fn open_url(
py: Python<'_>,
url: &str,
options: &clawhdf5_remote::Options,
) -> PyResult<Arc<Self>> {
let (file, storage) = py.detach(|| {
crate::no_panic(|| {
let storage = clawhdf5_remote::storage_for_url(url, options).map_err(remote_err)?;
let file = File::open_storage(storage.clone()).map_err(to_py_err)?;
Ok((file, storage))
})
})?;
Ok(Self::new(
file,
Source::Remote {
url: url.to_string(),
storage,
},
None,
))
}
/// Run `f` on the file with the GIL released (a remote read may wait
/// on the network; other Python threads run meanwhile). `f` must not
/// touch Python.
pub(crate) fn with<R: Send>(
&self,
py: Python<'_>,
f: impl FnOnce(&File) -> PyResult<R> + Send,
) -> PyResult<R> {
py.detach(|| self.with_detached(f))
}
/// [`with`](Self::with) for code that already runs without the GIL.
pub(crate) fn with_detached<R>(&self, f: impl FnOnce(&File) -> PyResult<R>) -> PyResult<R> {
crate::no_panic(|| {
let guard = self.file.read().unwrap_or_else(PoisonError::into_inner);
let file = guard.as_ref().ok_or_else(closed_after_failed_reopen)?;
f(file)
})
}
/// Edits so far: objects that cache something an edit can change (a
/// dataset's shape, an object's attributes) re-read it when this moved.
pub(crate) fn generation(&self) -> u64 {
self.generation.load(Ordering::Acquire)
}
/// Whether the file was opened for editing (`'r+'`), even if closed since.
pub(crate) fn is_writable(&self) -> bool {
self.editor.is_some()
}
/// The remote file's block cache.
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
match &self.source {
Source::Remote { storage, .. } => Some(storage),
Source::Local => None,
}
}
/// The URL of a remote file, credentials and query values redacted.
pub(crate) fn redacted_url(&self) -> Option<String> {
match &self.source {
Source::Remote { url, .. } => Some(clawhdf5_remote::redact_url(url)),
Source::Local => None,
}
}
/// Release the editor, and with it the file's lock. Objects still
/// open keep reading the file as it was last written; an edit through
/// them is an error.
pub(crate) fn close(&self) {
if let Some(ed) = &self.editor {
ed.lock().unwrap_or_else(PoisonError::into_inner).take();
}
}
/// Apply one edit with the GIL released. No read runs while it writes,
/// and the file is reopened afterwards — also after a failed edit, since
/// a commit that failed part-way may have changed the file.
pub(crate) fn edit<R: Send>(
&self,
py: Python<'_>,
f: impl FnOnce(&mut FileEditor) -> Result<R, clawhdf5_rs::Error> + Send,
) -> PyResult<R> {
let Some(editor) = &self.editor else {
return Err(PyOSError::new_err(match self.source {
Source::Remote { .. } => "remote files are read-only",
Source::Local => "the file is open read-only; open it with mode 'r+' to change it",
}));
};
if matches!(self.source, Source::Remote { .. }) {
return Err(PyOSError::new_err("remote files are read-only"));
}
py.detach(|| {
let mut ed = editor.lock().unwrap_or_else(PoisonError::into_inner);
let ed = ed
.as_mut()
.ok_or_else(|| PyOSError::new_err("the file is closed"))?;
let mut file = self.file.write().unwrap_or_else(PoisonError::into_inner);
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| f(ed)));
// Drop the old mapping and its chunk cache before reopening.
*file = None;
// Through the editor's file, not the path (see `open_editable`).
let reopened = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| ed.reader()));
self.generation.fetch_add(1, Ordering::AcqRel);
match reopened {
Ok(Ok(f)) => *file = Some(f),
Ok(Err(e)) => return Err(to_py_err(e)),
Err(_) => return Err(closed_after_failed_reopen()),
}
drop(file);
match result {
Ok(r) => r.map_err(to_py_err),
Err(p) => Err(crate::InternalError::new_err(format!(
"clawhdf5 internal error (please report it): {}",
panic_text(&*p)
))),
}
})
}
}
/// A `clawhdf5_remote::Error` as a Python exception: the network side
/// (unreachable, a status, no range support, a changed file) is `OSError`,
/// a file that is not HDF5 is what `to_py_err` makes of it.
pub(crate) fn remote_err(e: clawhdf5_remote::Error) -> PyErr {
match e {
clawhdf5_remote::Error::Hdf5(e) => to_py_err(e),
other => PyOSError::new_err(other.to_string()),
}
}
+14 -1
View File
@@ -7,13 +7,22 @@
//!
//! with clawhdf5.File('data.h5', 'r') as f:
//! data = f['dataset_name'][:]
//!
//! with clawhdf5.File('http://host/data.h5') as f: # range requests
//! block = f['dataset_name'][10:20]
//!
//! with clawhdf5.File('data.h5', 'r+') as f: # in-place edits
//! f['dataset_name'][0] = 1.5
//! f.attrs['note'] = 'edited'
//! ```
mod attrs;
mod convert;
mod dataset;
mod edit;
mod file;
mod group;
mod handle;
mod node;
mod select;
@@ -63,7 +72,7 @@ fn _panic_for_test() -> PyResult<()> {
/// Convert a `clawhdf5_rs::Error` into a `PyErr`.
///
/// Maps different error variants to more specific Python exception types:
/// - I/O errors -> `PyIOError`
/// - I/O errors, and failed reads of a remote file -> `PyIOError`/`PyOSError`
/// - Format/parsing errors -> `PyValueError`
/// - Missing dataset/path errors -> `PyKeyError`
/// - Invalid arguments -> `PyValueError`
@@ -73,6 +82,10 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr {
use clawhdf5_rs::Error;
match &e {
Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()),
// A failed read of the storage: a network error on a remote file.
Error::Format(clawhdf5_format::error::FormatError::Storage(_)) => {
PyErr::new::<pyo3::exceptions::PyOSError, _>(e.to_string())
}
Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()),
Error::NotADataset(_) | Error::MissingMessage(_) => {
PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string())
+72 -51
View File
@@ -1,16 +1,24 @@
//! Resolving paths to objects in a file opened for reading.
//!
//! Everything here parses through `File::storage()` (the `clawhdf5_format`
//! `*_in` functions), never `File::as_bytes()`, so it works the same on a
//! memory-mapped local file and on a remote one; and it runs inside
//! `Handle::with`, without the GIL.
use std::sync::Arc;
use clawhdf5_format::attribute::AttributeMessage;
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
use clawhdf5_format::error::FormatError;
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
use clawhdf5_rs::File;
use pyo3::exceptions::{PyKeyError, PyOSError, PyTypeError, PyValueError};
use pyo3::prelude::*;
use crate::dataset::PyDataset;
use crate::dataset::{DatasetMeta, PyDataset};
use crate::group::PyGroup;
use crate::handle::Handle;
/// Join `key` onto the group path `base` the way h5py does: an absolute key
/// starts from the root, a relative one from `base`. Paths are kept without
@@ -33,42 +41,46 @@ pub(crate) fn name(path: &str) -> String {
format!("/{path}")
}
/// A format error met at `path`: a failed read of the storage (a network
/// error on a remote file) is an `OSError`, anything else `other(message)`.
pub(crate) fn format_err(path: &str, e: FormatError, other: fn(String) -> PyErr) -> PyErr {
let msg = format!("{}: {e}", name(path));
match e {
FormatError::Storage(_) => PyOSError::new_err(msg),
_ => other(msg),
}
}
fn value_err(msg: String) -> PyErr {
PyValueError::new_err(msg)
}
/// The address of the object at `path`, resolved from the root group.
pub(crate) fn address(file: &clawhdf5_rs::File, path: &str) -> PyResult<u64> {
pub(crate) fn address(file: &File, path: &str) -> PyResult<u64> {
resolve_from(file, file.superblock().root_group_address, path, path)
}
/// The address of `rel` resolved from the group at `group` (`full` is the
/// resulting path, for the error message).
pub(crate) fn resolve_from(
file: &clawhdf5_rs::File,
group: u64,
rel: &str,
full: &str,
) -> PyResult<u64> {
pub(crate) fn resolve_from(file: &File, group: u64, rel: &str, full: &str) -> PyResult<u64> {
if rel.is_empty() {
return Ok(group);
}
crate::no_panic(|| {
clawhdf5_format::group_v2::resolve_path_from(file.as_bytes(), file.superblock(), group, rel)
.map_err(|e| {
PyKeyError::new_err(format!(
clawhdf5_format::group_v2::resolve_path_from_in(file.storage(), file.superblock(), group, rel)
.map_err(|e| match e {
FormatError::Storage(_) => format_err(full, e, value_err),
e => PyKeyError::new_err(format!(
"Unable to open object (object '{}' doesn't exist): {e}",
name(full)
))
})
)),
})
}
/// The object header at `addr` (the object at `path`).
pub(crate) fn header_at(file: &clawhdf5_rs::File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
crate::no_panic(|| {
pub(crate) fn header_at(file: &File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
let sb = file.superblock();
let at = usize::try_from(addr)
.map_err(|_| PyValueError::new_err(format!("{}: address out of range", name(path))))?;
ObjectHeader::parse(file.as_bytes(), at, sb.offset_size, sb.length_size)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))
})
ObjectHeader::parse_in(file.storage(), addr, sb.offset_size, sb.length_size)
.map_err(|e| format_err(path, e, value_err))
}
/// What an object header describes.
@@ -96,29 +108,50 @@ pub(crate) fn kind(hdr: &ObjectHeader) -> Option<Kind> {
}
}
/// The kind of the object at `addr`, from its header.
pub(crate) fn kind_at(file: &File, addr: u64, path: &str) -> PyResult<Option<Kind>> {
Ok(kind(&header_at(file, addr, path)?))
}
/// What opening an object found, read without the GIL.
enum Found {
Dataset(DatasetMeta),
Group,
Datatype,
Other,
}
/// Open the object at `addr` (whose path is `path`) as a `Dataset` or
/// `Group`. Both keep the address, so later reads resolve nothing.
pub(crate) fn open(
py: Python<'_>,
file: &Arc<clawhdf5_rs::File>,
handle: &Arc<Handle>,
path: String,
addr: u64,
) -> PyResult<Py<PyAny>> {
let hdr = header_at(file, addr, &path)?;
match kind(&hdr) {
Some(Kind::Dataset) => Ok(PyDataset::open(py, Arc::clone(file), path, addr, &hdr)?
let found = handle.with(py, |f| {
let hdr = header_at(f, addr, &path)?;
Ok(match kind(&hdr) {
Some(Kind::Dataset) => Found::Dataset(DatasetMeta::load(f, addr, &hdr, &path)?),
Some(Kind::Group) => Found::Group,
Some(Kind::Datatype) => Found::Datatype,
None => Found::Other,
})
})?;
match found {
Found::Dataset(meta) => Ok(PyDataset::new(py, Arc::clone(handle), path, addr, meta)
.into_pyobject(py)?
.into_any()
.unbind()),
Some(Kind::Group) => Ok(PyGroup::from_read(Arc::clone(file), path, addr)
Found::Group => Ok(PyGroup::from_read(Arc::clone(handle), path, addr)
.into_pyobject(py)?
.into_any()
.unbind()),
Some(Kind::Datatype) => Err(PyTypeError::new_err(format!(
Found::Datatype => Err(PyTypeError::new_err(format!(
"{}: committed (named) datatypes are not supported by clawhdf5",
name(&path)
))),
None => Err(PyValueError::new_err(format!(
Found::Other => Err(PyValueError::new_err(format!(
"{}: not a dataset, group or datatype",
name(&path)
))),
@@ -126,32 +159,26 @@ pub(crate) fn open(
}
/// The dataspace message of an object header.
pub(crate) fn dataspace(file: &clawhdf5_rs::File, hdr: &ObjectHeader) -> PyResult<Dataspace> {
crate::no_panic(|| {
pub(crate) fn dataspace(file: &File, hdr: &ObjectHeader, path: &str) -> PyResult<Dataspace> {
let sb = file.superblock();
let msg = hdr
.messages
.iter()
.find(|m| m.msg_type == MessageType::Dataspace)
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
let data = clawhdf5_format::shared_message::message_data(
file.as_bytes(),
let data = clawhdf5_format::shared_message::message_data_in(
file.storage(),
msg,
sb.offset_size,
sb.length_size,
)
.map_err(|e| PyValueError::new_err(e.to_string()))?;
Dataspace::parse(&data, sb.length_size).map_err(|e| PyValueError::new_err(e.to_string()))
})
.map_err(|e| format_err(path, e, value_err))?;
Dataspace::parse(&data, sb.length_size).map_err(|e| format_err(path, e, value_err))
}
/// The chunk shape of a chunked dataset (one entry per dataset dimension),
/// or `None` for other layouts or a layout message that does not parse.
pub(crate) fn chunk_shape(
file: &clawhdf5_rs::File,
hdr: &ObjectHeader,
rank: usize,
) -> Option<Vec<u64>> {
pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Option<Vec<u64>> {
let sb = file.superblock();
let msg = hdr
.messages
@@ -179,24 +206,18 @@ pub(crate) fn is_null(space: &Dataspace) -> bool {
/// The attributes of the object at `addr` (whose path is `path`), sorted by
/// name (h5py's order). Attributes whose messages cannot be parsed are left
/// out, as the facade's `attrs()` does.
pub(crate) fn attributes(
file: &clawhdf5_rs::File,
addr: u64,
path: &str,
) -> PyResult<Vec<AttributeMessage>> {
pub(crate) fn attributes(file: &File, addr: u64, path: &str) -> PyResult<Vec<AttributeMessage>> {
let hdr = header_at(file, addr, path)?;
crate::no_panic(|| {
let sb = file.superblock();
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant(
file.as_bytes(),
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
file.storage(),
&hdr,
sb.offset_size,
sb.length_size,
)
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))?;
.map_err(|e| format_err(path, e, value_err))?;
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
Ok(attrs)
})
}
#[cfg(test)]
+27 -4
View File
@@ -7,11 +7,12 @@
//! (negative from the end) drop their axis, slices must have a positive
//! step, one `Ellipsis` fills the unmentioned axes, a single increasing list
//! of integers may index one axis, and strings name compound fields.
//! Everything else (`None`/`np.newaxis`, boolean masks, several index lists)
//! is refused with the error h5py gives.
//! Everything else (`None`/`np.newaxis`, several index lists) is refused
//! with the error h5py gives; boolean masks, which h5py supports, raise
//! `NotImplementedError`.
use clawhdf5_format::selection::Selection;
use pyo3::exceptions::{PyIndexError, PyTypeError, PyValueError};
use pyo3::exceptions::{PyIndexError, PyNotImplementedError, PyTypeError, PyValueError};
use pyo3::prelude::*;
use pyo3::types::{PyEllipsis, PySlice, PyString, PyTuple};
@@ -254,6 +255,17 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
}
}
// A mask of the dataset's whole shape (`ds[ds[()] > 0]`).
if let [a] = args.as_slice() {
let np = key.py().import("numpy")?;
if a.is_instance(&np.getattr("ndarray")?)?
&& a.getattr("dtype")?.getattr("kind")?.extract::<String>()? == "b"
&& a.getattr("ndim")?.extract::<usize>()? > 1
&& a.getattr("shape")?.extract::<Vec<u64>>()? == dims
{
return Err(mask_unsupported());
}
}
if args.iter().any(|a| a.is_none()) {
return Err(PyTypeError::new_err(
"Indexing with None (or np.newaxis) is not supported",
@@ -332,6 +344,10 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
})
}
fn mask_unsupported() -> PyErr {
PyNotImplementedError::new_err("boolean mask indexing is not supported by clawhdf5")
}
fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
if a.is_none() {
return Err(PyTypeError::new_err(
@@ -379,8 +395,15 @@ fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
let arr = np.call_method1("asarray", (a,))?;
let kind: String = arr.getattr("dtype")?.getattr("kind")?.extract()?;
if kind == "b" {
// A mask along this axis (h5py supports them; clawhdf5 does
// not, for reads or writes: an unsupported operation). A mask
// of any other shape is a wrong key, as in h5py.
let shape: Vec<u64> = arr.getattr("shape")?.extract()?;
if shape == [n] {
return Err(mask_unsupported());
}
return Err(PyTypeError::new_err(
"Boolean mask indexing is not supported by clawhdf5",
"Boolean indexing array has incompatible shape",
));
}
let ndim: usize = arr.getattr("ndim")?.extract()?;
+135
View File
@@ -1,6 +1,10 @@
"""Shared fixtures for the clawhdf5 Python binding tests."""
import os
import re
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import pytest
@@ -16,3 +20,134 @@ def h5py():
pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable")
pytest.skip("h5py not installed")
return mod
# ---------------------------------------------------------------------------
# An HTTP server for remote reads
# ---------------------------------------------------------------------------
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
class RangeServer:
"""A static file server on 127.0.0.1, in a thread of this process, that
answers `Range: bytes=a-b` with 206 and `Content-Range` (the way S3 and
common web servers do), sends an ETag and honours `If-Match`.
- `ranges=False`: ignores `Range` and answers 200 with the whole file,
like a server without range support.
- `down` (set by `close()`): hang up on every request.
- `delay`: seconds to wait before answering each request after the
first `delay_after` ones (a slow network).
- `log`: every request as `(method, path, range header)`.
"""
def __init__(self, root, ranges=True):
self.root = str(root)
self.ranges = ranges
self.delay = 0.0
self.delay_after = 0
self.down = False
self.log = []
self._lock = threading.Lock()
server = self
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def log_message(self, *args): # quiet
pass
def do_HEAD(self):
self._serve(body=False)
def do_GET(self):
self._serve(body=True)
def _serve(self, body):
with server._lock:
server.log.append((self.command, self.path, self.headers.get("Range")))
n = len(server.log)
if server.delay and n > server.delay_after:
time.sleep(server.delay)
if server.down:
# Hang up without an answer (open keep-alive
# connections outlive shutdown(), so close() sets this).
self.close_connection = True
return
path = os.path.join(server.root, self.path.lstrip("/").split("?")[0])
if not os.path.isfile(path):
self.send_response(404)
self.send_header("Content-Length", "0")
self.end_headers()
return
with open(path, "rb") as fh:
data = fh.read()
st = os.stat(path)
etag = f'"{st.st_mtime_ns:x}-{st.st_size:x}"'
want = self.headers.get("If-Match")
if want is not None and want != etag and want != "*":
self.send_response(412)
self.send_header("Content-Length", "0")
self.end_headers()
return
rng = self.headers.get("Range") if server.ranges else None
m = _RANGE.match(rng.strip()) if rng else None
if m and (m.group(1) or m.group(2)):
size = len(data)
if m.group(1):
start = int(m.group(1))
end = int(m.group(2)) if m.group(2) else size - 1
else:
start = max(0, size - int(m.group(2)))
end = size - 1
if start >= size:
self.send_response(416)
self.send_header("Content-Range", f"bytes */{size}")
self.send_header("Content-Length", "0")
self.end_headers()
return
end = min(end, size - 1)
part = data[start : end + 1]
self.send_response(206)
self.send_header("Content-Range", f"bytes {start}-{end}/{size}")
else:
part = data
self.send_response(200)
if server.ranges:
self.send_header("Accept-Ranges", "bytes")
self.send_header("ETag", etag)
self.send_header("Content-Length", str(len(part)))
self.send_header("Content-Type", "application/x-hdf5")
self.end_headers()
if body:
try:
self.wfile.write(part)
except (BrokenPipeError, ConnectionResetError):
pass
self.httpd = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
self.httpd.daemon_threads = True
self.port = self.httpd.server_address[1]
self.thread = threading.Thread(target=self.httpd.serve_forever, daemon=True)
self.thread.start()
def url(self, name):
return f"http://127.0.0.1:{self.port}/{name}"
def requests(self):
with self._lock:
return len(self.log)
def close(self):
self.down = True
self.httpd.shutdown()
self.httpd.server_close()
@pytest.fixture
def range_server(tmp_path):
"""A range-capable server over `tmp_path`."""
server = RangeServer(tmp_path)
yield server
server.close()
+940
View File
@@ -0,0 +1,940 @@
"""In-place editing: clawhdf5.File(path, 'r+') against h5py.
Every edit is applied twice, to two copies of the same file: once through
h5py (libhdf5) and once through clawhdf5 (FileEditor). After every edit both
files are read back with h5py and must hold the same shapes, values and
attributes; clawhdf5's own view must agree; when h5py refuses an edit,
clawhdf5 must refuse it too and leave its file as it was. Files are written
by h5py (libver earliest and latest, so every chunk index kind) and by
clawhdf5; `h5dump` must read every result."""
import io
import os
import shutil
import subprocess
import threading
import numpy as np
import pytest
import clawhdf5
# ---------------------------------------------------------------------------
# Files
# ---------------------------------------------------------------------------
ENUM = {"RED": 0, "GREEN": 1, "BLUE": 7}
def _h5py_file(h5py, path, libver):
rng = np.random.default_rng(1)
with h5py.File(path, "w", libver=libver) as f:
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
f.create_dataset("be_i2", data=np.arange(24, dtype=">i2").reshape(4, 6))
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
f.create_dataset("f2", data=rng.standard_normal(12).astype("<f2"))
f.create_dataset("f4_2d", data=rng.standard_normal((8, 9)).astype("<f4"))
f.create_dataset("f8_3d", data=rng.standard_normal((4, 5, 6)))
f.create_dataset("u8", data=np.arange(10, dtype="<u8"))
f.create_dataset("c8", data=(np.arange(6) + 1j * np.arange(6)).astype("<c8"))
f.create_dataset("bool", data=np.array([True, False, True, True]))
f.create_dataset("enum", data=np.array([0, 1, 7, 0], dtype="i1"),
dtype=h5py.enum_dtype(ENUM, basetype="i1"))
f.create_dataset("s5", data=np.array([b"ab", b"cdefg", b""], dtype="S5"))
cmp_dt = np.dtype([("id", "<i4"), ("x", "<f8"), ("tag", "S3")])
f.create_dataset("cmp", data=np.array([(i, i / 2, b"t%d" % i) for i in range(5)], dtype=cmp_dt))
f.create_dataset("scalar", data=np.float64(3.5))
# Chunked: fixed maxshape (v4 fixed array under latest), one
# unlimited dimension (extensible array), two (v2 B-tree), one chunk.
f.create_dataset("chunk_fixed", data=np.arange(100, dtype="<i8").reshape(10, 10), chunks=(3, 4))
f.create_dataset("chunk_ext", data=rng.standard_normal((12, 7)), chunks=(5, 7), maxshape=(None, 7))
f.create_dataset("chunk_bt2", data=np.arange(30, dtype="<i4").reshape(5, 6), chunks=(2, 2),
maxshape=(None, None))
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20), chunks=(6, 6),
compression="gzip", maxshape=(40, 40))
f.create_dataset("chunk_single", data=np.arange(12, dtype="<u2").reshape(3, 4), chunks=(3, 4),
maxshape=(3, 4))
f.create_dataset("chunk_fill", shape=(8,), dtype="<i4", chunks=(3,), maxshape=(20,), fillvalue=-1)
f.create_dataset("vlen", data=["a", "bb"], dtype=h5py.string_dtype())
# Compact layout (low-level API).
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
dcpl.set_layout(h5py.h5d.COMPACT)
space = h5py.h5s.create_simple((7,))
dsid = h5py.h5d.create(f.id, b"compact", h5py.h5t.STD_I32LE, space, dcpl=dcpl)
dsid.write(h5py.h5s.ALL, h5py.h5s.ALL, np.arange(7, dtype="<i4"))
g = f.create_group("grp")
g.create_dataset("leaf", data=np.arange(5.0))
g.attrs["units"] = "m"
f.attrs["version"] = np.int32(1)
def _clawhdf5_file(path):
with clawhdf5.File(str(path), "w") as f:
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
f.create_dataset("f8", data=np.linspace(0, 1, 30).reshape(5, 6))
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20),
chunks=[6, 6], compression="gzip")
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
g = f.create_group("grp")
g.create_dataset("leaf", data=np.arange(5.0))
g.attrs["units"] = "m"
f.attrs["version"] = 1
# h5py's libver: "earliest" (v1 B-tree chunk indexes), "v114" (the 1.10+
# indexes: fixed and extensible arrays, v2 B-trees, single chunk) and
# "latest" (HDF5 2.0's newest format, which h5dump 1.14 cannot read).
SOURCES = ["h5py-earliest", "h5py-v114", "h5py-latest", "clawhdf5"]
def _make(h5py, tmp_path, source):
base = tmp_path / f"base-{source}.h5"
if source == "clawhdf5":
_clawhdf5_file(base)
else:
_h5py_file(h5py, str(base), source.split("-")[1])
theirs = tmp_path / f"theirs-{source}.h5"
ours = tmp_path / f"ours-{source}.h5"
shutil.copy(base, theirs)
shutil.copy(base, ours)
return str(theirs), str(ours), str(base)
# ---------------------------------------------------------------------------
# Comparing files through h5py
# ---------------------------------------------------------------------------
def _norm_attr(v):
"""Attribute values comparable across the two writers: clawhdf5 stores
`str` as fixed-length UTF-8 (h5py reads bytes), h5py as variable-length
(h5py reads str)."""
if isinstance(v, bytes):
return ("str", v.decode("utf-8"))
if isinstance(v, str):
return ("str", v)
arr = np.asarray(v)
if arr.dtype.kind in "SO":
return ("strs", [x.decode() if isinstance(x, bytes) else x for x in arr.ravel().tolist()], arr.shape)
return (arr.dtype.str, arr.shape, arr.tobytes())
def snapshot(h5py, path):
"""What h5py sees in the file: every dataset's shape, dtype, bytes and
attributes (read without locking: clawhdf5 may hold the file open)."""
out = {}
with h5py.File(path, "r", locking=False) as f:
def visit(name, obj):
attrs = {k: _norm_attr(obj.attrs[k]) for k in obj.attrs}
if isinstance(obj, h5py.Dataset):
if obj.dtype.kind == "O":
data = [x for x in obj[...].ravel().tolist()]
else:
data = obj[()].tobytes() if obj.shape is not None else None
out[name] = (obj.shape, obj.dtype.str, obj.maxshape, data, attrs)
else:
out[name] = ("group", attrs)
visit("/", f)
f.visititems(visit)
return out
def assert_same_files(h5py, theirs, ours, what):
a, b = snapshot(h5py, theirs), snapshot(h5py, ours)
assert a.keys() == b.keys(), what
for k in a:
assert a[k] == b[k], f"{what}: {k} differs\n h5py: {a[k]}\n clawhdf5: {b[k]}"
def assert_ours_reads_like_h5py(h5py, f, path, what):
"""clawhdf5's own view of the file it is editing matches h5py's."""
with h5py.File(path, "r", locking=False) as t:
for name in ["i4", "chunk_ext", "chunk_bt2", "chunk_gzip", "f8_3d", "cmp", "bool", "enum", "scalar"]:
if name not in t:
continue
o = f[name]
assert o.shape == t[name].shape, f"{what}: {name} shape"
assert o.maxshape == t[name].maxshape, f"{what}: {name} maxshape"
np.testing.assert_array_equal(o[()], t[name][()], err_msg=f"{what}: {name}")
for obj in ["/", "grp"]:
assert sorted(f[obj].attrs.keys()) == sorted(t[obj].attrs.keys()), what
for k in t[obj].attrs:
assert _norm_attr(f[obj].attrs[k]) == _norm_attr(t[obj].attrs[k]), f"{what}: {obj}.attrs[{k}]"
def h5dump_reads(path, base=None):
"""h5dump (libhdf5 1.14) reads every object and value of `path` — when it
reads the unedited `base` (it cannot read HDF5 2.0's newest format)."""
exe = shutil.which("h5dump")
if exe is None:
if os.environ.get("CLAWHDF5_REQUIRE_INTEROP") == "1":
pytest.fail("h5dump is required (CLAWHDF5_REQUIRE_INTEROP=1)")
return
h5rs = os.environ.get("CLAWHDF5_H5RS")
if h5rs:
# clawhdf5's structural and checksum validator (scripts/ci-test.sh
# points this at the h5rs it built).
r = subprocess.run([h5rs, "check", path], capture_output=True, text=True)
assert r.returncode == 0, (r.stdout + r.stderr)[-2000:]
if base is not None and subprocess.run([exe, "-H", base], capture_output=True).returncode != 0:
return
r = subprocess.run([exe, path], capture_output=True, text=True)
assert r.returncode == 0, r.stderr[-2000:]
# ---------------------------------------------------------------------------
# Applying one edit both ways
# ---------------------------------------------------------------------------
def _native_conversion(h5py, value, ds_dtype):
"""`value` as libhdf5 converts it to `ds_dtype` in native byte order.
libhdf5 2.0 (h5py 3.16) converts numbers differently when either side is
not in native byte order (its "soft" conversions): a float in (-1, 0)
becomes the integer type's minimum instead of 0, and an unsigned integer
too large for the signed type of the same size wraps instead of
saturating. clawhdf5 applies the native-order results to every byte
order, so the reference is h5py converting into a native dataset of the
same kind; the result then reaches the real dataset by a plain byte
swap."""
if not isinstance(value, np.ndarray):
return value
if value.dtype.kind not in "biuf" or ds_dtype.kind not in "biuf":
return value
if value.dtype.isnative and ds_dtype.isnative:
return value
with h5py.File(io.BytesIO(), "w") as tmp:
d = tmp.create_dataset("t", shape=value.shape, dtype=ds_dtype.newbyteorder("="))
d[...] = value.astype(value.dtype.newbyteorder("="))
return np.asarray(d[()])
def _apply(f, op, h5py=None):
"""Apply `op` to `f`; with `h5py`, `f` is an h5py file and a numpy array
value is first converted as libhdf5 converts in native byte order (see
`_native_conversion`)."""
kind = op[0]
if kind == "set":
_, name, key, value = op
if h5py is not None:
value = _native_conversion(h5py, value, f[name].dtype)
f[name][key] = value
elif kind == "resize":
_, name, size, axis = op
if axis is None:
f[name].resize(size)
else:
f[name].resize(size, axis=axis)
elif kind == "attr":
_, obj, name, value = op
f[obj].attrs[name] = value
else:
raise AssertionError(op)
def edit_both(h5py, theirs, ours_path, ours, op):
"""Apply `op` with h5py and with clawhdf5 (`ours`, open 'r+'); the two
files must then read the same through h5py. If h5py refuses, clawhdf5
must refuse and its file must be unchanged. Returns h5py's error."""
before = snapshot(h5py, ours_path)
try:
with h5py.File(theirs, "r+") as t:
_apply(t, op, h5py)
except Exception as e: # noqa: BLE001 - h5py refuses: so must we
try:
_apply(ours, op)
except Exception: # noqa: BLE001
pass
else:
pytest.fail(f"{op!r:.300}: h5py refused ({type(e).__name__}: {e}), clawhdf5 did not")
assert snapshot(h5py, ours_path) == before, f"{op!r}: clawhdf5 changed the file while failing"
return e
_apply(ours, op)
assert_same_files(h5py, theirs, ours_path, repr(op)[:200])
return None
# ---------------------------------------------------------------------------
# Tests
# ---------------------------------------------------------------------------
@pytest.mark.parametrize("source", SOURCES)
def test_edit_sequence_matches_h5py(h5py, tmp_path, source):
theirs, ours_path, base = _make(h5py, tmp_path, source)
ops = [
("set", "i4", 0, 99),
("set", "i4", (slice(1, 5, 2), slice(None, None, 3)), np.array([[1.5, -2.5, 1e12, -1e12]])),
("set", "i4", (slice(None), 2), np.arange(6, dtype="<i8") * 1000),
("set", "i4", ([0, 2, 5], slice(4, 6)), np.array([[1, 2], [3, 4], [5, 6]])),
("set", "i4", (Ellipsis, -1), np.int8(-7)),
("set", "i4", (3, 3), 12345.9),
("set", "u1", slice(2, 8), np.array([-5, 0, 300, 255, 256, 1], dtype="<i4")),
("set", "u1", slice(0, 3), [1, 2, 3]),
("set", "chunk_gzip", (slice(0, 20, 7), slice(3, 17)), 42.25),
("set", "chunk_gzip", (slice(5, 11), slice(5, 11)), np.ones((6, 6), dtype="<f8") * np.pi),
("set", "grp/leaf", slice(None), np.array([1, 2, 3, 4, 5], dtype="<u2")),
("attr", "/", "version", np.int32(2)),
("attr", "/", "count", 7),
("attr", "grp", "scale", np.array([0.5, 0.25], dtype="<f4")),
("attr", "grp", "matrix", np.arange(6, dtype=">i8").reshape(2, 3)),
("attr", "i4", "flag", np.bool_(True)),
("attr", "i4", "z", np.complex64(1 - 2j)),
("attr", "i4", "raw", np.bytes_(b"abc")),
("set", "i4", slice(0, 2), np.zeros((3, 10))), # shape mismatch: refused by both
]
if source != "clawhdf5":
ops += [
("set", "be_i2", (slice(None), slice(1, 3)), np.array([70000, -70000], dtype="<i8")),
("set", "f2", slice(None, None, 4), np.array([1e6, -3.25, 0.1])),
("set", "f4_2d", (2, slice(None)), np.linspace(-1, 1, 9)),
("set", "f8_3d", (slice(1, 3), 2, slice(None, None, 2)), np.arange(3, dtype="<i2")),
("set", "u8", slice(None), np.array([-1, 0, 2**63, 1e30, -1e30, 5.5, 2, 3, 4, 5])),
("set", "c8", slice(1, 3), np.array([1 + 1j, 2 - 2j], dtype="<c16")),
("set", "c8", 0, np.float64(1.0)), # h5py: no conversion path
("set", "bool", slice(None), np.array([0, 3, 0, -1], dtype="<i4")),
("set", "bool", 1, np.array(True)),
("set", "enum", slice(0, 2), np.array([7, 1], dtype="<i4")),
("set", "s5", 0, np.bytes_(b"xyzuvw")),
("set", "s5", slice(1, 3), [b"q", b"rs"]),
("set", "s5", 2, np.array("uni")), # h5py: no conversion from 'U'
("set", "cmp", 2, np.array((9, 9.5, b"zz"), dtype=[("id", "<i4"), ("x", "<f8"), ("tag", "S3")])),
("set", "cmp", slice(3, 5), [(1, 0.5, b"a"), (2, 1.5, b"b")]),
("set", "scalar", (), 7.25),
("set", "scalar", Ellipsis, np.float32(-1.5)),
("set", "compact", slice(1, 6, 2), np.array([10, 20, 30])),
("set", "chunk_fixed", (slice(2, 9), slice(1, 10, 4)), np.arange(21).reshape(7, 3)),
("set", "chunk_single", (1, slice(None)), np.array([9, 8, 7, 6])),
("resize", "chunk_ext", (20, 7), None),
("set", "chunk_ext", slice(12, 20), np.full((8, 7), 2.5)),
("resize", "chunk_ext", 9, 0),
("resize", "chunk_ext", 16, 0),
("resize", "chunk_bt2", (9, 11), None),
("set", "chunk_bt2", (slice(4, 9), slice(5, 11)), np.arange(30).reshape(5, 6)),
("resize", "chunk_bt2", (3, 3), None),
("resize", "chunk_bt2", (7, 8), None),
("resize", "chunk_gzip", (40, 25), None),
("resize", "chunk_gzip", (41, 25), None), # beyond maxshape: refused
("resize", "chunk_fill", 15, None), # h5py: a size without axis must be a tuple
("resize", "chunk_fill", (15,), None),
("set", "chunk_fill", slice(10, 12), [1, 2]),
("resize", "chunk_fill", (4,), None),
("resize", "chunk_fill", (20,), None),
("resize", "i4", (7, 10), None), # not chunked: refused
]
with clawhdf5.File(ours_path, "r+") as ours:
assert ours.mode == "r+"
for i, op in enumerate(ops):
edit_both(h5py, theirs, ours_path, ours, op)
if i % 5 == 0:
assert_ours_reads_like_h5py(h5py, ours, ours_path, repr(op))
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
h5dump_reads(ours_path, base)
# Reopened, clawhdf5 reads what h5py reads.
with clawhdf5.File(ours_path, "r") as f:
assert f.mode == "r"
assert_ours_reads_like_h5py(h5py, f, ours_path, "reopened")
def _random_key(rng, shape):
key = []
for n in shape:
r = rng.random()
if n == 0 or r < 0.3:
start = int(rng.integers(0, n + 1)) if n else 0
stop = int(rng.integers(start, n + 1)) if n else 0
step = int(rng.integers(1, 4))
key.append(slice(start, stop, step))
elif r < 0.55:
key.append(int(rng.integers(-n, n)))
elif r < 0.7:
k = int(rng.integers(1, min(n, 4) + 1))
key.append(sorted(rng.choice(n, size=k, replace=False).tolist()))
else:
key.append(slice(None))
# Only one index list per key.
lists = [i for i, k in enumerate(key) if isinstance(k, list)]
for i in lists[1:]:
key[i] = slice(None)
return tuple(key)
def _selection_shape(key, shape):
out = []
fancy = False
for k, n in zip(key, shape):
if isinstance(k, slice):
out.append(len(range(*k.indices(n))))
elif isinstance(k, list):
out.append(len(k))
fancy = True
return tuple(out), fancy
def _random_value(rng, sel_shape, fancy, dtype):
r = rng.random()
if r < 0.2 or not sel_shape:
v = rng.standard_normal() * 1000
return np.float64(v) if rng.random() < 0.5 else int(v)
shape = list(sel_shape)
if not fancy and r < 0.35 and shape:
shape[0] = 1 # broadcast along the first axis
kind = rng.choice(["same", "f8", "i8", "u1", "f4"])
if kind == "same" and np.dtype(dtype).kind in "iuf":
dt = np.dtype(dtype)
else:
dt = np.dtype(str(kind) if kind != "same" else "f8")
base = rng.standard_normal(size=shape) * (10 ** rng.integers(0, 6))
with np.errstate(all="ignore"):
return base.astype(dt)
@pytest.mark.parametrize("source", SOURCES)
@pytest.mark.parametrize("seed", [0, 1, 2, 3])
def test_random_edits_match_h5py(h5py, tmp_path, source, seed):
"""Random writes (slices, steps, integers, index lists, broadcasts,
other dtypes and out-of-range values), resizes and attributes, each
compared with h5py applying the same edit."""
rng = np.random.default_rng(1000 * seed + SOURCES.index(source))
theirs, ours_path, base = _make(h5py, tmp_path, source)
with h5py.File(theirs, "r") as t:
names = [n for n in ["i4", "u1", "f8", "f4_2d", "f8_3d", "be_i2", "chunk_fixed", "chunk_ext",
"chunk_bt2", "chunk_gzip", "chunk_fill", "compact", "grp/leaf"] if n in t]
resizable = [n for n in ["chunk_ext", "chunk_bt2", "chunk_gzip", "chunk_fill"] if n in names]
refused = 0
with clawhdf5.File(ours_path, "r+") as ours:
for step in range(40):
r = rng.random()
if r < 0.15 and resizable:
name = str(rng.choice(resizable))
with h5py.File(theirs, "r") as t:
maxshape = t[name].maxshape
shape = t[name].shape
new = tuple(int(rng.integers(0, (m if m is not None else s + 10) + 1)) for m, s in zip(maxshape, shape))
op = ("resize", name, new, None)
elif r < 0.25:
obj = str(rng.choice(["/", "grp", names[0]]))
choices = [np.int16(rng.integers(-100, 100)), rng.standard_normal(3),
np.arange(int(rng.integers(1, 5)), dtype=">u4"), np.float32(0.5)]
value = choices[int(rng.integers(0, len(choices)))]
op = ("attr", obj, f"a{int(rng.integers(0, 4))}", value)
else:
name = str(rng.choice(names))
with h5py.File(theirs, "r") as t:
shape, dtype = t[name].shape, t[name].dtype
key = _random_key(rng, shape)
sel_shape, fancy = _selection_shape(key, shape)
op = ("set", name, key, _random_value(rng, sel_shape, fancy, dtype))
before = None
if op[0] == "resize":
with h5py.File(ours_path, "r", locking=False) as o:
before, fill = o[op[1]][()], o[op[1]].fillvalue
if edit_both(h5py, theirs, ours_path, ours, op) is not None:
refused += 1
elif before is not None:
# Independently of h5py (which a wrong index layout fools
# the same way): the kept elements keep their values.
want = resized_model(before, op[2], fill)
np.testing.assert_array_equal(ours[op[1]][()], want, err_msg=repr(op))
with h5py.File(ours_path, "r", locking=False) as o:
np.testing.assert_array_equal(o[op[1]][()], want, err_msg=repr(op))
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
assert refused < 30
h5dump_reads(ours_path, base)
@pytest.mark.parametrize("seed", range(10, 40))
def test_random_edits_on_clawhdf5_files(h5py, tmp_path, seed):
"""More random sequences on a clawhdf5-written file, whose resizes to
zero extents once left files `h5rs check` could not read."""
test_random_edits_match_h5py(h5py, tmp_path, "clawhdf5", seed)
def resized_model(before, shape, fill):
"""`before` resized to `shape` as HDF5 resizes: elements inside both
extents keep their values, the others read as the fill value."""
out = np.full(shape, fill, dtype=before.dtype)
common = tuple(slice(0, min(a, b)) for a, b in zip(before.shape, shape))
out[common] = before[common]
return out
RESIZES = [(15, 15), (3, 2), (20, 20), (1, 1), (1, 0), (0, 0), (7, 20), (20, 13), (20, 20)]
def _check_resizes(h5py, path, name, orig, maxshape=(20, 20), base=None):
"""Resize `name` through RESIZES in 'r+', each checked against a numpy
model with clawhdf5 and h5py and the file with `h5rs check`."""
model = orig
with clawhdf5.File(path, "r+") as f:
ds = f[name]
for shape in RESIZES:
ds.resize(shape)
model = resized_model(model, shape, 0)
np.testing.assert_array_equal(ds[()], model, err_msg=f"{name} {shape}")
with h5py.File(path, "r", locking=False) as t:
np.testing.assert_array_equal(t[name][()], model, err_msg=f"h5py: {name} {shape}")
assert t[name].maxshape == maxshape
h5rs = os.environ.get("CLAWHDF5_H5RS")
if h5rs: # h5dump would wait for the editor's lock
r = subprocess.run([h5rs, "check", "--data", path], capture_output=True, text=True)
assert r.returncode == 0, f"{name} {shape}: " + (r.stdout + r.stderr)[-2000:]
with pytest.raises(ValueError):
ds.resize((maxshape[0] + 1, 20))
# Written values survive a shrink.
ds[...] = orig
ds.resize((15, 15))
with h5py.File(path, "r") as t:
np.testing.assert_array_equal(t[name][()], orig[:15, :15])
h5dump_reads(path, base)
@pytest.mark.parametrize("source", SOURCES)
def test_resizes_keep_values(h5py, tmp_path, source):
"""Shrinking, zero extents and growing back keep the values a numpy
model keeps, on every source (on clawhdf5's own files a shrink once
moved every chunk: h5py read the same wrong values)."""
_, ours, base = _make(h5py, tmp_path, source)
with clawhdf5.File(ours, "r+") as f:
f["chunk_gzip"].resize((20, 20))
maxshape = (20, 20) if source == "clawhdf5" else (40, 40)
_check_resizes(h5py, ours, "chunk_gzip", np.arange(400, dtype="<f4").reshape(20, 20), maxshape, base)
FIXTURES = os.path.join(os.path.dirname(__file__), "..", "..", "clawhdf5", "tests", "fixtures")
@pytest.mark.parametrize("name", ["d", "z"])
def test_resizes_of_a_file_with_no_recorded_maxshape(h5py, tmp_path, name):
"""A file clawhdf5 2.7.0 wrote records no maximum dimensions for chunked
datasets; resizing it must keep the Fixed Array index's layout."""
path = str(tmp_path / "old.h5")
shutil.copy(os.path.join(FIXTURES, "chunked_no_maxshape_v2_7_0.h5"), path)
with h5py.File(path, "r") as t:
orig = t[name][()]
_check_resizes(h5py, path, name, orig)
NUMERIC = ["<i1", "<u1", "<i2", ">u2", "<i4", "<u4", ">i8", "<u8", "<f2", "<f4", ">f8"]
def _quiet(f):
with np.errstate(all="ignore"):
return f()
def _libhdf5_undefined(vals, target):
"""The values whose conversion to `target` libhdf5 2.0 (h5py 3.16) gets
wrong even in native byte order, where its C casts are undefined
behaviour; clawhdf5 saturates them as libhdf5's range handling
intends (docs/known-issues.md):
- half floats into unsigned integers: negatives wrap (-1 -> 65535) and
+inf becomes 0; into signed integers, +-inf becomes the minimum;
- a float equal to the integer maximum rounded up in the float's
precision (float32(2**31 - 1) == 2**31 -> int32, float64(2**64 - 1)
-> uint64) becomes the minimum (or 0);
- a double between 65504 and 65520 into a half float becomes infinity
(IEEE rounds it down to 65504, as numpy does).
"""
bad = np.zeros(vals.shape, dtype=bool)
if vals.dtype.kind != "f":
return bad
if target.kind in "iu":
if vals.dtype.itemsize == 2:
bad |= np.isinf(vals)
if target.kind == "u":
bad |= vals <= -1
top = _quiet(lambda: np.array(np.iinfo(target).max).astype(vals.dtype))
if float(top) > np.iinfo(target).max:
bad |= vals == top
if target.kind == "f" and target.itemsize == 2 and vals.dtype.itemsize > 2:
bad |= (np.abs(vals) > 65504) & (np.abs(vals) < 65520)
return bad
def test_numeric_conversions_match_h5py(h5py, tmp_path):
"""Every numeric source dtype into every numeric dataset dtype, with
values at and beyond the targets' limits, as libhdf5 converts them."""
edge = np.array([0, 1, -1, -0.3, 2.5, -2.5, 3.7, -3.7, 127.9, -128.9, 200.5, 255.5, 256, -129,
32767.5, 40000, 65504, 70000, -70000, 2**31 - 1, 2**31, -2**31 - 1,
4e9, 1e15, -1e15, 1e19, 1e300, -1e300, np.inf, -np.inf])
sources = {
"f8": edge,
"f4": _quiet(lambda: edge.astype("<f4")),
"f2": np.array([0, 1, -1, 2.5, -3.5, 65504, -65504, np.inf, -np.inf, 100.5], dtype="<f2"),
"i8": np.array([0, 1, -1, 127, 128, -129, 255, 256, 32768, -32769, 65536, 2**31, -2**31 - 1,
2**32, 2**62, -2**63, 2**63 - 1], dtype="<i8"),
"u8": np.array([0, 1, 127, 128, 255, 256, 65535, 65536, 2**31, 2**32, 2**63, 2**64 - 1], dtype="<u8"),
"i1": np.array([-128, -1, 0, 1, 127], dtype="i1"),
"u2": np.array([0, 255, 256, 65535], dtype=">u2"),
"b": np.array([True, False, True]),
}
for target in NUMERIC:
for sname, src in sources.items():
vals = src[~_libhdf5_undefined(src, np.dtype(target))]
path_t = str(tmp_path / f"t_{target[1:]}_{sname}.h5")
path_o = str(tmp_path / f"o_{target[1:]}_{sname}.h5")
with h5py.File(path_t, "w") as f:
f.create_dataset("d", shape=vals.shape, dtype=target)
shutil.copy(path_t, path_o)
what = f"{vals.dtype} -> {target}"
try:
with h5py.File(path_t, "r+") as f:
f["d"][...] = _native_conversion(h5py, vals, f["d"].dtype)
except Exception: # noqa: BLE001
with clawhdf5.File(path_o, "r+") as f, pytest.raises(Exception):
f["d"][...] = vals
continue
with clawhdf5.File(path_o, "r+") as f:
f["d"][...] = vals
with h5py.File(path_t, "r") as a, h5py.File(path_o, "r") as b:
assert a["d"][...].tobytes() == b["d"][...].tobytes(), (
f"{what}: h5py {a['d'][...].tolist()} clawhdf5 {b['d'][...].tolist()}"
)
def test_nan_into_an_integer_dataset_is_refused(h5py, tmp_path):
"""libhdf5 stores NaN as an arbitrary integer (0, the minimum or 2**63,
depending on the type); clawhdf5 refuses and writes nothing."""
path = str(tmp_path / "nan.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
with clawhdf5.File(path, "r+") as f:
with pytest.raises(ValueError, match="NaN"):
f["d"][...] = np.array([1.0, np.nan, 2.0, 3.0])
# A Python list goes through numpy, which refuses NaN too.
with pytest.raises(ValueError):
f["d"][0:2] = [np.nan, 1.0]
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["d"][...], np.arange(4))
def test_unsupported_edits_are_clear_errors(h5py, tmp_path):
path = str(tmp_path / "u.h5")
_h5py_file(h5py, path, "earliest")
before = snapshot(h5py, path)
with clawhdf5.File(path, "r+") as f:
with pytest.raises(NotImplementedError, match="delet"):
del f.attrs["version"]
with pytest.raises(NotImplementedError, match="delet"):
del f["i4"]
with pytest.raises(NotImplementedError, match="delet"):
del f["grp"]["leaf"]
with pytest.raises(NotImplementedError):
f.create_dataset("new", data=np.arange(3.0))
with pytest.raises(NotImplementedError):
f.create_group("newgrp")
with pytest.raises(NotImplementedError):
f["grp"].create_dataset("new", data=np.arange(3.0))
with pytest.raises(NotImplementedError, match="variable-length"):
f["vlen"][0] = "x"
with pytest.raises(NotImplementedError, match="field"):
f["cmp"]["id"] = np.arange(5)
with pytest.raises(NotImplementedError):
f.attrs["empty"] = clawhdf5.Empty("f8")
with pytest.raises(TypeError, match="chunked"):
f["i4"].resize((7, 10))
with pytest.raises(ValueError):
f["chunk_gzip"].resize((41, 20))
with pytest.raises(ValueError, match="axis"):
f["chunk_ext"].resize(3, axis=2)
with pytest.raises(TypeError):
f["i4"][0] = np.array(["a"] * 10)
# h5py writes through boolean masks; clawhdf5 does not.
with pytest.raises(NotImplementedError, match="mask"):
f["u1"][np.arange(16) % 2 == 0] = 5
with pytest.raises(NotImplementedError, match="mask"):
f["i4"][f["i4"][()] > 30] = 0
with pytest.raises(NotImplementedError, match="mask"):
f["i4"][np.ones(6, dtype=bool), 2] = 0
assert snapshot(h5py, path) == before
h5dump_reads(path)
def test_read_only_files_and_modes(h5py, tmp_path):
path = str(tmp_path / "m.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"), chunks=(2,), maxshape=(None,))
with clawhdf5.File(path, "r") as f:
with pytest.raises(OSError, match="r\\+"):
f["d"][0] = 1
with pytest.raises(OSError):
f["d"].resize((8,))
with pytest.raises(OSError):
f.attrs["x"] = 1
with pytest.raises(NotImplementedError, match="does not exist"):
clawhdf5.File(str(tmp_path / "missing.h5"), "a")
with pytest.raises(ValueError, match="mode"):
clawhdf5.File(path, "rw")
with clawhdf5.File(path, "a") as f:
assert f.mode == "r+"
f["d"][1] = 10
with h5py.File(path, "r") as f:
assert f["d"][1] == 10
def test_the_file_is_locked_while_open_for_editing(h5py, tmp_path):
path = str(tmp_path / "lock.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
f = clawhdf5.File(path, "r+")
with pytest.raises(OSError):
clawhdf5.File(path, "r+")
with pytest.raises(OSError):
h5py.File(path, "r+")
f.close()
with h5py.File(path, "r+") as g:
g["d"][0] = 5
with clawhdf5.File(path, "r+") as g:
g["d"][1] = 6
with h5py.File(path, "r") as g:
np.testing.assert_array_equal(g["d"][...], [5, 6, 2, 3])
def test_objects_see_edits_made_through_others(h5py, tmp_path):
"""A dataset or attrs object taken before an edit reports the file as
it is after it: the new shape, the new attribute."""
path = str(tmp_path / "live.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(6.0), chunks=(4,), maxshape=(None,))
with clawhdf5.File(path, "r+") as f:
d1 = f["d"]
d2 = f["d"]
attrs = d1.attrs
assert len(attrs) == 0 and "u" not in attrs
d2.resize((10,))
assert d1.shape == (10,) and d1.size == 10 and len(d1) == 10
np.testing.assert_array_equal(d1[6:], np.zeros(4))
d2.attrs["u"] = "m/s"
assert "u" in attrs and attrs["u"] == b"m/s" and len(attrs) == 1
f.attrs.create("shaped", np.arange(6), shape=(2, 3), dtype="<i2")
f.attrs.modify("shaped2", [1.5, 2.5])
with h5py.File(path, "r") as f:
assert f["d"].shape == (10,)
assert f["d"].attrs["u"] == b"m/s"
assert f.attrs["shaped"].dtype == np.dtype("<i2") and f.attrs["shaped"].shape == (2, 3)
np.testing.assert_array_equal(f.attrs["shaped2"], [1.5, 2.5])
def test_attribute_types_as_h5py_reads_them(h5py, tmp_path):
path = str(tmp_path / "attrs.h5")
with h5py.File(path, "w") as f:
f.create_group("g")
values = {
"i8": 5,
"f8": 2.5,
"i1": np.int8(-3),
"u8": np.uint64(2**64 - 1),
">f4": np.array([1.5, 2.5], dtype=">f4"),
"f2": np.float16(0.5),
"b": True,
"barr": np.array([True, False]),
"c16": np.complex128(1 + 2j),
"bytes": b"raw",
"sarr": np.array([b"a", b"bcd"]),
"str": "héllo",
"strs": ["x", "yz"],
"2d": np.arange(12, dtype="<u2").reshape(3, 4),
"empty": np.zeros((0,), dtype="<i4"),
}
with clawhdf5.File(path, "r+") as f:
for k, v in values.items():
f["g"].attrs[k] = v
# Replace one, with another type and size.
f["g"].attrs["i8"] = np.arange(100.0)
with h5py.File(path, "r") as f:
a = f["g"].attrs
np.testing.assert_array_equal(a["i8"], np.arange(100.0))
assert a["f8"] == 2.5 and a["f8"].dtype == np.float64
assert a["i1"] == -3 and a["i1"].dtype == np.int8
assert a["u8"] == 2**64 - 1 and a["u8"].dtype == np.uint64
assert a[">f4"].dtype == np.dtype(">f4")
assert a["f2"].dtype == np.float16
assert a["b"] is np.True_ or a["b"] == True # noqa: E712
assert a["barr"].dtype == np.bool_
assert a["c16"] == 1 + 2j
assert a["bytes"] == b"raw"
assert list(a["sarr"]) == [b"a", b"bcd"]
# str is stored as fixed-length UTF-8: h5py reads bytes.
assert a["str"].decode("utf-8") == "héllo"
assert [x.decode() for x in a["strs"]] == ["x", "yz"]
assert a["2d"].shape == (3, 4) and a["2d"].dtype == np.dtype("<u2")
assert a["empty"].shape == (0,)
with clawhdf5.File(path, "r") as f:
assert f["g"].attrs["c16"] == 1 + 2j
assert f["g"].attrs["str"].decode("utf-8") == "héllo"
h5dump_reads(path)
def test_many_attributes_move_to_dense_storage(h5py, tmp_path):
"""Past the compact limit (8 attributes under libver v114) the object's
attributes move to dense storage; h5py reads all of them."""
path = str(tmp_path / "dense.h5")
with h5py.File(path, "w", libver="v114") as f:
f.create_dataset("d", data=np.arange(3))
with clawhdf5.File(path, "r+") as f:
for i in range(20):
f["d"].attrs[f"a{i:02d}"] = np.full(i + 1, i, dtype="<i2")
with h5py.File(path, "r") as f:
assert sorted(f["d"].attrs.keys()) == [f"a{i:02d}" for i in range(20)]
for i in range(20):
np.testing.assert_array_equal(f["d"].attrs[f"a{i:02d}"], np.full(i + 1, i))
h5dump_reads(path)
def test_reads_never_see_a_half_written_edit(h5py, tmp_path):
"""Readers on other threads while one thread rewrites a dataset: every
read returns one whole version (all elements equal), never a mix."""
path = str(tmp_path / "race.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.zeros((64, 64)), chunks=(16, 16), compression="gzip")
errors = []
with clawhdf5.File(path, "r+") as f:
stop = threading.Event()
def read():
ds = f["d"]
while not stop.is_set():
a = ds[...]
if not (a == a.flat[0]).all():
errors.append(a)
return
readers = [threading.Thread(target=read) for _ in range(3)]
for t in readers:
t.start()
try:
for k in range(1, 25):
f["d"][...] = float(k)
finally:
stop.set()
for t in readers:
t.join()
np.testing.assert_array_equal(f["d"][...], np.full((64, 64), 24.0))
assert not errors, "a read saw a partly written dataset"
def test_close_releases_the_file(h5py, tmp_path):
path = str(tmp_path / "close.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
f = clawhdf5.File(path, "r+")
ds = f["d"]
f.close()
# The handle still reads the file as last written, but cannot edit it.
np.testing.assert_array_equal(ds[...], np.arange(4))
with pytest.raises(OSError, match="closed"):
ds[0] = 1
with h5py.File(path, "r+") as g:
g["d"][0] = 9
def _two_files(h5py, tmp_path):
"""d1/f.h5 and d2/f.h5: same name, different layouts (the review's
repro)."""
(tmp_path / "d1").mkdir()
(tmp_path / "d2").mkdir()
with h5py.File(tmp_path / "d1" / "f.h5", "w") as f:
f.create_dataset("x", data=np.arange(10, dtype="<i4"))
f.create_dataset("big", data=np.full(5000, 1.5))
with h5py.File(tmp_path / "d2" / "f.h5", "w") as f:
f.create_dataset("pad", data=np.full(3000, 2.5))
f.create_dataset("x", data=np.arange(10, dtype="<i4") + 500)
return tmp_path / "d1" / "f.h5", tmp_path / "d2" / "f.h5"
def _held_file_edited(h5py, held, other, other_bytes):
"""The edits went to `held`, planned from its own metadata; `other`
was not touched."""
assert other.read_bytes() == other_bytes, "the other file changed"
with h5py.File(held, "r") as f:
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
np.testing.assert_array_equal(f.attrs["note"], np.arange(50.0))
h5dump_reads(str(held))
def test_relative_path_and_chdir(h5py, tmp_path, monkeypatch):
"""A file opened by a relative path keeps being the file edited and read
after os.chdir (an edit was planned from the file the path named in the
new directory and written into the one open, corrupting it)."""
held, other = _two_files(h5py, tmp_path)
other_bytes = other.read_bytes()
monkeypatch.chdir(held.parent)
f = clawhdf5.File("f.h5", "r+")
ds = f["x"]
monkeypatch.chdir(other.parent)
ds[:] = np.full(10, 7, "<i4")
np.testing.assert_array_equal(ds[:5], np.full(5, 7)) # read from the held file
f.attrs["note"] = np.arange(50.0)
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
assert "pad" not in f
f.close()
_held_file_edited(h5py, held, other, other_bytes)
# 'w' writes where the path named when the file was opened.
monkeypatch.chdir(held.parent)
w = clawhdf5.File("new.h5", "w")
monkeypatch.chdir(other.parent)
w.create_dataset("d", data=np.arange(3))
w.close()
assert (held.parent / "new.h5").exists() and not (other.parent / "new.h5").exists()
def test_path_replaced_between_edits(h5py, tmp_path):
"""The path renamed away and another file put in its place between
edits: the edits go to the file held open, never mixed with the other."""
held, other = _two_files(h5py, tmp_path)
path = str(tmp_path / "f.h5")
os.replace(held, path)
moved = tmp_path / "moved.h5"
with clawhdf5.File(path, "r+") as f:
ds = f["x"]
ds[0] = 7
os.replace(path, moved)
shutil.copy(other, path)
other_bytes = open(path, "rb").read()
ds[:] = np.full(10, 7, "<i4")
f.attrs["note"] = np.arange(50.0)
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
assert "pad" not in f
from pathlib import Path
_held_file_edited(h5py, moved, Path(path), other_bytes)
def test_an_edit_releases_the_gil(h5py, tmp_path):
"""Another Python thread keeps running while a large edit is written:
had the edit held the GIL, the other thread would stall for the whole
edit (Rust code never yields it)."""
import time
path = str(tmp_path / "big.h5")
with h5py.File(path, "w") as f:
f.create_dataset("d", shape=(1024, 1024), dtype="<f8", chunks=(64, 64), compression="gzip")
value = np.random.default_rng(0).standard_normal((1024, 1024))
stamps = []
stop = threading.Event()
def spin():
while not stop.is_set():
stamps.append(time.perf_counter())
with clawhdf5.File(path, "r+") as f:
ds = f["d"]
t = threading.Thread(target=spin)
t.start()
try:
time.sleep(0.05)
t0 = time.perf_counter()
ds[...] = value
t1 = time.perf_counter()
finally:
stop.set()
t.join()
np.testing.assert_array_equal(ds[...], value)
during = [x for x in stamps if t0 <= x <= t1]
gaps = np.diff([t0] + during + [t1])
assert t1 - t0 > 0.1, "the edit is too quick to tell"
assert gaps.max() < 0.5 * (t1 - t0), (
f"the other thread stalled for {gaps.max():.3f} s of a {t1 - t0:.3f} s edit")
+28 -4
View File
@@ -178,15 +178,39 @@ def _write_fixture(h5py, path):
g.attrs["depth"] = np.int8(3)
@pytest.fixture(scope="module")
def pair(h5py, tmp_path_factory):
path = str(tmp_path_factory.mktemp("h5") / "fixture.h5")
@pytest.fixture(scope="module", params=["local", "r+", "http", "http-1k-blocks"])
def pair(request, h5py, tmp_path_factory):
"""The fixture file through h5py and through clawhdf5: opened locally
(read-only, and for editing: a copy, since editing locks the file), and
over HTTP range requests (a local server in this process) with the
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
through many small ranges."""
import shutil
from conftest import RangeServer
root = tmp_path_factory.mktemp("h5")
path = str(root / "fixture.h5")
_write_fixture(h5py, path)
theirs = h5py.File(path, "r")
server = None
if request.param == "local":
ours = clawhdf5.File(path, "r")
elif request.param == "r+":
copy = str(root / "editable.h5")
shutil.copy(path, copy)
ours = clawhdf5.File(copy, "r+")
else:
server = RangeServer(root)
if request.param == "http":
ours = clawhdf5.File(server.url("fixture.h5"))
else:
ours = clawhdf5.File.open_url(server.url("fixture.h5"), block_size=1024)
yield ours, theirs, path
theirs.close()
ours.close()
if server is not None:
server.close()
def _all_datasets(h5py, f):
@@ -416,7 +440,7 @@ def test_unsupported_types_are_errors_not_data(pair):
def test_boolean_masks_are_refused(pair):
ours, _, _ = pair
with pytest.raises(TypeError):
with pytest.raises(NotImplementedError, match="mask"):
ours["num/le_i4_1d"][np.ones(37, dtype=bool)]
+244
View File
@@ -0,0 +1,244 @@
"""Remote files: `clawhdf5.File(url)` / `File.open_url(url, ...)` read over
HTTP range requests (clawhdf5-remote's block cache), against a server in this
process (conftest.RangeServer). Values are compared with h5py reading the
same file locally; the rest checks what the server saw (only the blocks a
read needs are fetched), the failure modes (no range support, a missing
file, a file that changes, a server that goes away: errors, never wrong
data), and that the GIL is released while a read waits on the network."""
import os
import sys
import threading
import time
import numpy as np
import pytest
import clawhdf5
from conftest import RangeServer
def _write(h5py, path):
rng = np.random.default_rng(7)
with h5py.File(path, "w") as f:
f.create_dataset("contig", data=rng.standard_normal((400, 300)))
f.create_dataset(
"chunked",
data=rng.integers(0, 1000, size=(512, 512), dtype="<i4"),
chunks=(64, 64),
compression="gzip",
)
f.create_dataset("strings", data=["alpha", "beta", "gamma"], dtype=h5py.string_dtype())
g = f.create_group("grp")
g.create_dataset("small", data=np.arange(10, dtype="<u2"))
g.attrs["units"] = "m/s"
f.attrs["title"] = "remote test"
@pytest.fixture
def remote_file(h5py, tmp_path, range_server):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
return path, range_server.url("remote.h5"), range_server
def test_remote_reads_match_h5py(h5py, remote_file):
path, url, server = remote_file
with h5py.File(path, "r") as theirs, clawhdf5.File(url) as ours:
assert ours.filename == url
assert list(ours.keys()) == list(theirs.keys())
assert "grp/small" in ours and "nope" not in ours
for name in ["contig", "chunked", "grp/small"]:
np.testing.assert_array_equal(ours[name][...], theirs[name][...])
np.testing.assert_array_equal(ours[name][3:7], theirs[name][3:7])
np.testing.assert_array_equal(ours["chunked"][100:130, 200:300:3], theirs["chunked"][100:130, 200:300:3])
np.testing.assert_array_equal(ours["chunked"][[1, 70, 300], 5], theirs["chunked"][[1, 70, 300], 5])
assert list(ours["strings"][...]) == list(theirs["strings"][...])
assert ours["grp"].attrs["units"] == theirs["grp"].attrs["units"]
assert ours.attrs["title"] == theirs.attrs["title"]
assert ours["chunked"].maxshape == theirs["chunked"].maxshape
assert server.requests() >= 2
def test_a_small_read_fetches_only_its_blocks(h5py, remote_file):
"""With 4 KiB blocks, opening and reading one chunk of a 1 MB chunked
dataset costs a handful of requests and a few blocks, not the file."""
path, url, server = remote_file
size = os.path.getsize(path)
f = clawhdf5.File.open_url(url, block_size=4096)
opened = server.requests()
assert opened == 1, server.log
ds = f["chunked"]
got = ds[0:10, 0:10]
with h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(got, theirs["chunked"][0:10, 0:10])
stats = f.remote_stats
assert stats["bytes_fetched"] < size / 4, (stats, size)
assert server.requests() - opened <= 12, server.log
# A second read of the same region is served by the cache.
before = server.requests()
ds[0:10, 0:10]
assert server.requests() == before
assert f.remote_stats["hits"] > stats["hits"]
assert clawhdf5.File(str(path), "r").remote_stats is None
def test_server_without_range_support(h5py, tmp_path):
"""A server that ignores Range answers 200 with the whole file: that is
an OSError by default, and a whole download when allowed."""
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path, ranges=False)
try:
url = server.url("remote.h5")
with pytest.raises(OSError, match="range"):
clawhdf5.File(url)
with clawhdf5.File.open_url(url, allow_full_download=True) as ours, h5py.File(path, "r") as theirs:
np.testing.assert_array_equal(ours["chunked"][...], theirs["chunked"][...])
np.testing.assert_array_equal(ours["contig"][5], theirs["contig"][5])
with pytest.raises(OSError):
clawhdf5.File.open_url(url, allow_full_download=True, max_full_download=1000)
finally:
server.close()
def test_errors_are_oserrors(remote_file):
_, url, server = remote_file
with pytest.raises(OSError, match="404"):
clawhdf5.File(server.url("missing.h5"))
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "r+")
with pytest.raises(ValueError, match="read-only"):
clawhdf5.File(url, "w")
with pytest.raises(OSError, match="unsupported URL"):
clawhdf5.File("nosuchscheme://x/y.h5")
with pytest.raises(ValueError):
clawhdf5.File.open_url(url, block_size=0)
with pytest.raises(TypeError):
clawhdf5.File.open_url(url, no_such_option=1)
def test_object_store_urls_need_their_features():
"""The default wheel has no S3/GCS/Azure clients (aws-lc-rs builds C):
such a URL is an OSError naming the build feature."""
for url, feature in [("s3://bucket/k.h5", "s3"), ("gs://b/k.h5", "gcs"), ("az://c/k.h5", "azure")]:
try:
clawhdf5.File(url)
except OSError as e:
if "feature" in str(e):
assert f"`{feature}`" in str(e), str(e)
else:
pytest.fail(f"{url} opened")
def test_https_needs_the_https_feature():
"""The default wheel has no TLS stack (rustls needs ring, which builds C):
an https URL is an OSError that names the build feature."""
with pytest.raises(OSError) as e:
clawhdf5.File("https://127.0.0.1:1/x.h5")
msg = str(e.value)
# Built with `--features https` the error is the refused connection.
assert "https" in msg or "connect" in msg.lower() or "refused" in msg.lower(), msg
def test_a_changed_file_is_an_error_not_mixed_data(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=1024)
first = f["grp/small"][...]
# Rewrite the file with other values: new ETag, same name.
time.sleep(0.01)
with h5py.File(path, "w") as g:
g.create_dataset("contig", data=np.zeros((400, 300)))
with pytest.raises(OSError, match="changed"):
f["contig"][...]
np.testing.assert_array_equal(first, np.arange(10, dtype="<u2"))
def test_a_server_that_goes_away_is_an_error(h5py, tmp_path):
path = tmp_path / "remote.h5"
_write(h5py, str(path))
server = RangeServer(tmp_path)
f = clawhdf5.File.open_url(server.url("remote.h5"), block_size=1024, retries=0, timeout=2)
ds = f["contig"]
server.close()
with pytest.raises(OSError):
ds[...]
def test_threads_read_one_remote_file(h5py, remote_file):
path, url, _ = remote_file
f = clawhdf5.File.open_url(url, block_size=2048)
with h5py.File(path, "r") as theirs:
expected = theirs["chunked"][...]
errors = []
def work(i):
try:
rows = slice((i * 37) % 400, (i * 37) % 400 + 64)
np.testing.assert_array_equal(f["chunked"][rows], expected[rows])
except Exception as e: # noqa: BLE001
errors.append(e)
threads = [threading.Thread(target=work, args=(i,)) for i in range(16)]
for t in threads:
t.start()
for t in threads:
t.join()
assert not errors, errors[:3]
def test_remote_reads_release_the_gil(h5py, remote_file):
"""A read waiting on a slow server lets other Python threads run: a
thread counting in a loop keeps counting (and never stalls for long)
while the main thread reads through requests that each take 0.2 s."""
_, url, server = remote_file
f = clawhdf5.File.open_url(url, block_size=1024, max_parallel=1)
ds = f["contig"]
server.delay = 0.2
server.delay_after = server.requests()
old = sys.getswitchinterval()
sys.setswitchinterval(0.001)
stop = threading.Event()
progress = {"n": 0, "worst": 0.0}
def spin():
last = time.perf_counter()
while not stop.is_set():
now = time.perf_counter()
progress["worst"] = max(progress["worst"], now - last)
last = now
progress["n"] += 1
t = threading.Thread(target=spin)
try:
t.start()
time.sleep(0.02)
t0 = time.perf_counter()
before = server.requests()
ds[0:2]
took = time.perf_counter() - t0
stop.set()
t.join()
finally:
sys.setswitchinterval(old)
assert server.requests() > before
assert took >= 0.2, took
assert progress["n"] > 1000, progress
# Held across a 0.2 s request, the spinner would stall that long.
assert progress["worst"] < 0.1, (progress, took)
def test_a_clawhdf5_written_file_reads_the_same_remotely(tmp_path, range_server):
path = tmp_path / "ours.h5"
data = np.arange(3000, dtype="<f8").reshape(100, 30)
with clawhdf5.File(str(path), "w") as f:
f.create_dataset("d", data=data, chunks=[10, 30], compression="gzip")
g = f.create_group("g")
g.create_dataset("i", data=np.arange(5, dtype="<i4"))
f.attrs["k"] = 3
with clawhdf5.File(range_server.url("ours.h5")) as f, clawhdf5.File(str(path)) as local:
np.testing.assert_array_equal(f["d"][...], data)
np.testing.assert_array_equal(f["d"][5:9, ::4], local["d"][5:9, ::4])
np.testing.assert_array_equal(f["g/i"][...], np.arange(5))
assert f.attrs["k"] == local.attrs["k"]
assert "file (read" in repr(f).lower() and "127.0.0.1" in repr(f)
@@ -1509,3 +1509,68 @@ fn unencodable_filters_are_unsupported() {
"a refused edit changed the file"
);
}
fn fixture(dir: &Path, name: &str) -> std::path::PathBuf {
let path = dir.join(name);
std::fs::copy(
Path::new(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name),
&path,
)
.unwrap();
path
}
/// Zero extents on chunked datasets with no recorded maximum. The unfixed
/// editor left `chunk_zero_extent_no_maxshape.h5` (a 2.7.0-written file
/// resized to 1x0): a Fixed Array whose maximum, taken from the current
/// dimensions, has no chunks along one dimension, so every stride before it
/// is 0 — `h5rs check` panicked dividing by it and the next resize failed
/// with an internal error. Such a file must check clean and resize on; a
/// 2.7.0-written file taken through zero extents by the fixed editor must
/// check clean at every step and read the fill value where it grew.
#[test]
fn zero_extent_resizes_without_a_recorded_maximum() {
if !tools_ok() {
return;
}
let dir = tmpdir();
let path = fixture(dir.path(), "chunk_zero_extent_no_maxshape.h5");
check_tools(&path, true);
let mut ed = FileEditor::open(&path).unwrap();
ed.resize("d", &[0, 0]).unwrap();
ed.resize("z", &[0, 0]).unwrap();
// Their maximum is now what the index was laid out by (1 x 0).
ed.resize("d", &[1, 0]).unwrap();
assert!(matches!(
ed.resize("d", &[1, 1]),
Err(Error::InvalidArgument(_))
));
drop(ed);
check_tools(&path, true);
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
for shape in [[15, 15], [3, 2], [1, 1], [1, 0], [0, 0], [0, 20], [20, 20]] {
let mut ed = FileEditor::open(&path).unwrap();
ed.resize("d", &shape).unwrap();
ed.resize("z", &shape).unwrap();
drop(ed);
check_tools(&path, true);
}
let f = File::open(&path).unwrap();
for name in ["d", "z"] {
let d = f.dataset(name).unwrap();
assert_eq!(d.shape().unwrap(), [20, 20]);
assert!(d.read_f32().unwrap().iter().all(|&v| v == 0.0), "{name}");
}
assert_eq!(
py(&format!(
"import h5py\n\
with h5py.File({:?}) as f:\n\
\x20 print(int(abs(f['d'][()]).sum() + abs(f['z'][()]).sum()), f['d'].maxshape)",
path.to_str().unwrap()
)),
"0 (20, 20)"
);
}
+2
View File
@@ -24,6 +24,8 @@ clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
# Must match the wasm-bindgen CLI exactly; build.sh checks.
wasm-bindgen = "0.2.129"
js-sys = "0.3.106"
# Promises for openUrl and RemoteFile (pure Rust over js-sys).
wasm-bindgen-futures = "0.4.79"
[dev-dependencies]
serde_json = "1"
+235
View File
@@ -0,0 +1,235 @@
// HTTP for clawhdf5-wasm's openUrl (see src/lib.rs and src/lazy.rs).
//
// The Rust side decides which byte ranges a read needs; this file fetches
// them with `fetch` and `Range` headers and checks every answer, so a server
// that ignores the range, answers with other bytes, or serves a file that
// changed since it was opened is an error, never data. wasm-bindgen copies
// it into the package (pkg/snippets/...).
const DEFAULT_MAX_DOWNLOAD = 512 * 1024 * 1024;
const DEFAULT_PARALLEL = 6;
function fetcher(opts) {
const f = opts?.fetch ?? globalThis.fetch;
if (typeof f !== "function") {
throw new Error("openUrl: no fetch() in this environment (pass opts.fetch)");
}
return f;
}
// The caller's headers (`opts.headers`: a Headers, [name, value] pairs or a
// plain object, as fetch takes them; names come out in lower case) with
// `extra` over them. A Range of the caller's is dropped: this file asks for
// the ranges.
function init(opts, extra, method = "GET", signal = undefined) {
const headers = {};
if (opts?.headers != null) {
for (const [name, value] of new Headers(opts.headers)) {
if (name !== "range") headers[name] = value;
}
}
Object.assign(headers, extra);
return { method, headers, credentials: opts?.credentials, signal };
}
// `opts.parallel`: range requests in flight at once.
function parallelism(opts) {
const p = opts?.parallel ?? DEFAULT_PARALLEL;
if (!Number.isSafeInteger(p) || p < 1) {
throw new Error(`openUrl: parallel must be a positive integer, got ${String(p)}`);
}
return p;
}
// "bytes a-b/total" -> { start, end (exclusive), total | null }; null when
// the page cannot see the header (cross-origin, not exposed).
function contentRange(resp, url) {
const v = resp.headers.get("Content-Range");
if (v === null) return null;
const m = /^bytes (\d+)-(\d+)\/(\d+|\*)$/.exec(v.trim());
if (!m) throw new Error(`${url}: the server sent an unusable Content-Range: ${v}`);
return { start: Number(m[1]), end: Number(m[2]) + 1, total: m[3] === "*" ? null : Number(m[3]) };
}
// What pins the file: its ETag, else its Last-Modified (null if neither is
// visible to this page).
function validatorOf(resp) {
return resp.headers.get("ETag") ?? resp.headers.get("Last-Modified");
}
async function discard(resp) {
try {
await resp.body?.cancel();
} catch {
// Nothing to release.
}
}
// The body, refusing more than `limit` bytes as they arrive: it is piped
// through a TransformStream that errors the moment the count passes the
// limit, which cancels the body and so aborts the request. Whatever the
// server declares or sends, the page never holds more than `limit` bytes of
// it. (`tooBig(n)` makes the error; n is the count so far.) A reader loop
// would do the same, but it stalls on small bodies in headless Chromium
// under --virtual-time-budget, which the page test uses; a pipe does not.
async function readCapped(resp, limit, tooBig) {
const declared = resp.headers.get("Content-Length");
if (declared !== null && Number(declared) > limit) {
await discard(resp);
throw tooBig(declared);
}
if (!resp.body) {
// No stream to read from (some fetch implementations): all at once.
const all = new Uint8Array(await resp.arrayBuffer());
if (all.length > limit) throw tooBig(all.length);
return all;
}
let n = 0;
let over = null;
const capped = resp.body.pipeThrough(new TransformStream({
transform(chunk, ctl) {
n += chunk.length;
if (n > limit) {
over = tooBig(`over ${limit}`);
ctl.error(over);
return;
}
ctl.enqueue(chunk);
},
}));
try {
return new Uint8Array(await new Response(capped).arrayBuffer());
} catch (e) {
throw over ?? e;
}
}
// The whole body of a 200 answer (a server without range support), at
// most `limit` (maxDownload) bytes.
function readAll(resp, limit, url) {
return readCapped(resp, limit, (n) =>
new Error(`${url} is ${n} bytes, more than maxDownload (${limit}); ` +
"the server does not support range requests, so the whole file would have to be downloaded"));
}
// The body of a 206 answer, which may not be longer than the `limit` bytes
// asked for at `start` (the caller checks the exact length).
function readLimited(resp, limit, url, start) {
return readCapped(resp, limit, () =>
new Error(`${url}: asked for ${limit} bytes at offset ${start}, the server sent more`));
}
/**
* Ask for the file's first `firstLen` bytes. A server that honours the
* range (206) gives `{ length, first, validator, requests }`; one that
* answers 200 sends the whole file, which is kept (`{ whole, requests }`)
* when `opts.fallback` is "download" (the default) and the file is at most
* `opts.maxDownload` bytes, and is an error otherwise.
*/
export async function probe(url, firstLen, opts) {
const f = fetcher(opts);
parallelism(opts);
const resp = await f(url, init(opts, { Range: `bytes=0-${firstLen - 1}` }));
if (resp.status === 206) {
const cr = contentRange(resp, url);
if (cr && cr.start !== 0) {
await discard(resp);
throw new Error(`${url}: asked for bytes from 0, the server sent bytes from ${cr.start}`);
}
const first = await readLimited(resp, firstLen, url, 0);
let length = cr?.total ?? null;
let requests = 1;
if (length === null) {
// Content-Range is not readable here: a cross-origin server that does
// not list it in Access-Control-Expose-Headers. Content-Length of a
// HEAD request is always readable.
const head = await f(url, init(opts, {}, "HEAD"));
requests++;
const cl = head.headers.get("Content-Length");
if (!head.ok || cl === null) {
throw new Error(`${url}: cannot learn the file's size (a cross-origin server must send ` +
"Access-Control-Expose-Headers: Content-Range, or answer HEAD with Content-Length)");
}
length = Number(cl);
}
if (!Number.isSafeInteger(length) || length < 0) {
throw new Error(`${url}: the server gave a file size of ${length} bytes; openUrl reads files ` +
"of up to 2^53 - 1 bytes (the largest offset a JavaScript number holds exactly)");
}
if (first.length !== Math.min(firstLen, length)) {
throw new Error(`${url}: asked for the first ${firstLen} bytes of ${length}, got ${first.length}`);
}
return { length, first, validator: validatorOf(resp), requests };
}
if (resp.status === 200) {
if ((opts?.fallback ?? "download") !== "download") {
await discard(resp);
throw new Error(`${url}: the server does not support HTTP range requests (it answered 200 ` +
"to a Range request); open it with { fallback: \"download\" } to download the whole file");
}
const whole = await readAll(resp, opts?.maxDownload ?? DEFAULT_MAX_DOWNLOAD, url);
return { whole, requests: 1 };
}
await discard(resp);
throw new Error(`${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
}
/**
* Fetch `ranges` ([start0, end0, start1, end1, ...], ends exclusive) of a
* file opened by `probe`, at most `opts.parallel` (default 6) at a time.
* Every answer must be a 206 with exactly the bytes asked for, from the same
* file (validator and length). When one request fails, the others in
* flight are aborted and no more are made; that failure is the error.
*/
export async function fetchRanges(url, ranges, opts, validator, length) {
const f = fetcher(opts);
const parallel = parallelism(opts);
const n = ranges.length / 2;
const out = new Array(n);
const abort = new AbortController();
let next = 0;
async function one(i) {
const start = ranges[2 * i];
const end = ranges[2 * i + 1];
const resp = await f(url, init(opts, { Range: `bytes=${start}-${end - 1}` }, "GET", abort.signal));
if (resp.status !== 206) {
await discard(resp);
throw new Error(resp.status === 200
? `${url}: the server stopped honouring range requests`
: `${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
}
const cr = contentRange(resp, url);
const v = validatorOf(resp);
if ((validator != null && v !== null && v !== validator) ||
(cr?.total != null && cr.total !== length)) {
await discard(resp);
throw new Error(`${url} changed on the server since it was opened`);
}
if (cr && (cr.start !== start || cr.end !== end)) {
await discard(resp);
throw new Error(`${url}: asked for bytes ${start}-${end - 1}, the server sent ${cr.start}-${cr.end - 1}`);
}
const body = await readLimited(resp, end - start, url, start);
if (body.length !== end - start) {
throw new Error(`${url}: asked for ${end - start} bytes at offset ${start}, got ${body.length}`);
}
out[i] = body;
}
async function worker() {
while (next < n && !abort.signal.aborted) {
try {
await one(next++);
} catch (e) {
// The first failure stops the rest: requests in flight are aborted
// (their AbortErrors are not reported) and no new ones start.
if (!abort.signal.aborted) {
abort.abort();
throw e;
}
return;
}
}
}
await Promise.all(Array.from({ length: Math.min(parallel, n) }, worker));
return out;
}
+110 -25
View File
@@ -6,11 +6,25 @@
//! with no typed-array mapping (compound, reference, opaque, ...) is refused
//! with a message naming it, never returned as reinterpreted bytes.
use std::sync::Arc;
use clawhdf5::{AttrValue, File, Selection};
use clawhdf5_format::data_read;
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::storage::Storage;
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
/// The most memory one read may use while it decodes: the stored bytes,
/// the values at 64 bits (integers are widened first) and the values
/// returned. A larger read fails with an error naming `readHyperslab`,
/// before anything is read: on wasm32 a buffer past 2 GiB cannot be
/// allocated at all, and failing to allocate aborts the module (every open
/// file on the page with it). 1 GiB leaves room in wasm32's 4 GiB for the
/// file's cached blocks and the JavaScript copy of the result.
pub const MAX_READ_BYTES: u64 = 1 << 30;
/// Errors are reported to JavaScript as messages.
pub type Result<T> = std::result::Result<T, String>;
@@ -119,7 +133,9 @@ pub struct Attr {
pub value: AttrValue,
}
/// An open file, held in memory.
/// An open file: held in memory ([`Reader::open`]) or read through a
/// [`Storage`] ([`Reader::open_storage`], such as a
/// [`LazyStorage`](crate::lazy::LazyStorage)).
pub struct Reader {
file: File,
}
@@ -132,6 +148,14 @@ impl Reader {
})
}
/// Open a file read through `storage` (the file's bytes from offset 0,
/// user block included, as [`File::open_storage`] takes them).
pub fn open_storage(storage: Arc<dyn Storage + Send + Sync>) -> Result<Self> {
Ok(Self {
file: File::open_storage(storage).map_err(err)?,
})
}
/// Whether `path` names a group or a dataset.
pub fn kind(&self, path: &str) -> Result<Kind> {
match self.file.dataset(path) {
@@ -144,31 +168,54 @@ impl Reader {
/// The groups, then the datasets, in the group at `path` (`/` is the
/// root). Soft links are listed as their targets; external and dangling
/// links, and named datatypes, are left out.
///
/// What [`Group::groups`](clawhdf5::Group::groups) and `datasets` list,
/// but every child's object header is read before an error ends the
/// listing (the first error, in listing order, is the one returned, as
/// there). Over a [`LazyStorage`](crate::lazy::LazyStorage) that makes
/// one pass ask for all the headers it is missing at once, instead of
/// one pass, and one round trip, per header.
pub fn list(&self, path: &str) -> Result<Vec<Child>> {
if self.kind(path)? != Kind::Group {
return Err(format!("not a group: {path}"));
}
let group = self.file.group(path).map_err(err)?;
let mut out: Vec<Child> = group
.groups()
.map_err(err)?
.into_iter()
.map(|name| Child {
name,
let entries = group.entries().map_err(err)?;
let sb = self.file.superblock();
let storage = self.file.storage();
let mut groups = Vec::new();
let mut datasets = Vec::new();
let mut first_error = None;
for (name, address) in entries {
match ObjectHeader::parse_in(storage, address, sb.offset_size, sb.length_size) {
Ok(header) => {
let has = |t: MessageType| header.messages.iter().any(|m| m.msg_type == t);
if has(MessageType::LinkInfo)
|| has(MessageType::Link)
|| has(MessageType::SymbolTable)
{
groups.push(Child {
name: name.clone(),
kind: Kind::Group,
})
.collect();
out.extend(
group
.datasets()
.map_err(err)?
.into_iter()
.map(|name| Child {
});
}
if has(MessageType::DataLayout) {
datasets.push(Child {
name,
kind: Kind::Dataset,
}),
);
Ok(out)
});
}
}
Err(e) => {
first_error.get_or_insert(e);
}
}
}
if let Some(e) = first_error {
return Err(err(clawhdf5::Error::from(e)));
}
groups.extend(datasets);
Ok(groups)
}
/// Shape, max shape and datatype of the dataset at `path`.
@@ -228,13 +275,27 @@ impl Reader {
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
}
let raw = ds.read_selection(&selection).map_err(err)?;
let data = self.decode(&raw, &dt)?;
out_shape.extend(element_shape(&dt));
let expected = out_shape
.iter()
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
.ok_or("selection size overflows")?;
let cost = expected.saturating_mul(bytes_per_value(&dt));
if cost > MAX_READ_BYTES {
return Err(format!(
"reading {path}{} would take about {} MiB of memory, more than the {} MiB \
one read may use; read it in parts (readHyperslab)",
if slab.is_some() {
" (this selection)"
} else {
" whole"
},
cost >> 20,
MAX_READ_BYTES >> 20
));
}
let raw = ds.read_selection(&selection).map_err(err)?;
let data = self.decode(&raw, &dt)?;
if data.len() as u64 != expected {
return Err(format!(
"read {} values for shape {out_shape:?} ({expected} expected)",
@@ -281,11 +342,14 @@ impl Reader {
// string ends at its first NUL and a heap object of the
// wrong size is an error, as in libhdf5 and h5py.
let sb = self.file.superblock();
Data::Strings(
VlResolver::new(self.file.as_bytes(), sb.offset_size, sb.length_size)
.strings(raw)
.map_err(err)?,
)
let strings = match self.file.contiguous_bytes() {
Some(bytes) => {
VlResolver::new(bytes, sb.offset_size, sb.length_size).strings(raw)
}
None => VlResolver::new_in(self.file.storage(), sb.offset_size, sb.length_size)
.strings(raw),
};
Data::Strings(strings.map_err(err)?)
}
Datatype::Enumeration { .. } if !is_array => {
Data::Strings(data_read::read_enum_names(raw, dt).map_err(err)?)
@@ -300,6 +364,27 @@ impl Reader {
}
}
/// Memory one value of type `dt` takes while [`Reader::read`] decodes it
/// (an array type's elements count as values): its stored bytes, plus what
/// [`Reader::decode`] builds from them. A string counts its `String` (24
/// bytes on 64-bit targets, less on wasm32) and, for a fixed-length one,
/// its text; a variable-length string's text lives in the heap and is
/// bounded by the storage's own read limit.
fn bytes_per_value(dt: &Datatype) -> u64 {
let base = array_base(dt);
let stored = u64::from(base.type_size());
stored
+ match base {
Datatype::FloatingPoint { size, .. } if *size <= 4 => 4,
Datatype::FloatingPoint { .. } => 8,
// Widened to 64 bits, then narrowed to a new vector.
Datatype::FixedPoint { .. } => 8 + stored,
Datatype::String { .. } => 24 + stored,
Datatype::VariableLength { .. } | Datatype::Enumeration { .. } => 24,
_ => 0,
}
}
/// Narrow integers read at 64 bits to the dataset's own width. The source is
/// that width, so this cannot fail on correct input; it is checked anyway.
fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T>> {
+815
View File
@@ -0,0 +1,815 @@
//! Reading a file that is not all here, when no read may wait for the
//! network: the restartable "NeedBytes" mode of `docs/design/range-reads.md`
//! (milestone M4).
//!
//! A browser's main thread cannot block on `fetch`, and the parsers are
//! synchronous. So an operation (open, list a group, read a dataset) runs
//! as a *pass* over a [`LazyStorage`] that holds the blocks fetched so far:
//!
//! 1. [`LazyStorage::attempt`] runs the operation. A read whose blocks are
//! all present is served; a read that misses records the missing blocks
//! and fails with a storage error.
//! 2. If the pass missed anything, its result is thrown away — whatever it
//! is, since a parser may have caught the error and carried on (a
//! listing skips a link it cannot resolve) — and the caller gets the
//! byte ranges to fetch ([`Step::Need`]).
//! 3. The caller fetches them (asynchronously, with HTTP `Range` requests),
//! hands them over with [`LazyStorage::supply`] and runs the operation
//! again.
//!
//! A pass is pure over the storage: the facade only caches what completed
//! reads decoded (its chunk cache), so re-running it is safe. Every pass
//! that does not finish asks for at least one block not yet present, and no
//! block is evicted while an operation is in flight
//! ([`LazyStorage::operation`]), so an operation finishes after at most one
//! pass per block it needs. In practice it is one pass per *wave* of
//! misses: a chunked read asks for all the chunks of a batch at once.
//!
//! Blocks are kept in an LRU cache with a byte budget, trimmed only when no
//! operation is in flight. Blocks fetched for bulk reads (raw data: a
//! `read_ranges` call, or a read longer than a block) go first, so reading
//! a large dataset does not evict the metadata.
use std::borrow::Cow;
use std::collections::{BTreeSet, HashMap};
use std::ops::Range;
use std::sync::{Arc, Mutex, MutexGuard};
use clawhdf5_format::error::FormatError;
use clawhdf5_format::storage::Storage;
/// Default block size: 1 MiB, as `clawhdf5-remote`'s block cache (the size
/// `docs/design/range-reads.md` §2 measured).
pub const DEFAULT_BLOCK_SIZE: u64 = 1 << 20;
/// Default of [`LazyConfig::max_fetch`]: 512 MiB, the same as `openUrl`'s
/// `maxDownload` for a server without range support.
pub const DEFAULT_MAX_FETCH: u64 = 512 << 20;
/// The message of the error a read that misses returns. It never reaches
/// the caller of [`LazyStorage::attempt`]: a pass that missed is re-run.
pub const NEED_BYTES: &str = "bytes not fetched yet (restartable read)";
/// Settings of a [`LazyStorage`].
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct LazyConfig {
/// Size of a block in bytes (at least 512); fetches are whole, aligned
/// blocks (the file's last block is shorter).
pub block_size: u64,
/// Byte budget of cached blocks between operations. An operation keeps
/// every block it needs until it finishes, whatever the budget.
pub capacity: u64,
/// Largest single range asked for, in bytes (whole blocks, at least
/// one); longer runs are split so they can be fetched in parallel.
pub max_request: u64,
/// Most bytes one operation may fetch (at least one block), and so the
/// longest single read: a read longer than this fails at once, before
/// anything is fetched, and so does an operation whose passes would
/// fetch more. The file's length comes from the server, so without
/// this a hostile file (a heap "collection" claiming 2 GiB) makes the
/// reader fetch and hold whatever it names; on wasm32 a buffer past
/// 2 GiB cannot even be allocated.
pub max_fetch: u64,
}
impl Default for LazyConfig {
fn default() -> Self {
LazyConfig {
block_size: DEFAULT_BLOCK_SIZE,
capacity: 64 << 20,
max_request: 8 << 20,
max_fetch: DEFAULT_MAX_FETCH,
}
}
}
/// What a [`LazyStorage`] has done so far.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub struct LazyStats {
/// Passes run by [`LazyStorage::attempt`].
pub passes: u64,
/// Ranges handed to [`LazyStorage::supply`]: one HTTP request each.
pub requests: u64,
/// Bytes handed to [`LazyStorage::supply`].
pub bytes_fetched: u64,
/// Blocks evicted to stay within the budget.
pub evictions: u64,
/// Bytes cached now.
pub cached_bytes: u64,
}
/// The outcome of one pass.
#[derive(Debug)]
pub enum Step<T> {
/// The pass read only bytes that were present: its result stands.
Done(T),
/// The pass missed: fetch these byte ranges (sorted, disjoint, block
/// aligned), [`supply`](LazyStorage::supply) them and run it again.
Need(Vec<Range<u64>>),
}
struct Block {
data: Arc<[u8]>,
/// Eviction order: bulk blocks (`false`) before metadata (`true`),
/// then least recently used first.
key: (bool, u64),
}
#[derive(Default)]
struct State {
blocks: HashMap<u64, Block>,
/// `(metadata?, tick, block index)`, in eviction order.
order: BTreeSet<(bool, u64, u64)>,
tick: u64,
bytes: u64,
/// Blocks the current pass missed, and whether a small read wanted
/// them (metadata).
missing: HashMap<u64, bool>,
/// Blocks a bulk read missed that have not been supplied yet: kept
/// as bulk when they arrive.
bulk_pending: BTreeSet<u64>,
/// Operations in flight: no eviction while any is.
active: u32,
stats: LazyStats,
}
/// A [`Storage`] over the blocks of a file fetched so far; a read of
/// anything else fails and is recorded, so the pass can be re-run once the
/// bytes arrive. See the [module documentation](self).
pub struct LazyStorage {
len: u64,
config: LazyConfig,
state: Mutex<State>,
}
fn lock(m: &Mutex<State>) -> MutexGuard<'_, State> {
m.lock().unwrap_or_else(std::sync::PoisonError::into_inner)
}
/// Keeps an operation's blocks cached until it is dropped; see
/// [`LazyStorage::operation`].
pub struct Operation<'a> {
storage: &'a LazyStorage,
/// Bytes fetched for this operation so far.
fetched: std::cell::Cell<u64>,
}
impl Operation<'_> {
/// Count `ranges` against the operation's budget
/// ([`LazyConfig::max_fetch`]) before they are fetched: an error, and
/// nothing counted, if they would take it past the budget.
pub fn charge(&self, ranges: &[Range<u64>]) -> Result<(), String> {
let max = self.storage.config.max_fetch;
let total = ranges.iter().fold(self.fetched.get(), |n, r| {
n.saturating_add(r.end.saturating_sub(r.start))
});
if total > max {
return Err(format!(
"this call would fetch more than {max} bytes of the file (the maxFetch limit); \
read less at a time (readHyperslab) or raise maxFetch"
));
}
self.fetched.set(total);
Ok(())
}
}
impl Drop for Operation<'_> {
fn drop(&mut self) {
let mut st = lock(&self.storage.state);
st.active = st.active.saturating_sub(1);
if st.active == 0 {
self.storage.evict(&mut st);
}
}
}
impl LazyStorage {
/// An empty cache for a file of `len` bytes.
pub fn new(len: u64, mut config: LazyConfig) -> Self {
config.block_size = config.block_size.max(512);
config.max_request = (config.max_request / config.block_size).max(1) * config.block_size;
config.max_fetch = config.max_fetch.max(config.block_size);
LazyStorage {
len,
config,
state: Mutex::new(State::default()),
}
}
/// The settings in use (after rounding).
pub fn config(&self) -> &LazyConfig {
&self.config
}
/// Counters since the storage was made.
pub fn stats(&self) -> LazyStats {
let st = lock(&self.state);
LazyStats {
cached_bytes: st.bytes,
..st.stats
}
}
/// Mark an operation in flight until the guard is dropped: no block is
/// evicted meanwhile, so re-running its passes always makes progress.
/// Hold it across every pass of one operation.
pub fn operation(&self) -> Operation<'_> {
lock(&self.state).active += 1;
Operation {
storage: self,
fetched: std::cell::Cell::new(0),
}
}
/// Run one pass of `f` over this storage. `Done` when `f` read nothing
/// that is missing; otherwise `Need` with the ranges to fetch, and `f`'s
/// result is dropped (it may be an error caused by the miss, or a
/// result built around one).
pub fn attempt<T>(&self, f: impl FnOnce() -> T) -> Step<T> {
{
let mut st = lock(&self.state);
st.missing.clear();
st.stats.passes += 1;
}
let out = f();
let missing = std::mem::take(&mut lock(&self.state).missing);
if missing.is_empty() {
return Step::Done(out);
}
drop(out);
Step::Need(self.runs(missing))
}
/// The bytes of the file at `offset`, fetched for a range a pass asked
/// for. `offset` must be block aligned and the bytes whole blocks (the
/// last block of the file may be short) inside the file, or this is an
/// error and nothing is kept. Blocks already present are left alone.
pub fn supply(&self, offset: u64, bytes: &[u8]) -> Result<(), String> {
let bs = self.config.block_size;
let end = offset
.checked_add(bytes.len() as u64)
.filter(|&e| e <= self.len)
.ok_or_else(|| {
format!(
"{} bytes at offset {offset} run past the end of the {}-byte file",
bytes.len(),
self.len
)
})?;
if !offset.is_multiple_of(bs) || (!end.is_multiple_of(bs) && end != self.len) {
return Err(format!(
"{} bytes at offset {offset} are not whole {bs}-byte blocks",
bytes.len()
));
}
let mut st = lock(&self.state);
st.stats.requests += 1;
st.stats.bytes_fetched += bytes.len() as u64;
let mut start = offset;
while start < end {
let i = start / bs;
let stop = (start + bs).min(end);
if !st.blocks.contains_key(&i) {
let rel = (start - offset) as usize..(stop - offset) as usize;
let metadata = !st.bulk_pending.remove(&i);
self.keep(&mut st, i, Arc::from(&bytes[rel]), metadata);
}
start = stop;
}
if st.active == 0 {
self.evict(&mut st);
}
Ok(())
}
/// [`supply`](Self::supply) the bytes fetched for `range`, one of the
/// ranges a [`Step::Need`] asked for: anything but exactly its length
/// (a server that answered with more or less) is an error.
pub fn supply_range(&self, range: &Range<u64>, bytes: &[u8]) -> Result<(), String> {
let want = range.end.saturating_sub(range.start);
if bytes.len() as u64 != want {
return Err(format!(
"asked for {want} bytes at offset {}, got {}",
range.start,
bytes.len()
));
}
self.supply(range.start, bytes)
}
/// Run `f` to completion, fetching what its passes miss with `fetch`
/// (a byte range to its bytes). The blocking driver, for native code
/// and tests; the browser's is the same loop with an `await` between
/// passes.
pub fn run_blocking<T>(
&self,
mut f: impl FnMut() -> T,
mut fetch: impl FnMut(Range<u64>) -> Result<Vec<u8>, String>,
) -> Result<T, String> {
let op = self.operation();
loop {
match self.attempt(&mut f) {
Step::Done(v) => return Ok(v),
Step::Need(ranges) => {
op.charge(&ranges)?;
for r in ranges {
let bytes = fetch(r.clone())?;
self.supply_range(&r, &bytes)?;
}
}
}
}
}
/// Cache block `i`.
fn keep(&self, st: &mut State, i: u64, data: Arc<[u8]>, metadata: bool) {
st.tick += 1;
let key = (metadata, st.tick);
st.bytes += <[u8]>::len(&data) as u64;
st.order.insert((key.0, key.1, i));
if let Some(old) = st.blocks.insert(i, Block { data, key }) {
st.order.remove(&(old.key.0, old.key.1, i));
st.bytes -= <[u8]>::len(&old.data) as u64;
}
}
fn evict(&self, st: &mut State) {
while st.bytes > self.config.capacity {
let Some((_, _, i)) = st.order.pop_first() else {
break;
};
if let Some(b) = st.blocks.remove(&i) {
st.bytes -= <[u8]>::len(&b.data) as u64;
st.stats.evictions += 1;
}
}
}
/// Byte ranges covering the missing blocks: runs of consecutive
/// blocks, a one-block hole between two runs filled so they merge
/// (unless the hole is cached: it would be fetched again), each at
/// most `max_request` long.
fn runs(&self, missing: HashMap<u64, bool>) -> Vec<Range<u64>> {
let bs = self.config.block_size;
let mut wanted: Vec<u64> = missing.keys().copied().collect();
wanted.sort_unstable();
let mut st = lock(&self.state);
// Remember which blocks only bulk reads asked for: they are kept
// as bulk once supplied.
for (&i, &metadata) in &missing {
if metadata {
st.bulk_pending.remove(&i);
} else {
st.bulk_pending.insert(i);
}
}
let per_request = self.config.max_request / bs;
let mut runs: Vec<(u64, u64)> = Vec::new();
for i in wanted {
match runs.last_mut() {
Some((first, last))
if (i == *last + 1
|| (i == *last + 2 && !st.blocks.contains_key(&(i - 1))))
&& i - *first < per_request =>
{
*last = i
}
_ => runs.push((i, i)),
}
}
drop(st);
runs.into_iter()
.map(|(a, b)| a * bs..((b + 1) * bs).min(self.len))
.collect()
}
/// Block indices covering `[offset, offset + len)`, clamped to the file.
fn span(&self, offset: u64, len: u64) -> Option<Range<u64>> {
let end = offset.saturating_add(len).min(self.len);
if offset >= end {
return None;
}
let bs = self.config.block_size;
Some(offset / bs..(end - 1) / bs + 1)
}
/// The blocks of `spans` if all are present (touching them), else
/// record the missing ones and fail.
fn blocks(
&self,
spans: &[Range<u64>],
metadata: bool,
) -> Result<HashMap<u64, Arc<[u8]>>, FormatError> {
let mut st = lock(&self.state);
let mut have = HashMap::new();
let mut missed = false;
for span in spans {
for i in span.clone() {
if have.contains_key(&i) {
continue;
}
match st.blocks.get(&i) {
Some(b) => {
have.insert(i, b.data.clone());
}
None => {
missed = true;
let m = st.missing.entry(i).or_insert(metadata);
*m |= metadata;
}
}
}
}
if missed {
return Err(FormatError::Storage(NEED_BYTES.into()));
}
// Touch: most recently used last; a small read promotes a bulk
// block to metadata.
for &i in have.keys() {
st.tick += 1;
let tick = st.tick;
let Some(b) = st.blocks.get_mut(&i) else {
continue;
};
let old = b.key;
b.key = (old.0 || metadata, tick);
let new = b.key;
st.order.remove(&(old.0, old.1, i));
st.order.insert((new.0, new.1, i));
}
Ok(have)
}
/// Refuse a read of `n` bytes longer than an operation may fetch
/// ([`LazyConfig::max_fetch`]), before its blocks are asked for.
fn check_len(&self, n: u64) -> Result<(), FormatError> {
let max = self.config.max_fetch;
if n > max {
return Err(FormatError::Storage(format!(
"a read of {n} bytes is more than one call may fetch ({max} bytes, the maxFetch limit)"
)));
}
Ok(())
}
/// The bytes `offset..end` from `blocks`, which hold every block of
/// that span. The buffer is reserved fallibly: a length the address
/// space cannot hold (past `isize::MAX` on wasm32) is an error, never
/// an abort.
fn assemble(
&self,
offset: u64,
end: u64,
blocks: &HashMap<u64, Arc<[u8]>>,
) -> Result<Vec<u8>, FormatError> {
let bs = self.config.block_size;
let n = end - offset;
let too_long =
|| FormatError::Storage(format!("cannot hold a read of {n} bytes in memory"));
let mut out = Vec::new();
out.try_reserve_exact(usize::try_from(n).map_err(|_| too_long())?)
.map_err(|_| too_long())?;
let mut pos = offset;
while pos < end {
let i = pos / bs;
let block = &blocks[&i];
let from = (pos - i * bs) as usize;
let to = ((end - i * bs) as usize).min(<[u8]>::len(block));
out.extend_from_slice(&block[from..to]);
pos = i * bs + to as u64;
}
Ok(out)
}
}
impl Storage for LazyStorage {
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
let Some(span) = self.span(offset, len as u64) else {
return Ok(Cow::Owned(Vec::new()));
};
let end = offset.saturating_add(len as u64).min(self.len);
self.check_len(end - offset)?;
let metadata = len as u64 <= self.config.block_size;
let blocks = self.blocks(std::slice::from_ref(&span), metadata)?;
Ok(Cow::Owned(self.assemble(offset, end, &blocks)?))
}
fn len(&self) -> u64 {
self.len
}
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
let mut spans = Vec::with_capacity(ranges.len());
let mut total = 0u64;
for r in ranges {
if r.end < r.start {
return Err(FormatError::Storage(
"read range ends before it starts".into(),
));
}
total = total.saturating_add(r.end.min(self.len).saturating_sub(r.start));
spans.extend(self.span(r.start, r.end - r.start));
}
self.check_len(total)?;
let blocks = self.blocks(&spans, false)?;
ranges
.iter()
.map(|r| {
let end = r.end.min(self.len);
if r.start >= end {
Ok(Cow::Owned(Vec::new()))
} else {
self.assemble(r.start, end, &blocks).map(Cow::Owned)
}
})
.collect()
}
}
#[cfg(test)]
#[allow(clippy::single_range_in_vec_init)]
mod tests {
use super::*;
fn file(n: usize) -> Vec<u8> {
(0..n).map(|i| (i * 7 + i / 251) as u8).collect()
}
fn config(block: u64, capacity: u64) -> LazyConfig {
LazyConfig {
block_size: block,
capacity,
max_request: 4 * block,
max_fetch: DEFAULT_MAX_FETCH,
}
}
/// Supply every range of `need` from `data`.
fn serve(s: &LazyStorage, data: &[u8], need: &[Range<u64>]) {
for r in need {
s.supply_range(r, &data[r.start as usize..r.end as usize])
.unwrap();
}
}
fn owned(r: Result<Cow<'_, [u8]>, FormatError>) -> Result<Vec<u8>, FormatError> {
r.map(Cow::into_owned)
}
#[test]
fn a_miss_asks_for_whole_blocks_then_the_rerun_reads_them() {
let data = file(10_000);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let Step::Need(need) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
panic!("nothing is cached yet");
};
assert_eq!(need, vec![1024..3072]);
serve(&s, &data, &need);
let Step::Done(got) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
panic!("the blocks were supplied");
};
assert_eq!(got.unwrap(), &data[1500..2500]);
// Past the end: short, then empty, as for a slice.
serve(&s, &data, &[9216..10_000]);
let Step::Done(tail) = s.attempt(|| owned(s.read_at(9_990, 100))) else {
panic!("the last block was supplied");
};
assert_eq!(tail.unwrap(), &data[9_990..]);
assert!(matches!(
s.attempt(|| s.read_at(20_000, 10).map(|c| c.len())),
Step::Done(Ok(0))
));
let st = s.stats();
assert_eq!(
(st.passes, st.requests, st.bytes_fetched),
(4, 2, 2048 + 784)
);
}
#[test]
fn a_pass_that_swallowed_the_miss_is_still_rerun() {
// A parser that catches the error and returns something anyway
// (a listing skipping a link it cannot resolve) must not have its
// result used.
let data = file(4096);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let step = s.attempt(|| s.read_at(0, 4).map(|b| b.len()).unwrap_or(0));
assert!(
matches!(step, Step::Need(ref n) if n == &vec![0..1024]),
"{step:?}"
);
}
#[test]
fn read_ranges_asks_for_every_missing_block_in_one_pass() {
let data = file(64 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let ranges = [
100..200,
5000..5100,
5200..5300,
30_000..33_000,
60_000..60_010,
];
let read = || {
s.read_ranges(&ranges)
.map(|v| v.into_iter().map(Cow::into_owned).collect::<Vec<_>>())
};
let Step::Need(need) = s.attempt(read) else {
panic!("nothing is cached yet");
};
// 5000..5300 is blocks 4 and 5; 30_000..33_000 is blocks 29..=32.
assert_eq!(
need,
vec![
0..1024,
4096..6144,
29 * 1024..33 * 1024,
58 * 1024..59 * 1024
]
);
serve(&s, &data, &need);
let Step::Done(got) = s.attempt(read) else {
panic!("one pass fetched everything");
};
for (r, g) in ranges.iter().zip(&got.unwrap()) {
assert_eq!(g, &data[r.start as usize..r.end as usize]);
}
}
#[test]
fn runs_merge_one_block_holes_and_split_long_runs() {
let data = file(32 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
// Blocks 0 and 2 (a hole of one: merged), 5..=14 (split in fours).
let Step::Need(need) = s.attempt(|| {
let _ = s.read_at(0, 10);
let _ = s.read_at(2048, 10);
s.read_ranges(&[5120..15 * 1024]).map(|_| ())
}) else {
panic!("nothing is cached yet");
};
assert_eq!(
need,
vec![0..3072, 5120..9216, 9216..13_312, 13_312..15_360]
);
}
#[test]
fn a_cached_hole_is_not_fetched_again() {
let data = file(8 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
serve(&s, &data, &[1024..2048]);
// Blocks 0 and 2 missing, 1 cached: two requests, not 0..3072.
let Step::Need(need) = s.attempt(|| {
let _ = s.read_at(0, 10);
s.read_at(2048, 10).map(|_| ())
}) else {
panic!("blocks 0 and 2 are missing");
};
assert_eq!(need, vec![0..1024, 2048..3072]);
}
#[test]
fn supply_refuses_what_was_not_asked_for() {
let data = file(10_000);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
assert!(s.supply(1, &data[1..1025]).unwrap_err().contains("whole"));
assert!(s.supply(0, &data[..1000]).unwrap_err().contains("whole"));
assert!(
s.supply(9216, &[0u8; 1024])
.unwrap_err()
.contains("past the end")
);
for wrong in [&data[..1023], &data[..2048]] {
assert!(
s.supply_range(&(0..1024), wrong)
.unwrap_err()
.contains("asked for 1024 bytes")
);
}
assert_eq!(s.stats().cached_bytes, 0);
// The file's last, short block is whole.
s.supply(9216, &data[9216..]).unwrap();
assert_eq!(s.stats().cached_bytes, 784);
}
#[test]
fn an_operation_keeps_its_blocks_whatever_the_budget() {
// A budget of one block, an operation that needs eight: without
// the operation guard every pass would evict what the last one
// fetched and never finish.
let data = file(8 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1024));
let got = s
.run_blocking(
|| {
(0..8)
.map(|i| owned(s.read_at(i * 1024 + 3, 10)))
.collect::<Result<Vec<_>, _>>()
},
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap()
.unwrap();
for (i, g) in got.iter().enumerate() {
assert_eq!(g, &data[i * 1024 + 3..i * 1024 + 13]);
}
// Trimmed to the budget once the operation is over.
let st = s.stats();
assert_eq!(st.cached_bytes, 1024);
assert_eq!(st.evictions, 7);
assert_eq!(st.passes, 9, "one pass per block, then the one that ends");
}
#[test]
fn bulk_blocks_are_evicted_before_metadata() {
let data = file(16 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 4 * 1024));
let fetch = |r: Range<u64>| Ok(data[r.start as usize..r.end as usize].to_vec());
// Metadata: a small read of block 0.
s.run_blocking(|| s.read_at(0, 16).map(|_| ()), fetch)
.unwrap()
.unwrap();
// Bulk: raw data over blocks 4..12, more than the budget.
s.run_blocking(|| s.read_ranges(&[4096..12 * 1024]).map(|_| ()), fetch)
.unwrap()
.unwrap();
// The metadata block survived: reading it again is a hit.
let before = s.stats();
assert!(before.cached_bytes <= 4 * 1024);
assert!(matches!(
s.attempt(|| s.read_at(0, 16).map(|_| ())),
Step::Done(Ok(()))
));
assert_eq!(s.stats().requests, before.requests);
}
#[test]
fn a_read_longer_than_max_fetch_fails_without_fetching() {
// A hostile file names a 2 GiB heap collection in a "file" the
// server claims is 1 TiB: the read is refused before any block is
// asked for (on wasm32 its buffer could not even be allocated).
let s = LazyStorage::new(1 << 40, LazyConfig::default());
let step = s.attempt(|| s.read_at(4096, (1usize << 31) + 4096).map(|b| b.len()));
match step {
Step::Done(Err(e)) => assert!(e.to_string().contains("maxFetch"), "{e}"),
other => panic!("expected a refusal, got {other:?}"),
}
let step = s.attempt(|| s.read_ranges(&[0..(600 << 20)]).map(|v| v.len()));
assert!(matches!(step, Step::Done(Err(_))), "{step:?}");
assert_eq!(s.stats().requests, 0);
// At the limit it is an ordinary miss.
let s = LazyStorage::new(1 << 40, config(1024, 1 << 20));
let step = s.attempt(|| s.read_at(0, DEFAULT_MAX_FETCH as usize).map(|b| b.len()));
assert!(matches!(step, Step::Need(_)), "{step:?}");
}
#[test]
fn an_operation_stops_at_its_fetch_budget() {
// Many small reads, none over the limit, that together would fetch
// more than the budget: the operation fails before fetching past it.
let data = file(64 * 1024);
let mut c = config(1024, 1 << 20);
c.max_fetch = 8 * 1024;
let s = LazyStorage::new(data.len() as u64, c);
let e = s
.run_blocking(
|| {
(0..64)
.map(|i| owned(s.read_at(i * 1024, 8)))
.collect::<Result<Vec<_>, _>>()
},
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap_err();
assert!(e.contains("maxFetch"), "{e}");
assert!(s.stats().bytes_fetched <= 8 * 1024, "{:?}", s.stats());
// Within the budget it completes, and the budget is per operation.
for _ in 0..3 {
s.run_blocking(
|| owned(s.read_at(10 * 1024, 3000)),
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap()
.unwrap();
}
}
#[test]
fn a_failed_fetch_is_an_error_not_data() {
let data = file(4096);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let e = s
.run_blocking(|| owned(s.read_at(0, 8)), |_| Err("HTTP 500".into()))
.unwrap_err();
assert_eq!(e, "HTTP 500");
let e = s
.run_blocking(|| owned(s.read_at(0, 8)), |_| Ok(vec![0; 10]))
.unwrap_err();
assert!(e.contains("asked for 1024 bytes"), "{e}");
assert_eq!(s.stats().cached_bytes, 0);
drop(data);
}
}
+473 -59
View File
@@ -1,7 +1,7 @@
//! clawhdf5's HDF5 reader for JavaScript, via `wasm-bindgen`.
//!
//! ```js
//! import init, { open } from "./pkg/clawhdf5_wasm.js";
//! import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
//! await init();
//! const file = open(new Uint8Array(await blob.arrayBuffer()));
//! file.list("/"); // [{ name, kind: "group" | "dataset" }]
@@ -10,6 +10,12 @@
//! file.read("/x"); // { shape, dtype, data: Float64Array | ... | string[] }
//! file.readHyperslab("/x", [0, 0], [10, 10]); // stride, block optional
//! file.free();
//!
//! // A file on a web server, read by HTTP range requests as needed: the
//! // same methods, returning promises.
//! const remote = await openUrl("https://example.org/data.h5");
//! await remote.list("/");
//! remote.stats(); // { requests, bytesFetched, size, ... }
//! ```
//!
//! Numeric data comes back in the typed array of the stored width
@@ -18,15 +24,23 @@
//! Anything else is a thrown `Error` naming the datatype. Only the reader is
//! exposed: nothing here writes files.
//!
//! The logic lives in [`core`], which is plain Rust and tested natively.
//! The logic lives in [`core`] and [`lazy`], which are plain Rust and tested
//! natively; `js/remote.js` does the HTTP.
pub mod core;
pub mod lazy;
use std::ops::Range;
use std::rc::Rc;
use std::sync::Arc;
use clawhdf5::AttrValue;
use js_sys::{Array, Object, Reflect};
use js_sys::{Array, Object, Promise, Reflect, Uint8Array};
use wasm_bindgen::prelude::*;
use wasm_bindgen_futures::future_to_promise;
use crate::core::{Data, Hyperslab, Reader};
use crate::core::{Attr, Child, Data, DatasetInfo, Hyperslab, Reader};
use crate::lazy::{LazyConfig, LazyStorage, Step};
/// JavaScript numbers are exact up to 2^53.
const MAX_SAFE_INTEGER: f64 = 9_007_199_254_740_991.0;
@@ -35,11 +49,30 @@ fn js_err(msg: String) -> JsError {
JsError::new(&msg)
}
/// A JavaScript exception (from `fetch`, or `js/remote.js`) as a `JsError`
/// with its message.
fn js_exception(e: JsValue) -> JsError {
let msg = e
.dyn_ref::<js_sys::Error>()
.map(|e| String::from(e.message()))
.or_else(|| e.as_string())
.unwrap_or_else(|| format!("{e:?}"));
JsError::new(&msg)
}
fn set(obj: &Object, key: &str, value: impl Into<JsValue>) {
// Defining a property on a fresh plain object cannot fail.
Reflect::set(obj, &JsValue::from_str(key), &value.into()).unwrap_throw();
}
fn get(obj: &JsValue, key: &str) -> JsValue {
if obj.is_object() {
Reflect::get(obj, &JsValue::from_str(key)).unwrap_or(JsValue::UNDEFINED)
} else {
JsValue::UNDEFINED
}
}
fn shape_to_js(shape: &[u64]) -> Array {
shape.iter().map(|&d| JsValue::from_f64(d as f64)).collect()
}
@@ -58,6 +91,20 @@ fn indices_from_js(what: &str, v: &[f64]) -> Result<Vec<u64>, JsError> {
.collect()
}
fn slab_from_js(
start: &[f64],
count: &[f64],
stride: Option<Vec<f64>>,
block: Option<Vec<f64>>,
) -> Result<Hyperslab, JsError> {
Ok(Hyperslab {
start: indices_from_js("start", start)?,
count: indices_from_js("count", count)?,
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
})
}
fn data_to_js(data: Data) -> JsValue {
match data {
Data::F32(v) => js_sys::Float32Array::from(&v[..]).into(),
@@ -110,6 +157,72 @@ fn attr_to_js(value: AttrValue) -> (JsValue, Option<String>) {
}
}
fn list_to_js(children: Vec<Child>) -> Array {
children
.into_iter()
.map(|c| {
let o = Object::new();
set(&o, "name", c.name);
set(&o, "kind", c.kind.as_str());
JsValue::from(o)
})
.collect()
}
fn info_to_js(i: DatasetInfo) -> Object {
let o = Object::new();
set(&o, "shape", shape_to_js(&i.shape));
let max: JsValue = match i.maxshape {
None => JsValue::NULL,
Some(dims) => dims
.into_iter()
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
.collect::<Array>()
.into(),
};
set(&o, "maxshape", max);
set(&o, "dtype", i.dtype);
set(&o, "elementShape", shape_to_js(&i.element_shape));
o
}
fn attrs_to_js(attrs: Vec<Attr>) -> Array {
attrs
.into_iter()
.map(|a| {
let o = Object::new();
set(&o, "name", a.name);
let (value, dtype) = attr_to_js(a.value);
set(&o, "value", value);
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
JsValue::from(o)
})
.collect()
}
fn errors_to_js(errors: Vec<String>) -> Array {
errors.into_iter().map(JsValue::from).collect()
}
/// A read's values with the dataset's datatype.
fn values_to_js((dtype, a): (String, core::Array)) -> Object {
let o = Object::new();
set(&o, "shape", shape_to_js(&a.shape));
set(&o, "dtype", dtype);
set(&o, "data", data_to_js(a.data));
o
}
/// Read the dataset at `path` (whole, or `slab`) with its datatype.
fn read_values(
r: &Reader,
path: &str,
slab: Option<&Hyperslab>,
) -> core::Result<(String, core::Array)> {
let dtype = r.info(path)?.dtype;
Ok((dtype, r.read(path, slab)?))
}
/// An open HDF5 (or NetCDF-4) file.
#[wasm_bindgen]
pub struct H5File {
@@ -145,38 +258,13 @@ impl H5File {
/// The group's members: `[{ name, kind }]`, groups first.
pub fn list(&self, path: &str) -> Result<Array, JsError> {
Ok(self
.inner
.list(path)
.map_err(js_err)?
.into_iter()
.map(|c| {
let o = Object::new();
set(&o, "name", c.name);
set(&o, "kind", c.kind.as_str());
JsValue::from(o)
})
.collect())
Ok(list_to_js(self.inner.list(path).map_err(js_err)?))
}
/// `{ shape, maxshape, dtype, elementShape }`. `maxshape` is `null`
/// when not recorded, with `null` for each unlimited dimension.
pub fn info(&self, path: &str) -> Result<Object, JsError> {
let i = self.inner.info(path).map_err(js_err)?;
let o = Object::new();
set(&o, "shape", shape_to_js(&i.shape));
let max: JsValue = match i.maxshape {
None => JsValue::NULL,
Some(dims) => dims
.into_iter()
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
.collect::<Array>()
.into(),
};
set(&o, "maxshape", max);
set(&o, "dtype", i.dtype);
set(&o, "elementShape", shape_to_js(&i.element_shape));
Ok(o)
Ok(info_to_js(self.inner.info(path).map_err(js_err)?))
}
/// `[{ name, value, dtype }]`, sorted by name. Scalars are `number`
@@ -185,31 +273,21 @@ impl H5File {
/// and its `dtype`; one that could not be read at all is reported by
/// [`attrErrors`](Self::attr_errors).
pub fn attrs(&self, path: &str) -> Result<Array, JsError> {
let (attrs, _) = self.inner.attrs(path).map_err(js_err)?;
Ok(attrs
.into_iter()
.map(|a| {
let o = Object::new();
set(&o, "name", a.name);
let (value, dtype) = attr_to_js(a.value);
set(&o, "value", value);
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
JsValue::from(o)
})
.collect())
Ok(attrs_to_js(self.inner.attrs(path).map_err(js_err)?.0))
}
/// Messages for attributes that could not be read.
#[wasm_bindgen(js_name = attrErrors)]
pub fn attr_errors(&self, path: &str) -> Result<Array, JsError> {
let (_, errors) = self.inner.attrs(path).map_err(js_err)?;
Ok(errors.into_iter().map(JsValue::from).collect())
Ok(errors_to_js(self.inner.attrs(path).map_err(js_err)?.1))
}
/// The whole dataset: `{ shape, dtype, data }`, `data` in row-major
/// order.
pub fn read(&self, path: &str) -> Result<Object, JsError> {
self.read_impl(path, None)
Ok(values_to_js(
read_values(&self.inner, path, None).map_err(js_err)?,
))
}
/// A regular hyperslab (`H5Sselect_hyperslab`): `stride` and `block`
@@ -223,22 +301,358 @@ impl H5File {
stride: Option<Vec<f64>>,
block: Option<Vec<f64>>,
) -> Result<Object, JsError> {
let slab = Hyperslab {
start: indices_from_js("start", &start)?,
count: indices_from_js("count", &count)?,
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
let slab = slab_from_js(&start, &count, stride, block)?;
Ok(values_to_js(
read_values(&self.inner, path, Some(&slab)).map_err(js_err)?,
))
}
}
// ---------------------------------------------------------------------------
// Remote files: openUrl.
#[wasm_bindgen(module = "/js/remote.js")]
extern "C" {
#[wasm_bindgen(catch)]
async fn probe(url: &str, first_len: f64, opts: &JsValue) -> Result<JsValue, JsValue>;
#[wasm_bindgen(catch, js_name = fetchRanges)]
async fn fetch_ranges(
url: &str,
ranges: Vec<f64>,
opts: &JsValue,
validator: &JsValue,
length: f64,
) -> Result<JsValue, JsValue>;
}
/// Where a remote file's bytes come from.
struct Http {
url: String,
opts: JsValue,
/// ETag or Last-Modified at open (`null` if the server sent neither).
validator: JsValue,
length: u64,
/// Requests the probe made (1, or 2 with a HEAD for the length).
probe_requests: u64,
}
enum Source {
/// Read by range requests through a restartable cache.
Lazy {
http: Http,
storage: Arc<LazyStorage>,
reader: Reader,
},
/// The server ignored `Range`: the whole file, downloaded at open.
Whole {
reader: Reader,
size: u64,
requests: u64,
},
}
impl Http {
/// Fetch `ranges` and hand them to `storage`.
async fn fetch(&self, storage: &LazyStorage, ranges: &[Range<u64>]) -> Result<(), JsError> {
let flat: Vec<f64> = ranges
.iter()
.flat_map(|r| [r.start as f64, r.end as f64])
.collect();
let got = fetch_ranges(
&self.url,
flat,
&self.opts,
&self.validator,
self.length as f64,
)
.await
.map_err(js_exception)?;
let got = Array::from(&got);
if got.length() as usize != ranges.len() {
return Err(js_err(format!(
"fetchRanges returned {} ranges for {}",
got.length(),
ranges.len()
)));
}
for (r, bytes) in ranges.iter().zip(got.iter()) {
let bytes = Uint8Array::new(&bytes).to_vec();
storage.supply_range(r, &bytes).map_err(js_err)?;
}
Ok(())
}
/// Run `f` over `storage` until it has every byte it reads.
async fn drive<T>(
&self,
storage: &LazyStorage,
mut f: impl FnMut() -> T,
) -> Result<T, JsError> {
let op = storage.operation();
loop {
match storage.attempt(&mut f) {
Step::Done(v) => return Ok(v),
Step::Need(ranges) => {
op.charge(&ranges).map_err(js_err)?;
self.fetch(storage, &ranges).await?
}
}
}
}
}
impl Source {
async fn run<T>(&self, op: impl Fn(&Reader) -> core::Result<T>) -> Result<T, JsError> {
match self {
Source::Whole { reader, .. } => op(reader).map_err(js_err),
Source::Lazy {
http,
storage,
reader,
} => http.drive(storage, || op(reader)).await?.map_err(js_err),
}
}
}
/// A non-negative integer option, or `None` when not given.
fn int_opt(opts: &JsValue, key: &str, min: f64, max: f64) -> Result<Option<u64>, JsError> {
let v = get(opts, key);
if v.is_undefined() || v.is_null() {
return Ok(None);
}
match v.as_f64() {
Some(x) if x.fract() == 0.0 && (min..=max).contains(&x) => Ok(Some(x as u64)),
_ => Err(js_err(format!(
"openUrl: {key} must be an integer from {min} to {max}"
))),
}
}
/// Most bytes `maxFetch` and `maxDownload` may allow: 1 GiB. wasm32 has
/// 4 GiB of memory and no buffer past 2 GiB, and what is fetched is held
/// while it is decoded.
const MAX_FETCH_LIMIT: u64 = 1 << 30;
/// The largest file `openUrl` reads by ranges: on wasm32, 4 GiB - 1 bytes.
/// The format code turns file offsets into `usize` to use them (with a
/// clean error past it, see scripts/check-32bit-casts.sh), so on a 32-bit
/// target nothing at 4 GiB or beyond can be read; a larger file is refused
/// at open rather than failing on whichever read reaches past 4 GiB. On
/// 64-bit targets it is 2^53 - 1, the largest offset a JavaScript number
/// holds exactly.
const MAX_REMOTE_LENGTH: u64 = if (usize::MAX as u64) < MAX_SAFE_INTEGER as u64 {
usize::MAX as u64
} else {
MAX_SAFE_INTEGER as u64
};
fn config_from(opts: &JsValue) -> Result<LazyConfig, JsError> {
let mut c = LazyConfig::default();
if let Some(b) = int_opt(opts, "blockSize", 512.0, (64u64 << 20) as f64)? {
c.block_size = b;
}
if let Some(n) = int_opt(opts, "cacheSize", 0.0, MAX_SAFE_INTEGER)? {
c.capacity = n;
}
if let Some(n) = int_opt(opts, "maxFetch", 512.0, MAX_FETCH_LIMIT as f64)? {
c.max_fetch = n;
}
// Read by remote.js; checked here so a value wasm32 cannot hold is an
// option error rather than a download that cannot be kept.
int_opt(opts, "maxDownload", 0.0, MAX_FETCH_LIMIT as f64)?;
// Also read by remote.js (which checks it too, for direct callers).
int_opt(opts, "parallel", 1.0, 1024.0)?;
Ok(c)
}
/// Open the HDF5 file at `url` without downloading it: its bytes are
/// fetched with HTTP `Range` requests as the methods of the returned
/// [`RemoteFile`] need them, through a block cache.
///
/// `opts` (all optional):
/// - `blockSize` — bytes per request block, 512 to 64 MiB (default 1 MiB);
/// - `cacheSize` — bytes of blocks kept between calls (default 64 MiB);
/// - `maxFetch` — most bytes one call may fetch, and so the longest single
/// read, up to 1 GiB (default 512 MiB): a call that would fetch more
/// fails before fetching it;
/// - `fallback` — `"download"` (default) reads the whole file when the
/// server ignores `Range` (answers 200), up to `maxDownload` bytes
/// (default 512 MiB, at most 1 GiB); `"error"` refuses such a server;
/// - `headers`, `credentials` — passed to every `fetch` (`headers` as
/// `fetch` takes them: a `Headers`, `[name, value]` pairs or an object);
/// - `parallel` — range requests in flight at once, 1 to 1024 (default
/// 6); when one fails the others are aborted;
/// - `fetch` — a `fetch`-compatible function to use instead of the global.
///
/// Cross-origin servers must allow CORS and expose `Content-Range` (or
/// answer `HEAD` with `Content-Length`). A file may be up to 4 GiB - 1
/// bytes long (wasm32 offsets); a longer one is refused at open. A whole-dataset `read` that would use more than 1 GiB of
/// memory ([`core::MAX_READ_BYTES`]) is refused: read it in parts with
/// `readHyperslab`.
#[wasm_bindgen(js_name = openUrl)]
pub async fn open_url(url: String, opts: JsValue) -> Result<RemoteFile, JsError> {
let config = config_from(&opts)?;
let p = probe(&url, config.block_size as f64, &opts)
.await
.map_err(js_exception)?;
let requests = get(&p, "requests").as_f64().unwrap_or(1.0) as u64;
let whole = get(&p, "whole");
if !whole.is_undefined() {
let bytes = Uint8Array::new(&whole).to_vec();
let size = bytes.len() as u64;
let reader = Reader::open(bytes).map_err(js_err)?;
return Ok(RemoteFile {
inner: Rc::new(Source::Whole {
reader,
size,
requests,
}),
});
}
let length = get(&p, "length")
.as_f64()
.filter(|x| x.fract() == 0.0 && (0.0..=MAX_SAFE_INTEGER).contains(x))
.ok_or_else(|| js_err(format!("{url}: the server gave no usable file size")))?;
if length as u64 > MAX_REMOTE_LENGTH {
return Err(js_err(format!(
"{url} is {length} bytes; openUrl reads files of up to {MAX_REMOTE_LENGTH} bytes \
(4 GiB - 1: the WebAssembly reader addresses a file with 32-bit offsets)"
)));
}
let http = Http {
url,
opts,
validator: get(&p, "validator"),
length: length as u64,
probe_requests: requests,
};
self.read_impl(path, Some(&slab))
let storage = Arc::new(LazyStorage::new(http.length, config));
let first = Uint8Array::new(&get(&p, "first")).to_vec();
storage.supply(0, &first).map_err(js_err)?;
let s = storage.clone();
let reader = http
.drive(&storage, || Reader::open_storage(s.clone()))
.await?
.map_err(js_err)?;
Ok(RemoteFile {
inner: Rc::new(Source::Lazy {
http,
storage,
reader,
}),
})
}
/// A file opened with [`openUrl`](open_url): the methods of [`H5File`],
/// each returning a `Promise` (it may have to fetch bytes first).
#[wasm_bindgen]
pub struct RemoteFile {
inner: Rc<Source>,
}
impl RemoteFile {
/// Run `op` (fetching what it needs) and convert its result.
fn call<T: 'static>(
&self,
op: impl Fn(&Reader) -> core::Result<T> + 'static,
to_js: impl FnOnce(T) -> JsValue + 'static,
) -> Promise {
let inner = self.inner.clone();
future_to_promise(async move {
let v = inner.run(op).await.map_err(JsValue::from)?;
Ok(to_js(v))
})
}
}
#[wasm_bindgen]
impl RemoteFile {
/// `"group"` or `"dataset"`.
#[wasm_bindgen(unchecked_return_type = "Promise<string>")]
pub fn kind(&self, path: String) -> Promise {
self.call(move |r| r.kind(&path), |k| k.as_str().into())
}
/// The group's members: `[{ name, kind }]`, groups first.
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
pub fn list(&self, path: String) -> Promise {
self.call(move |r| r.list(&path), |c| list_to_js(c).into())
}
/// `{ shape, maxshape, dtype, elementShape }`, as [`H5File::info`].
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
pub fn info(&self, path: String) -> Promise {
self.call(move |r| r.info(&path), |i| info_to_js(i).into())
}
/// `[{ name, value, dtype }]`, as [`H5File::attrs`].
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
pub fn attrs(&self, path: String) -> Promise {
self.call(move |r| r.attrs(&path), |a| attrs_to_js(a.0).into())
}
/// Messages for attributes that could not be read.
#[wasm_bindgen(js_name = attrErrors, unchecked_return_type = "Promise<Array<string>>")]
pub fn attr_errors(&self, path: String) -> Promise {
self.call(move |r| r.attrs(&path), |a| errors_to_js(a.1).into())
}
/// The whole dataset: `{ shape, dtype, data }`, as [`H5File::read`].
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
pub fn read(&self, path: String) -> Promise {
self.call(
move |r| read_values(r, &path, None),
|v| values_to_js(v).into(),
)
}
/// A regular hyperslab, as [`H5File::read_hyperslab`]. Only the chunks
/// (or the contiguous runs) the selection touches are fetched.
#[wasm_bindgen(js_name = readHyperslab, unchecked_return_type = "Promise<any>")]
pub fn read_hyperslab(
&self,
path: String,
start: Vec<f64>,
count: Vec<f64>,
stride: Option<Vec<f64>>,
block: Option<Vec<f64>>,
) -> Result<Promise, JsError> {
let slab = slab_from_js(&start, &count, stride, block)?;
Ok(self.call(
move |r| read_values(r, &path, Some(&slab)),
|v| values_to_js(v).into(),
))
}
fn read_impl(&self, path: &str, slab: Option<&Hyperslab>) -> Result<Object, JsError> {
let dtype = self.inner.info(path).map_err(js_err)?.dtype;
let a = self.inner.read(path, slab).map_err(js_err)?;
/// What reading this file has cost so far: `{ lazy, size, requests,
/// bytesFetched, cachedBytes, passes }`. `lazy` is false when the
/// server ignored `Range` and the file was downloaded whole.
pub fn stats(&self) -> Object {
let o = Object::new();
set(&o, "shape", shape_to_js(&a.shape));
set(&o, "dtype", dtype);
set(&o, "data", data_to_js(a.data));
Ok(o)
match &*self.inner {
Source::Lazy { http, storage, .. } => {
let st = storage.stats();
set(&o, "lazy", true);
set(&o, "size", http.length as f64);
set(
&o,
"requests",
(st.requests.saturating_sub(1) + http.probe_requests) as f64,
);
set(&o, "bytesFetched", st.bytes_fetched as f64);
set(&o, "cachedBytes", st.cached_bytes as f64);
set(&o, "passes", st.passes as f64);
}
Source::Whole { size, requests, .. } => {
set(&o, "lazy", false);
set(&o, "size", *size as f64);
set(&o, "requests", *requests as f64);
set(&o, "bytesFetched", *size as f64);
set(&o, "cachedBytes", *size as f64);
set(&o, "passes", 0.0);
}
}
o
}
}
+578
View File
@@ -0,0 +1,578 @@
//! The restartable ("NeedBytes") reader against the in-memory one: every
//! file must list, describe and read the same through a [`LazyStorage`]
//! that starts empty and is fed only the ranges its passes ask for, as the
//! browser's `openUrl` feeds it from HTTP range requests.
//!
//! - Files written here with `FileBuilder`, at several block sizes (512 B
//! blocks make almost every structure read a miss).
//! - The h5py/netCDF4 fixture of `examples/wasm-viewer/test/make_fixture.py`
//! (skipped without h5py, unless `CLAWHDF5_REQUIRE_INTEROP=1`;
//! `CLAWHDF5_PYTHON` names the interpreter).
//! - `CLAWHDF5_WASM_CORPUS=dir[:dir...]`: every HDF5 file under those
//! directories up to 64 MiB (e.g. `conformance/.cache/corpus`).
//!
//! Also the request budget: listing and reading one small dataset of a large
//! file fetches a few blocks, not the file.
use std::ops::Range;
use std::path::{Path, PathBuf};
use std::process::Command;
use std::sync::Arc;
use clawhdf5::{AttrValue, FileBuilder};
use clawhdf5_format::storage::CountingStorage;
use clawhdf5_wasm::core::{Hyperslab, Kind, Reader};
use clawhdf5_wasm::lazy::{LazyConfig, LazyStorage};
/// The API the JavaScript side calls, one operation at a time.
trait Api {
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T;
}
struct Local(Reader);
impl Api for Local {
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
op(&self.0)
}
}
/// A lazily read file and the "server" it fetches from.
struct Lazy {
data: Arc<Vec<u8>>,
storage: Arc<LazyStorage>,
reader: Reader,
}
fn fetch(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
Ok(data[r.start as usize..r.end as usize].to_vec())
}
impl Lazy {
/// Open as `openUrl` does: the first block comes with the probe that
/// learns the length, then the open is run until it has its bytes.
fn open(data: Vec<u8>, config: LazyConfig) -> Result<Lazy, String> {
let data = Arc::new(data);
let storage = Arc::new(LazyStorage::new(data.len() as u64, config));
let first = (storage.config().block_size as usize).min(data.len());
storage.supply(0, &data[..first])?;
let s = storage.clone();
let reader =
storage.run_blocking(|| Reader::open_storage(s.clone()), |r| fetch(&data, r))??;
Ok(Lazy {
data,
storage,
reader,
})
}
}
impl Api for Lazy {
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
self.storage
.run_blocking(|| op(&self.reader), |r| fetch(&self.data, r))
.expect("serving from memory cannot fail")
}
}
/// Everything the viewer can show of a file, as text: each object's kind,
/// listing, attributes (and attribute errors), dataset info, whole value
/// and a hyperslab — or the error each gives.
fn transcript(api: &impl Api) -> Vec<String> {
let mut out = Vec::new();
let mut todo = vec![("/".to_string(), 0usize)];
while let Some((path, depth)) = todo.pop() {
if out.len() > 4000 {
out.push("... (truncated)".into());
break;
}
let kind = api.call(|r| r.kind(&path));
out.push(format!("{path}: {kind:?}"));
out.push(format!("{path} attrs: {:?}", api.call(|r| r.attrs(&path))));
match kind {
Ok(Kind::Group) => {
let list = api.call(|r| r.list(&path));
out.push(format!("{path} list: {list:?}"));
if let Ok(children) = list
&& depth < 12
{
for c in children.into_iter().rev() {
let child = if path == "/" {
format!("/{}", c.name)
} else {
format!("{path}/{}", c.name)
};
todo.push((child, depth + 1));
}
}
}
Ok(Kind::Dataset) => {
let info = api.call(|r| r.info(&path));
out.push(format!("{path} info: {info:?}"));
let Ok(info) = info else { continue };
let n = info
.shape
.iter()
.chain(&info.element_shape)
.try_fold(1u64, |a, &d| a.checked_mul(d));
if n.is_none_or(|n| n > 4_000_000) {
out.push(format!("{path}: not read ({n:?} values)"));
continue;
}
out.push(format!(
"{path} read: {:?}",
api.call(|r| r.read(&path, None))
));
if !info.shape.is_empty() && info.shape.iter().all(|&d| d > 1) {
let slab = Hyperslab {
start: info.shape.iter().map(|_| 1).collect(),
count: info.shape.iter().map(|&d| d / 2).collect(),
stride: None,
block: None,
};
let part = api.call(|r| r.read(&path, Some(&slab)));
out.push(format!("{path} slab: {part:?}"));
}
}
Err(_) => {}
}
}
out
}
/// The lazy transcript of `data` at `block` bytes per block equals the
/// transcript of the same file through a range storage that has every byte
/// (`CountingStorage`: the facade's `Storage` path, the one the lazy reader
/// takes), and agrees with the in-memory one: the same values, and an error
/// wherever it has one (a malformed file can fail at a different check,
/// with a different message, when read by ranges). Returns what the lazy
/// reader fetched and its transcript.
fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec<String>) {
let ctx = format!("{name} (blocks of {block} B)");
let ranged = Reader::open_storage(Arc::new(CountingStorage::new(data.to_vec())));
let local = Reader::open(data.to_vec());
let lazy = Lazy::open(data.to_vec(), config(block));
let (ranged, local, lazy) = match (ranged, local, lazy) {
(Ok(r), Ok(l), Ok(z)) => (r, l, z),
(Err(r), Err(_), Err(z)) => {
assert_eq!(z, r, "{ctx}: open error");
return (0, 0, Vec::new());
}
(r, l, z) => panic!(
"{ctx}: opens differently: ranged {:?}, in memory {:?}, lazily {:?}",
r.err(),
l.err(),
z.err()
),
};
let got = transcript(&lazy);
let want = transcript(&Local(ranged));
for (i, (w, g)) in want.iter().zip(&got).enumerate() {
assert_eq!(g, w, "{ctx}, line {i}");
}
assert_eq!(got.len(), want.len(), "{ctx}: transcript length");
let local = transcript(&Local(local));
for (i, (l, g)) in local.iter().zip(&got).enumerate() {
let both_errors = match (l.split_once("Err("), g.split_once("Err(")) {
(Some((a, _)), Some((b, _))) => a == b,
_ => false,
};
assert!(
l == g || both_errors,
"{ctx}, line {i}: in memory\n {l}\nlazily\n {g}"
);
}
assert_eq!(got.len(), local.len(), "{ctx}: transcript length");
let st = lazy.storage.stats();
(st.requests, st.bytes_fetched, got)
}
fn config(block: u64) -> LazyConfig {
LazyConfig {
block_size: block,
// A small budget, so eviction between operations is exercised.
capacity: 16 * block,
max_request: 8 * block,
..LazyConfig::default()
}
}
fn builder_file() -> Vec<u8> {
let mut b = FileBuilder::new();
b.create_dataset("grid")
.with_f64_data(&(0..20_000).map(f64::from).collect::<Vec<_>>())
.with_shape(&[100, 200])
.with_chunks(&[10, 25])
.with_deflate(4);
b.create_dataset("contiguous")
.with_i32_data(&(0..50_000).collect::<Vec<_>>());
b.create_dataset("bytes").with_u8_data(&[1, 2, 250]);
let mut g = b.create_group("sensors");
for i in 0..40 {
g.create_dataset(&format!("t{i}"))
.with_f32_data(&[i as f32, 1.5, -2.25]);
}
g.set_attr("location", AttrValue::String("lab".into()));
b.add_group(g.finish());
b.set_attr("version", AttrValue::I64(3));
b.set_attr("scale", AttrValue::F64Array(vec![0.5, 2.0]));
b.finish().unwrap()
}
#[test]
fn builder_files_read_the_same_at_every_block_size() {
let data = builder_file();
for block in [512, 4096, 1 << 20] {
let (requests, _, lines) = check_equal("builder", &data, block);
assert!(requests > 0);
// The transcript covers every object, values included.
assert!(lines.iter().any(|l| l.starts_with("/grid read: Ok")));
assert!(lines.iter().any(|l| l.starts_with("/grid slab: Ok")));
assert!(lines.iter().any(|l| l.starts_with("/sensors/t39 read: Ok")));
}
}
#[test]
fn garbage_fails_to_open_as_in_memory() {
check_equal("zeros", &[0u8; 5000], 512);
check_equal("empty", &[], 512);
let mut cut = builder_file();
cut.truncate(cut.len() / 3);
check_equal("truncated", &cut, 512);
}
/// Listing a large file and reading one small dataset fetches a few blocks,
/// not the file.
#[test]
fn a_small_read_of_a_large_file_fetches_a_few_blocks() {
let mut b = FileBuilder::new();
b.create_dataset("small").with_f64_data(&[1.0, 2.0, 3.0]);
// 48 MB of raw data, written after the small dataset's metadata.
b.create_dataset("big")
.with_f64_data(&(0..6_000_000).map(f64::from).collect::<Vec<_>>());
let mut g = b.create_group("group");
g.create_dataset("inner").with_i32_data(&[7, 8]);
b.add_group(g.finish());
let data = b.finish().unwrap();
let lazy = Lazy::open(data.clone(), LazyConfig::default()).unwrap();
let list = lazy.call(|r| r.list("/")).unwrap();
assert_eq!(list.len(), 3);
assert_eq!(
format!("{:?}", lazy.call(|r| r.read("/small", None)).unwrap().data),
"F64([1.0, 2.0, 3.0])"
);
assert_eq!(
format!(
"{:?}",
lazy.call(|r| r.read("/group/inner", None)).unwrap().data
),
"I32([7, 8])"
);
// A window of the big dataset reads only its block(s).
let slab = Hyperslab {
start: vec![3_000_000],
count: vec![4],
stride: None,
block: None,
};
assert_eq!(
format!(
"{:?}",
lazy.call(|r| r.read("/big", Some(&slab))).unwrap().data
),
"F64([3000000.0, 3000001.0, 3000002.0, 3000003.0])"
);
let st = lazy.storage.stats();
eprintln!("{} bytes: {st:?}", data.len());
assert!(st.requests <= 6, "{st:?}");
assert!(st.bytes_fetched <= 6 << 20, "{st:?}");
assert!(st.bytes_fetched * 8 < data.len() as u64, "{st:?}");
}
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn python_available() -> bool {
Command::new(python())
.args(["-c", "import h5py, netCDF4, numpy"])
.output()
.is_ok_and(|o| o.status.success())
}
#[test]
fn h5py_and_netcdf4_files_read_the_same_lazily() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
python()
);
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
return;
}
let dir = fixture_dir();
for name in ["fixture.h5", "fixture.nc"] {
let data = std::fs::read(dir.path().join(name)).unwrap();
for block in [512, 64 * 1024] {
let (_, _, lines) = check_equal(name, &data, block);
assert!(lines.iter().filter(|l| l.contains(" read: Ok")).count() >= 2);
}
}
}
/// The bytes of `data` at `r`, zero past its end: a server that claims
/// the file is longer than it is.
fn fetch_padded(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
let mut out = vec![0u8; (r.end - r.start) as usize];
let len = data.len() as u64;
if r.start < len {
let end = r.end.min(len);
out[..(end - r.start) as usize].copy_from_slice(&data[r.start as usize..end as usize]);
}
Ok(out)
}
/// Sizes a hostile server or a large dataset can name are errors, never
/// allocations that abort the wasm module: make_fixture.py's limits.h5 and
/// hostile_vl.h5 (see write_limits there).
#[test]
fn size_limits_are_errors_not_aborts() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
python()
);
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
return;
}
let dir = fixture_dir();
// Read whole, /huge_u8 would widen 2^28 values to 64 bits (2 GiB): an
// error naming readHyperslab, before its chunks are read. A window of
// it reads.
let data = std::fs::read(dir.path().join("limits.h5")).unwrap();
let n = (1u64 << 28) + 1024;
let window = Hyperslab {
start: vec![n - 4],
count: vec![4],
stride: None,
block: None,
};
let local = Reader::open(data.clone()).unwrap();
let lazy = Lazy::open(data, LazyConfig::default()).unwrap();
let before = lazy.storage.stats().requests;
for e in [
local.read("/huge_u8", None).unwrap_err(),
lazy.call(|r| r.read("/huge_u8", None)).unwrap_err(),
] {
assert!(e.contains("readHyperslab"), "{e}");
}
assert_eq!(lazy.storage.stats().requests, before, "nothing fetched");
for part in [
local.read("/huge_u8", Some(&window)).unwrap(),
lazy.call(|r| r.read("/huge_u8", Some(&window))).unwrap(),
] {
assert_eq!(format!("{:?}", part.data), "U8([0, 0, 0, 7])");
}
// A server that claims 3 GiB and a heap collection of 2 GiB + 4 KiB:
// reading the strings fails at once, fetching a few blocks.
let data = std::fs::read(dir.path().join("hostile_vl.h5")).unwrap();
let storage = Arc::new(LazyStorage::new(3 << 30, LazyConfig::default()));
let s = storage.clone();
let reader = storage
.run_blocking(
|| Reader::open_storage(s.clone()),
|r| fetch_padded(&data, r),
)
.unwrap()
.unwrap();
let e = storage
.run_blocking(|| reader.read("/a", None), |r| fetch_padded(&data, r))
.unwrap()
.unwrap_err();
assert!(e.contains("maxFetch"), "{e}");
let st = storage.stats();
assert!(st.requests <= 4 && st.bytes_fetched <= 4 << 20, "{st:?}");
}
/// make_fixture.py's files, written to a temporary directory.
fn fixture_dir() -> tempfile::TempDir {
let dir = tempfile::tempdir().unwrap();
let generator = Path::new(env!("CARGO_MANIFEST_DIR"))
.join("../../examples/wasm-viewer/test/make_fixture.py");
let out = Command::new(python())
.arg(&generator)
.arg(dir.path())
.output()
.unwrap();
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
dir
}
fn hdf5_files(dir: &Path, out: &mut Vec<PathBuf>) {
let Ok(entries) = std::fs::read_dir(dir) else {
return;
};
for e in entries.flatten() {
let p = e.path();
if p.is_dir() {
hdf5_files(&p, out);
} else if std::fs::read(&p)
.ok()
.is_some_and(|b| b.len() <= 64 << 20 && is_hdf5(&b))
{
out.push(p);
}
}
}
/// The HDF5 signature at 0 or a power-of-two user-block offset.
fn is_hdf5(b: &[u8]) -> bool {
const SIG: &[u8] = b"\x89HDF\r\n\x1a\n";
let mut at = 0usize;
loop {
if b.get(at..at + 8) == Some(SIG) {
return true;
}
at = if at == 0 { 512 } else { at * 2 };
if at >= b.len() {
return false;
}
}
}
#[test]
fn corpus_files_read_the_same_lazily() {
let Ok(dirs) = std::env::var("CLAWHDF5_WASM_CORPUS") else {
eprintln!("CLAWHDF5_WASM_CORPUS not set; skipping the corpus");
return;
};
let mut files = Vec::new();
for d in std::env::split_paths(&dirs) {
hdf5_files(&d, &mut files);
}
files.sort();
assert!(!files.is_empty(), "no HDF5 files under {dirs}");
let (mut requests, mut bytes, mut total) = (0u64, 0u64, 0u64);
for f in &files {
let data = std::fs::read(f).unwrap();
total += data.len() as u64;
let (r, b, _) = check_equal(&f.display().to_string(), &data, 64 * 1024);
requests += r;
bytes += b;
}
eprintln!(
"{} files ({total} bytes): {requests} requests, {bytes} bytes fetched",
files.len()
);
}
/// Passes and requests `list(path)` takes on a file opened lazily at
/// `block`-byte blocks (the open not counted), checking the listing against
/// the in-memory one.
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
let lazy = Lazy::open(
data.to_vec(),
LazyConfig {
block_size: block,
..LazyConfig::default()
},
)
.unwrap();
let before = lazy.storage.stats();
assert_eq!(lazy.call(|r| r.list(path)).unwrap(), want);
let after = lazy.storage.stats();
(
after.passes - before.passes,
after.requests - before.requests,
)
}
/// Listing a group reads every child's object header, and its index (B-tree
/// and symbol table nodes, or B-tree v2 and heap blocks) before that. Each
/// pass asks for every node of a level it is missing, not the first one
/// only, so the passes (network round trips) grow with the depth of the
/// index, not with the number of children: 2000 children with headers
/// scattered over 512-byte blocks list in a handful of passes, where each
/// header block used to cost its own.
#[test]
fn listing_a_large_group_takes_a_few_passes() {
let mut b = FileBuilder::new();
let mut g = b.create_group("many");
for i in 0..600 {
g.create_dataset(&format!("d{i}")).with_i32_data(&[i; 64]);
}
b.add_group(g.finish());
let data = b.finish().unwrap();
let (passes, requests) = listing_cost(&data, "/many", 512);
eprintln!("FileBuilder, 600 children: {passes} passes, {requests} requests");
assert!(passes <= 6, "{passes} passes");
if !python_available() {
return;
}
let dir = tempfile::tempdir().unwrap();
for libver in ["earliest", "latest"] {
let path = dir.path().join(format!("{libver}.h5"));
let script = format!(
"import h5py, numpy as np\n\
with h5py.File({:?}, 'w', libver='{libver}') as f:\n\
\x20 for i in range(2000):\n\
\x20 f.create_dataset('d%d' % i, data=np.full(256, i, np.float32))\n",
path.display().to_string()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.unwrap();
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let data = std::fs::read(&path).unwrap();
let (passes, requests) = listing_cost(&data, "/", 512);
eprintln!("h5py libver={libver}, 2000 children: {passes} passes, {requests} requests");
assert!(passes <= 12, "{libver}: {passes} passes");
}
}
/// `CLAWHDF5_WASM_LIST_FILE=file.h5`: what listing the root group of that
/// file costs lazily, at 1 MiB and 64 KiB blocks (a measurement, printed).
#[test]
fn listing_cost_of_a_given_file() {
let Ok(path) = std::env::var("CLAWHDF5_WASM_LIST_FILE") else {
return;
};
let data = std::fs::read(&path).unwrap();
for block in [1 << 20, 64 << 10] {
let lazy = Lazy::open(
data.clone(),
LazyConfig {
block_size: block,
..LazyConfig::default()
},
)
.unwrap();
let open = lazy.storage.stats();
let n = lazy.call(|r| r.list("/")).unwrap().len();
let st = lazy.storage.stats();
eprintln!(
"{path} ({} bytes), {block}-byte blocks: open {} requests / {} passes; list('/') of {n}: {} passes, {} requests, {} bytes",
data.len(),
open.requests,
open.passes,
st.passes - open.passes,
st.requests - open.requests,
st.bytes_fetched - open.bytes_fetched
);
}
}
+116 -8
View File
@@ -65,7 +65,8 @@ const MSG_FLAG_DONTSHARE: u8 = 0x04;
/// that may be mid-update.
///
/// Every method is one self-contained edit: it re-reads the file's
/// metadata, applies the change, and syncs the file before returning.
/// metadata (from the file it holds open, never by path), applies the
/// change, and syncs the file before returning.
///
/// # What it can change
///
@@ -750,9 +751,17 @@ impl FileEditor {
/// consistent: a metadata cache image, paged or persistent free-space
/// management, a multi-file driver, a file another writer has marked
/// open (superblock version 3 consistency flags).
///
/// The path is only used to open the file: every edit is planned from
/// and written to the file opened here, even if the path is renamed,
/// replaced or (relative) resolved from another working directory
/// later. [`path`](Self::path) is the absolute path it had at open.
pub fn open<P: AsRef<Path>>(path: P) -> Result<Self, Error> {
let path = path.as_ref().to_path_buf();
let file = OpenOptions::new().read(true).write(true).open(&path)?;
let file = OpenOptions::new()
.read(true)
.write(true)
.open(path.as_ref())?;
let path = std::fs::canonicalize(path.as_ref())?;
match file.try_lock() {
Ok(()) => {}
Err(TryLockError::WouldBlock) => {
@@ -768,16 +777,81 @@ impl FileEditor {
file,
free: FreeList::default(),
};
let f = File::open(&ed.path)?;
let f = ed.plan_reader()?;
check_editable(&f)?;
Ok(ed)
}
/// The file's path.
/// The file's absolute path when it was opened (it may have been
/// renamed since; the editor keeps editing the file it opened).
pub fn path(&self) -> &Path {
&self.path
}
/// A reader over the file this editor holds, as last written: the file
/// opened by [`open`](Self::open), not whatever its path names now.
///
/// It opens the file anew (read-only), so it does not share the
/// editor's lock and stays usable after the editor is dropped: on Linux
/// through `/proc/self/fd`, which reaches the held file even after its
/// path was renamed or replaced; elsewhere by the path the file had at
/// open, refused with [`Error::Io`] when that path no longer names the
/// held file (on Unix, compared by device and inode; Windows cannot
/// check). With the `mmap` feature the reader maps the file: edits
/// through the editor change the bytes it sees, so take a new reader
/// after each edit rather than reading through an old one while an edit
/// runs.
pub fn reader(&self) -> Result<File, Error> {
let dir = self.path.parent().map(Path::to_path_buf);
File::from_std_file(self.reopen()?, dir)
}
/// A new read-only open file description of the held file (see
/// [`reader`](Self::reader)).
fn reopen(&self) -> Result<std::fs::File, Error> {
#[cfg(target_os = "linux")]
{
use std::os::fd::AsRawFd;
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
if let Ok(f) = std::fs::File::open(proc) {
return Ok(f);
}
}
let f = std::fs::File::open(&self.path)?;
#[cfg(unix)]
{
use std::os::unix::fs::MetadataExt;
let (a, b) = (self.file.metadata()?, f.metadata()?);
if (a.dev(), a.ino()) != (b.dev(), b.ino()) {
return Err(Error::Io(std::io::Error::other(format!(
"{} no longer names the file being edited (renamed or replaced)",
self.path.display()
))));
}
}
Ok(f)
}
/// A reader over the held file for planning an edit (dropped before the
/// edit writes). On Linux a new open file description through
/// `/proc/self/fd`: a mapping of a clone of the held descriptor would
/// share its `flock`, and a process forked meanwhile (any
/// `std::process::Command` on another thread) would briefly keep the
/// lock alive after the editor is dropped. Elsewhere a clone of the
/// held descriptor, which follows the file wherever its path goes.
fn plan_reader(&self) -> Result<File, Error> {
let dir = self.path.parent().map(Path::to_path_buf);
#[cfg(target_os = "linux")]
{
use std::os::fd::AsRawFd;
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
if let Ok(f) = std::fs::File::open(proc) {
return File::from_std_file(f, dir);
}
}
File::from_std_file(self.file.try_clone()?, dir)
}
/// Bytes earlier edits of this editor freed that later ones can still
/// reuse.
pub fn reusable_bytes(&self) -> u64 {
@@ -794,7 +868,7 @@ impl FileEditor {
&mut self,
op: impl FnOnce(&File, &mut Image<'_>) -> Result<R, Error>,
) -> Result<R, Error> {
let f = File::open(&self.path)?;
let f = self.plan_reader()?;
check_editable(&f)?;
let sb = f.superblock().clone();
let user_block = f.user_block_size();
@@ -877,7 +951,7 @@ impl FileEditor {
/// the fill value (`H5D__chunk_prune_by_extent`).
pub fn resize(&mut self, path: &str, shape: &[u64]) -> Result<(), Error> {
self.edit(|f, img| {
let t = Target::load(f, path)?;
let mut t = Target::load(f, path)?;
let dims = t.dims().to_vec();
if shape.len() != dims.len() {
return Err(Error::InvalidArgument(format!(
@@ -889,7 +963,12 @@ impl FileEditor {
if shape == dims.as_slice() {
return Ok(());
}
let max = t.ds.max_dimensions.clone().unwrap_or_else(|| dims.clone());
// No maximum recorded means the current dimensions (see below).
let record_max = t.ds.max_dimensions.is_none();
let max =
t.ds.max_dimensions
.get_or_insert_with(|| dims.clone())
.clone();
for d in 0..dims.len() {
if shape[d] > max[d] {
return Err(Error::InvalidArgument(format!(
@@ -928,7 +1007,36 @@ impl FileEditor {
}
put_uint(&mut dims_bytes[d * ls..], n, img.ls);
}
if !record_max {
hdr.patch(img, i, first, &dims_bytes)?;
} else {
// No maximum recorded (clawhdf5's writer, for a dataset
// created without a maxshape). libhdf5 never writes such a
// dataspace: `H5S_set_extent_simple` records the maximum,
// equal to the dimensions when none is given. Reading one,
// libhdf5 takes the maximum to be the *current* dimensions
// (`H5S_extent_get_dims`), so changing them would also
// change the maximum the chunk index was built with — the
// Fixed Array linearises chunks by it — and move every
// existing chunk. Record the maximum libhdf5 would have
// written, the dimensions before this resize, so the index
// keeps its layout and the dataset can grow back to them.
let body_len = first + dims.len() * ls;
if body.len() < body_len || body[2] & !0x01 != 0 {
return Err(Error::Unsupported("dataspace message layout".into()));
}
let mut new_body = body[..first].to_vec();
new_body[2] |= 0x01;
new_body.extend_from_slice(&dims_bytes);
let at = new_body.len();
new_body.resize(at + dims.len() * ls, 0);
for (d, &n) in dims.iter().enumerate() {
put_uint(&mut new_body[at + d * ls..], n, img.ls);
}
let (flags, corder) = (hdr.msgs[i].flags, hdr.msgs[i].corder);
hdr.delete(img, i)?;
hdr.insert(img, MSG_DATASPACE, flags, &new_body, corder)?;
}
let fill = fill_info(img, &hdr)?;
hdr.finish(img)?;
let expand = shape.iter().zip(&dims).any(|(n, o)| n > o);
+2
View File
@@ -47,6 +47,7 @@ pub mod lazy;
#[cfg(feature = "mmap")]
pub mod mmap_file;
pub mod reader;
mod swmr;
pub mod types;
pub mod vlen;
pub mod writer;
@@ -57,6 +58,7 @@ pub use lazy::{LazyDataset, LazyFile, LazyGroup};
#[cfg(feature = "mmap")]
pub use mmap_file::{MmapDataset, MmapFile, MmapGroup};
pub use reader::{Dataset, File, Group, SharedStorage, VdsResolver};
pub use swmr::{FileStorage, SWMR_READ_ATTEMPTS};
pub use types::{AttrValue, DType};
pub use vlen::VlenValue;
pub use writer::FileBuilder;
+427 -37
View File
@@ -28,7 +28,9 @@ use clawhdf5_format::superblock_ext::{self, CacheImageState};
use crate::cache_image::{self, ImageView};
use crate::error::Error;
use crate::types::{AttrValue, DType, classify_datatype, read_attr, read_attrs};
use crate::types::{
AttrValue, DType, classify_datatype, read_attr_reporting, read_attrs_reporting,
};
// ---------------------------------------------------------------------------
// FileData — internal storage for owned bytes, an mmap, or any Storage
@@ -115,6 +117,10 @@ struct FileData {
/// parser reads asks for it, and the patched/overlay checks and range
/// conversions behind it cost a local metadata walk a few percent.
contiguous: Option<WholeView>,
/// A file a SWMR writer may still be appending to
/// ([`File::open_swmr`]): reads are bounded by the storage's current
/// length, not by `end`, and nothing is ever read as one slice.
live: bool,
}
/// A borrow of the HDF5 data held by a [`FileData`]'s own `backing` or
@@ -138,7 +144,7 @@ impl FileData {
/// bytes past the recorded end of file are not read, as in libhdf5.
fn new(mut backing: Backing) -> Result<(Self, Superblock), Error> {
if let Backing::Storage(storage) = backing {
return Self::new_storage(storage);
return Self::new_storage(storage, false, &mut None);
}
let whole = backing.whole_file().unwrap_or_default();
let (user_block, hdf5) = signature::split_user_block(whole)?;
@@ -173,6 +179,7 @@ impl FileData {
overlay: Vec::new(),
image_error,
contiguous: None,
live: false,
};
data.contiguous = data.find_contiguous();
Ok((data, superblock))
@@ -180,7 +187,18 @@ impl FileData {
/// [`Self::new`] for a [`Storage`] backend: the same checks, through
/// reads of the storage.
fn new_storage(storage: SharedStorage) -> Result<(Self, Superblock), Error> {
///
/// With `swmr` ([`File::open_storage_swmr`]) the file is read live when
/// its superblock has the SWMR-write flag (version 3), as libhdf5's SWMR
/// reader reads it: reads end where the file ends at the time of each
/// read, and nothing is read as one slice. Any other file is read as
/// without `swmr`, bounded by its recorded end of file. Once the
/// superblock has been read, `flagged` says which it was.
fn new_storage(
storage: SharedStorage,
swmr: bool,
flagged: &mut Option<bool>,
) -> Result<(Self, Superblock), Error> {
let file_len = storage.len();
let base = signature::find_signature_in(&*storage)?;
let mut data = Self {
@@ -191,11 +209,16 @@ impl FileData {
overlay: Vec::new(),
image_error: None,
contiguous: None,
// Until the superblock says otherwise: its own reads are not
// bounded by an end of file it has not read yet.
live: swmr,
};
// Worked out again below, once the end of file and any cache image
// are known.
data.contiguous = data.find_contiguous();
let superblock = Superblock::parse_in(&data, 0)?;
data.live = swmr && superblock.version >= 3 && superblock.is_swmr_write();
*flagged = Some(data.live);
data.end = base + superblock.data_end(base, file_len)?;
match superblock_ext::cache_image_state_in(&data, &superblock)? {
CacheImageState::Absent => {}
@@ -230,6 +253,9 @@ impl FileData {
/// [`Self::contiguous`], worked out from `backing` and `patched`.
fn find_contiguous(&self) -> Option<WholeView> {
if self.live {
return None;
}
let bytes = self.compute_contiguous()?;
Some(WholeView {
ptr: bytes.as_ptr(),
@@ -295,6 +321,12 @@ impl Storage for FileData {
if let Some(all) = self.contiguous() {
return all.read_at(offset, len);
}
if self.live {
// No end fixed at open: the storage's reads end where the file
// ends now.
let bytes = cut_to(self.remote()?.read_at(self.base + offset, len)?, len);
return Ok(self.with_overlay(offset, bytes));
}
let size = self.end - self.base;
let len = usize::try_from(size.saturating_sub(offset)).map_or(len, |avail| avail.min(len));
if len == 0 {
@@ -305,6 +337,11 @@ impl Storage for FileData {
}
fn len(&self) -> u64 {
if self.live
&& let Backing::Storage(s) = &self.backing
{
return s.len().saturating_sub(self.base);
}
self.end - self.base
}
@@ -312,7 +349,7 @@ impl Storage for FileData {
if let Some(all) = self.contiguous() {
return all.read_ranges(ranges);
}
let size = self.end - self.base;
let size = Storage::len(self);
let shifted: Vec<Range<u64>> = ranges
.iter()
.map(|r| {
@@ -375,6 +412,9 @@ pub struct File {
/// Resolves external Virtual Dataset source files instead of
/// `base_dir` (see [`File::set_vds_resolver`]).
vds_resolver: Option<VdsResolver>,
/// Attempts per operation on a live file, and the retries made (see
/// [`File::open_swmr`]).
swmr: crate::swmr::Retries,
}
impl File {
@@ -394,6 +434,7 @@ impl File {
chunk_cache: ChunkCache::new(),
base_dir,
vds_resolver: None,
swmr: crate::swmr::Retries::default(),
})
}
#[cfg(not(feature = "mmap"))]
@@ -405,6 +446,39 @@ impl File {
}
}
/// A reader over an already open file (the file itself, not whatever
/// its path names now): mapped with the `mmap` feature, else read into
/// memory. `base_dir` resolves external Virtual Dataset sources.
pub(crate) fn from_std_file(
file: std::fs::File,
base_dir: Option<std::path::PathBuf>,
) -> Result<Self, Error> {
#[cfg(feature = "mmap")]
let mut f = {
let reader = clawhdf5_io::MmapReader::from_file(file).map_err(Error::Io)?;
let (data, superblock) = FileData::new(Backing::Mmap(reader))?;
Self {
data,
superblock,
chunk_cache: ChunkCache::new(),
base_dir: None,
vds_resolver: None,
swmr: crate::swmr::Retries::default(),
}
};
#[cfg(not(feature = "mmap"))]
let mut f = {
use std::io::{Read, Seek, SeekFrom};
let mut file = file;
let mut bytes = Vec::new();
file.seek(SeekFrom::Start(0)).map_err(Error::Io)?;
file.read_to_end(&mut bytes).map_err(Error::Io)?;
Self::from_bytes(bytes)?
};
f.base_dir = base_dir;
Ok(f)
}
/// Open an HDF5 file by reading it entirely into memory.
///
/// This is the pre-mmap behaviour and is useful when memory-mapping is
@@ -428,6 +502,7 @@ impl File {
chunk_cache: ChunkCache::new(),
base_dir: None,
vds_resolver: None,
swmr: crate::swmr::Retries::default(),
})
}
@@ -463,9 +538,205 @@ impl File {
chunk_cache: ChunkCache::new(),
base_dir: None,
vds_resolver: None,
swmr: crate::swmr::Retries::default(),
})
}
/// Open a file that a libhdf5 SWMR writer (h5py `f.swmr_mode = True`)
/// may still be appending to, as libhdf5's SWMR reader does
/// (`H5F_ACC_SWMR_READ`, h5py `File(path, "r", swmr=True)`).
///
/// The file is read with positioned reads ([`FileStorage`](crate::FileStorage)),
/// never mapped, and reads are bounded by the file's length at the
/// time of each read rather than a length fixed at open. The chunk cache
/// is not used, so every read reads the chunk index and chunks as they
/// are now. A [`Dataset`] handle keeps the extent it was opened (or
/// last [refreshed](Dataset::refresh)) with, as in libhdf5: call
/// [`Dataset::refresh`] to see the writer's appends.
///
/// An operation that fails with an error a concurrent write can cause
/// — the failures libhdf5's SWMR reader retries: a checksum mismatch, a
/// read past the file's current end, an object header prefix (signature,
/// version) that does not decode — is run again from the start, up to
/// [`swmr_read_attempts`](Self::swmr_read_attempts) times (100 by
/// default, libhdf5's default for SWMR readers), with a pause of 1 µs
/// doubling up to 10 ms between attempts. Data is only returned from an
/// attempt in which every structure read verified, so a torn read is
/// an error (after the last attempt), never data. Every other error is
/// returned at once.
///
/// All of this applies only to a file whose superblock (version 3) has
/// the SWMR-write flag set when it is opened: a file a SWMR writer has
/// open, or had and did not close. libhdf5's SWMR writer does not keep
/// the superblock's recorded end of file current, so for such a file it
/// is ignored, as libhdf5's SWMR reader ignores it. Any other file (one
/// whose writer has closed it, or that was never written in SWMR mode)
/// is read exactly as [`File::open`] reads it — bounded by its recorded
/// end of file, through the chunk cache, each operation tried once —
/// and [`is_swmr_read`](Self::is_swmr_read) is `false`. Which of the two
/// a handle is does not change after it is opened.
///
/// A file whose superblock says it is open for writing without SWMR is
/// refused with [`Error::Locked`], as libhdf5 refuses it: such a writer
/// does not order its writes for readers.
pub fn open_swmr<P: AsRef<std::path::Path>>(path: P) -> Result<Self, Error> {
let storage = crate::swmr::FileStorage::open(path.as_ref()).map_err(Error::Io)?;
let mut f = Self::open_storage_swmr(Arc::new(storage))?;
f.base_dir = path.as_ref().parent().map(|p| p.to_path_buf());
Ok(f)
}
/// [`File::open_swmr`] over any [`Storage`] whose [`Storage::len`]
/// follows the file as it grows. A storage that caches blocks, or pins
/// the file's length at open (`clawhdf5-remote`'s `BlockCache` and
/// `HttpStorage`), does not show the writer's appends.
pub fn open_storage_swmr(storage: SharedStorage) -> Result<Self, Error> {
let swmr = crate::swmr::Retries::default();
// Retried only while the superblock cannot be read, or says a SWMR
// writer has the file: an error opening any other file is final.
let (data, superblock) = swmr.retry(|| {
let mut flagged = None;
match FileData::new_storage(storage.clone(), true, &mut flagged) {
Ok(opened) => Ok(Ok(opened)),
Err(e) if flagged == Some(false) => Ok(Err(e)),
Err(e) => Err(e),
}
})??;
if superblock.version >= 3 && superblock.is_write_access() && !superblock.is_swmr_write() {
return Err(Error::Locked(
"the file is open for writing without SWMR (libhdf5: \"file is already open \
for write\"); a SWMR reader needs a SWMR writer"
.into(),
));
}
Ok(Self {
data,
superblock,
chunk_cache: ChunkCache::new(),
base_dir: None,
vds_resolver: None,
swmr,
})
}
/// Whether the file is read live: opened with [`File::open_swmr`] or
/// [`File::open_storage_swmr`] while its superblock had the SWMR-write
/// flag set. `false` for every other file, including one opened with
/// `open_swmr` whose writer had already closed it (it reads as
/// [`File::open`] reads it).
pub fn is_swmr_read(&self) -> bool {
self.data.live
}
/// How many times an operation on a SWMR-read file is tried (see
/// [`File::open_swmr`]); 100 unless set. Files not read live
/// ([`is_swmr_read`](Self::is_swmr_read) `false`) try every operation
/// once.
pub fn swmr_read_attempts(&self) -> u32 {
self.swmr.attempts
}
/// Set how many times an operation on a SWMR-read file is tried before
/// its error is returned (libhdf5's `H5Pset_metadata_read_attempts`); at
/// least 1.
pub fn set_swmr_read_attempts(&mut self, attempts: u32) {
self.swmr.attempts = attempts.max(1);
}
/// How many times an operation on this SWMR-read file has been run
/// again because a concurrent write made it fail (the counterpart of
/// libhdf5's `H5Fget_metadata_read_retry_info`), counting the attempts
/// made while opening it. Always 0 for other files.
pub fn swmr_retries(&self) -> u64 {
self.swmr.retries()
}
/// Whether a SWMR writer has the file open now, as far as the file
/// says: the superblock is read again and its SWMR-write flag returned.
/// libhdf5 clears the flag when the writer closes the file, so a reader
/// can stop following it then (and one last [`Dataset::refresh`] sees
/// the final extents). A writer that crashed or was killed never clears
/// it, so the flag alone is not a stop condition: give the loop another
/// one, such as a time without growth.
///
/// ```no_run
/// # fn main() -> Result<(), clawhdf5::Error> {
/// use std::time::{Duration, Instant};
///
/// let file = clawhdf5::File::open_swmr("live.h5")?;
/// let mut ds = file.dataset("samples")?;
/// let (mut seen, mut last_growth) = (0, Instant::now());
/// while file.swmr_writer_active()? && last_growth.elapsed() < Duration::from_secs(60) {
/// ds.refresh()?;
/// let n = ds.shape()?[0];
/// if n > seen {
/// // read rows seen..n ...
/// (seen, last_growth) = (n, Instant::now());
/// }
/// std::thread::sleep(Duration::from_millis(100));
/// }
/// ds.refresh()?; // the final extent
/// # Ok(())
/// # }
/// ```
pub fn swmr_writer_active(&self) -> Result<bool, Error> {
self.retry(|| {
let sb = Superblock::parse_in(&self.data, 0)?;
Ok(sb.version >= 3 && sb.is_swmr_write())
})
}
/// Run `op`, again on errors a SWMR writer can cause when the file is
/// live (see [`File::open_swmr`]); once otherwise.
fn retry<T>(&self, op: impl FnMut() -> Result<T, Error>) -> Result<T, Error> {
if self.data.live {
self.swmr.retry(op)
} else {
let mut op = op;
op()
}
}
/// An attribute read under [`Self::retry`]. Attribute reads leave out
/// (or return as [`AttrValue::Raw`]) an attribute they cannot read
/// rather than fail, so `op` also returns the errors behind those: on a
/// live file one a concurrent write can cause runs the read again too.
/// If every attempt meets one, the last attempt's result is returned,
/// as it is for a file that is not live.
fn retry_attrs<T>(
&self,
mut op: impl FnMut() -> Result<(T, Vec<FormatError>), Error>,
) -> Result<T, Error> {
let mut last = None;
let result = self.retry(|| {
last = None;
let (value, errors) = op()?;
match errors.into_iter().find(crate::swmr::is_transient_format) {
Some(e) if self.data.live => {
last = Some(value);
Err(Error::Format(e))
}
_ => Ok(value),
}
});
match (result, last) {
(Err(_), Some(value)) => Ok(value),
(result, _) => result,
}
}
/// The chunk cache reads of this file go through: the file's own, or
/// for a live file a fresh one per read, since a cached chunk index
/// would hide the writer's new chunks and a cached edge chunk would
/// read as fill where the writer has written since.
fn with_chunk_cache<T>(&self, op: impl FnOnce(&ChunkCache) -> T) -> T {
if self.data.live {
op(&ChunkCache::new())
} else {
op(&self.chunk_cache)
}
}
/// Resolve external Virtual Dataset source files (their names as the
/// mappings store them) with `resolver`, instead of reading them from
/// the directory of the file (for [`File::open`]) or refusing them (for
@@ -488,6 +759,7 @@ impl File {
///
/// The path uses `/` separators (e.g., `"group1/values"`).
pub fn dataset(&self, path: &str) -> Result<Dataset<'_>, Error> {
self.retry(|| {
let addr = with_bytes!(self.data.meta()?, |d| group_v2::resolve_path_any_in(
d,
&self.superblock,
@@ -499,9 +771,11 @@ impl File {
}
Dataset {
file: self,
address: addr,
header: hdr,
}
.check_open()
})
}
/// A `Dataset` handle for the object header at `address` (an address
@@ -511,15 +785,18 @@ impl File {
/// are scanned); keep the address instead to open the same dataset
/// repeatedly.
pub fn dataset_at(&self, address: u64) -> Result<Dataset<'_>, Error> {
self.retry(|| {
let hdr = self.parse_header(address)?;
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(format!("object at address {address}")));
}
Dataset {
file: self,
address,
header: hdr,
}
.check_open()
})
}
/// A `Group` handle for the object header at `address` (from
@@ -538,11 +815,11 @@ impl File {
/// The path uses `/` separators (e.g., `"sensors"`).
/// Use `"/"` or `""` for the root group.
pub fn group(&self, path: &str) -> Result<Group<'_>, Error> {
let addr = with_bytes!(self.data.meta()?, |d| group_v2::resolve_path_any_in(
d,
&self.superblock,
path
))?;
let addr = self.retry(|| {
Ok(with_bytes!(self.data.meta()?, |d| {
group_v2::resolve_path_any_in(d, &self.superblock, path)
})?)
})?;
Ok(Group {
file: self,
address: addr,
@@ -665,6 +942,7 @@ impl File {
/// [`AttrValue::Raw`] attribute. Variable-length strings are resolved in
/// this file's global heap; see [`Dataset::read_string`] for the values.
pub fn decode_strings(&self, datatype: &Datatype, raw: &[u8]) -> Result<Vec<String>, Error> {
self.retry(|| {
with_bytes!(&self.data, |d| crate::vlen::decode_strings(
d,
datatype,
@@ -672,6 +950,7 @@ impl File {
self.offset_size(),
self.length_size(),
))
})
}
/// Like [`decode_strings`](Self::decode_strings) for variable-length
@@ -682,6 +961,7 @@ impl File {
datatype: &Datatype,
raw: &[u8],
) -> Result<Vec<Vec<u8>>, Error> {
self.retry(|| {
with_bytes!(self.data.meta()?, |d| crate::vlen::decode_string_bytes(
d,
datatype,
@@ -689,6 +969,7 @@ impl File {
self.offset_size(),
self.length_size(),
))
})
}
/// Decode the variable-length sequences in `raw`, a buffer of elements
@@ -700,6 +981,7 @@ impl File {
datatype: &Datatype,
raw: &[u8],
) -> Result<Vec<Vec<T>>, Error> {
self.retry(|| {
with_bytes!(self.data.meta()?, |d| crate::vlen::decode_vlen(
d,
datatype,
@@ -707,6 +989,7 @@ impl File {
self.offset_size(),
self.length_size(),
))
})
}
fn parse_header(&self, address: u64) -> Result<ObjectHeader, FormatError> {
@@ -750,6 +1033,7 @@ pub struct Group<'f> {
impl<'f> Group<'f> {
/// List the names of datasets in this group.
pub fn datasets(&self) -> Result<Vec<String>, Error> {
self.file.retry(|| {
let entries = self.children()?;
let mut names = Vec::new();
for entry in &entries {
@@ -759,10 +1043,12 @@ impl<'f> Group<'f> {
}
}
Ok(names)
})
}
/// List the names of subgroups in this group.
pub fn groups(&self) -> Result<Vec<String>, Error> {
self.file.retry(|| {
let entries = self.children()?;
let mut names = Vec::new();
for entry in &entries {
@@ -772,6 +1058,7 @@ impl<'f> Group<'f> {
}
}
Ok(names)
})
}
/// Read all attributes of this group.
@@ -792,26 +1079,31 @@ impl<'f> Group<'f> {
pub fn attrs_with_errors(
&self,
) -> Result<(HashMap<String, AttrValue>, Vec<FormatError>), Error> {
self.file.retry_attrs(|| {
let hdr = self.file.parse_header(self.address)?;
with_bytes!(&self.file.data, |d| read_attrs(
d,
&hdr,
self.file.offset_size(),
self.file.length_size()
))
let (attrs, errors, read_errors) = with_bytes!(&self.file.data, |d| {
read_attrs_reporting(d, &hdr, self.file.offset_size(), self.file.length_size())
})?;
let transient = errors.iter().chain(&read_errors).cloned().collect();
Ok(((attrs, errors), transient))
})
}
/// Get a dataset within this group by name.
pub fn dataset(&self, name: &str) -> Result<Dataset<'f>, Error> {
let hdr = self.file.parse_header(self.child_address(name)?)?;
self.file.retry(|| {
let address = self.child_address(name)?;
let hdr = self.file.parse_header(address)?;
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(name.to_string()));
}
Dataset {
file: self.file,
address,
header: hdr,
}
.check_open()
})
}
/// Get a subgroup within this group by name.
@@ -827,14 +1119,16 @@ impl<'f> Group<'f> {
/// that name, found without reading the other attributes when they are
/// stored densely.
pub fn attr(&self, name: &str) -> Result<Option<AttrValue>, Error> {
self.file.retry_attrs(|| {
let hdr = self.file.parse_header(self.address)?;
with_bytes!(&self.file.data, |d| read_attr(
with_bytes!(&self.file.data, |d| read_attr_reporting(
d,
&hdr,
name,
self.file.offset_size(),
self.file.length_size(),
))
})
}
/// The object header address of the child called `name`: the entry of
@@ -842,6 +1136,7 @@ impl<'f> Group<'f> {
/// name index rather than by listing the group (see
/// [`group_v2::resolve_child`]).
fn child_address(&self, name: &str) -> Result<u64, Error> {
self.file.retry(|| {
with_bytes!(self.file.data.meta()?, |d| group_v2::resolve_child_in(
d,
&self.file.superblock,
@@ -849,6 +1144,7 @@ impl<'f> Group<'f> {
name
))
.map_err(Error::Format)
})
}
/// This group's children that can be opened, as `(name, object header
@@ -869,10 +1165,12 @@ impl<'f> Group<'f> {
/// [`group_v2::resolve_group_children`]); dangling, external and
/// user-defined links are left out.
fn children(&self) -> Result<Vec<GroupEntry>, Error> {
self.file.retry(|| {
with_bytes!(self.file.data.meta()?, |d| {
group_v2::resolve_group_children_in(d, &self.file.superblock, self.address)
})
.map_err(Error::Format)
})
}
}
@@ -884,6 +1182,8 @@ impl<'f> Group<'f> {
#[derive(Debug)]
pub struct Dataset<'f> {
file: &'f File,
/// Address of the object header, for [`Dataset::refresh`].
address: u64,
header: ObjectHeader,
}
@@ -906,6 +1206,31 @@ impl<'f> Dataset<'f> {
Ok(self)
}
/// Read the dataset's object header again, as libhdf5's `H5Drefresh`
/// (h5py `Dataset.refresh()`) does: afterwards [`shape`](Self::shape)
/// and every read use the dataset's extent as it is in the file now.
/// For a file a SWMR writer is appending to ([`File::open_swmr`]) this
/// is how a reader sees the appends; a transient failure is retried
/// (see there). On error the handle keeps its old header.
pub fn refresh(&mut self) -> Result<(), Error> {
let file = self.file;
let address = self.address;
let fresh = file.retry(|| {
let hdr = file.parse_header(address)?;
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(format!("object at address {address}")));
}
Dataset {
file,
address,
header: hdr,
}
.check_open()
})?;
self.header = fresh.header;
Ok(())
}
/// Returns the shape (dimensions) of the dataset.
pub fn shape(&self) -> Result<Vec<u64>, Error> {
let ds = self.dataspace()?;
@@ -938,9 +1263,11 @@ impl<'f> Dataset<'f> {
/// Read all data as `f64` values.
pub fn read_f64(&self) -> Result<Vec<f64>, Error> {
self.file.retry(|| {
let dt = self.datatype()?;
// A contiguous dataset is converted straight from the file bytes; going
// through `read_raw` first copied the whole dataset an extra time.
// A contiguous dataset is converted straight from the file bytes;
// going through `read_raw` first copied the whole dataset an
// extra time.
if let Ok(Some(bytes)) = self.contiguous_raw() {
return Ok(data_read::read_as_f64(&bytes, &dt)?);
}
@@ -949,6 +1276,7 @@ impl<'f> Dataset<'f> {
}
let raw = self.read_raw()?;
Ok(data_read::read_as_f64(&raw, &dt)?)
})
}
/// Zero-copy read of contiguous native-endian `f64` data.
@@ -958,9 +1286,11 @@ impl<'f> Dataset<'f> {
///
/// Read all data as `f32` values.
pub fn read_f32(&self) -> Result<Vec<f32>, Error> {
self.file.retry(|| {
let dt = self.datatype()?;
// A contiguous dataset is converted straight from the file bytes; going
// through `read_raw` first copied the whole dataset an extra time.
// A contiguous dataset is converted straight from the file bytes;
// going through `read_raw` first copied the whole dataset an
// extra time.
if let Ok(Some(bytes)) = self.contiguous_raw() {
return Ok(data_read::read_as_f32(&bytes, &dt)?);
}
@@ -969,13 +1299,16 @@ impl<'f> Dataset<'f> {
}
let raw = self.read_raw()?;
Ok(data_read::read_as_f32(&raw, &dt)?)
})
}
/// Read all data as `i32` values.
pub fn read_i32(&self) -> Result<Vec<i32>, Error> {
self.file.retry(|| {
let dt = self.datatype()?;
// A contiguous dataset is converted straight from the file bytes; going
// through `read_raw` first copied the whole dataset an extra time.
// A contiguous dataset is converted straight from the file bytes;
// going through `read_raw` first copied the whole dataset an
// extra time.
if let Ok(Some(bytes)) = self.contiguous_raw() {
return Ok(data_read::read_as_i32(&bytes, &dt)?);
}
@@ -984,13 +1317,16 @@ impl<'f> Dataset<'f> {
}
let raw = self.read_raw()?;
Ok(data_read::read_as_i32(&raw, &dt)?)
})
}
/// Read all data as `i64` values.
pub fn read_i64(&self) -> Result<Vec<i64>, Error> {
self.file.retry(|| {
let dt = self.datatype()?;
// A contiguous dataset is converted straight from the file bytes; going
// through `read_raw` first copied the whole dataset an extra time.
// A contiguous dataset is converted straight from the file bytes;
// going through `read_raw` first copied the whole dataset an
// extra time.
if let Ok(Some(bytes)) = self.contiguous_raw() {
return Ok(data_read::read_as_i64(&bytes, &dt)?);
}
@@ -999,13 +1335,16 @@ impl<'f> Dataset<'f> {
}
let raw = self.read_raw()?;
Ok(data_read::read_as_i64(&raw, &dt)?)
})
}
/// Read all data as `u64` values.
pub fn read_u64(&self) -> Result<Vec<u64>, Error> {
self.file.retry(|| {
let dt = self.datatype()?;
// A contiguous dataset is converted straight from the file bytes; going
// through `read_raw` first copied the whole dataset an extra time.
// A contiguous dataset is converted straight from the file bytes;
// going through `read_raw` first copied the whole dataset an
// extra time.
if let Ok(Some(bytes)) = self.contiguous_raw() {
return Ok(data_read::read_as_u64(&bytes, &dt)?);
}
@@ -1014,6 +1353,7 @@ impl<'f> Dataset<'f> {
}
let raw = self.read_raw()?;
Ok(data_read::read_as_u64(&raw, &dt)?)
})
}
/// Read all data as `String` values, in row-major order.
@@ -1024,17 +1364,21 @@ impl<'f> Dataset<'f> {
/// them; bytes that are not valid UTF-8 are replaced with U+FFFD — use
/// [`read_string_bytes`](Self::read_string_bytes) for the exact bytes.
pub fn read_string(&self) -> Result<Vec<String>, Error> {
self.file.retry(|| {
let raw = self.read_raw()?;
let dt = self.datatype()?;
self.file.decode_strings(&dt, &raw)
})
}
/// Read a variable-length string dataset as the exact bytes of each
/// string (what h5py's `Dataset[()]` returns), in row-major order.
pub fn read_string_bytes(&self) -> Result<Vec<Vec<u8>>, Error> {
self.file.retry(|| {
let raw = self.read_raw()?;
let dt = self.datatype()?;
self.file.decode_string_bytes(&dt, &raw)
})
}
/// Read the selected elements of a fixed- or variable-length string
@@ -1043,9 +1387,11 @@ impl<'f> Dataset<'f> {
&self,
selection: &clawhdf5_format::selection::Selection,
) -> Result<Vec<String>, Error> {
self.file.retry(|| {
let raw = self.read_selection(selection)?;
let dt = self.datatype()?;
self.file.decode_strings(&dt, &raw)
})
}
/// Read a variable-length sequence dataset (h5py
@@ -1054,9 +1400,11 @@ impl<'f> Dataset<'f> {
/// is converted to `T` as [`read_f64`](Self::read_f64) and the other
/// typed readers convert. A null element is an empty sequence.
pub fn read_vlen<T: crate::vlen::VlenValue>(&self) -> Result<Vec<Vec<T>>, Error> {
self.file.retry(|| {
let raw = self.read_raw()?;
let dt = self.datatype()?;
self.file.decode_vlen(&dt, &raw)
})
}
/// Read the selected elements of a variable-length sequence dataset
@@ -1065,9 +1413,11 @@ impl<'f> Dataset<'f> {
&self,
selection: &clawhdf5_format::selection::Selection,
) -> Result<Vec<Vec<T>>, Error> {
self.file.retry(|| {
let raw = self.read_selection(selection)?;
let dt = self.datatype()?;
self.file.decode_vlen(&dt, &raw)
})
}
// ----- Selection-based read methods -----
@@ -1096,6 +1446,15 @@ impl<'f> Dataset<'f> {
if matches!(selection, clawhdf5_format::selection::Selection::All) {
return self.read_raw();
}
self.file.retry(|| self.read_selection_once(selection))
}
/// [`read_selection`](Self::read_selection) of a selection other than
/// `All`, once.
fn read_selection_once(
&self,
selection: &clawhdf5_format::selection::Selection,
) -> Result<Vec<u8>, Error> {
let dt = self.datatype()?;
let ds = self.dataspace()?;
let dl = self.data_layout()?;
@@ -1182,6 +1541,17 @@ impl<'f> Dataset<'f> {
if matches!(selection, clawhdf5_format::selection::Selection::All) {
return full();
}
self.file
.retry(|| self.read_typed_selection_once(selection, convert))
}
/// [`read_typed_selection`](Self::read_typed_selection) of a selection
/// other than `All`, once.
fn read_typed_selection_once<T: data_read::NativeElement>(
&self,
selection: &clawhdf5_format::selection::Selection,
convert: fn(&[u8], &Datatype) -> Result<Vec<T>, FormatError>,
) -> Result<Vec<T>, Error> {
let dt = self.datatype()?;
if T::is_native(&dt) && self.file.data.contiguous().is_some() {
if let Ok(Some(raw)) = self.read_raw_ref() {
@@ -1445,12 +1815,18 @@ impl<'f> Dataset<'f> {
pub fn attrs_with_errors(
&self,
) -> Result<(HashMap<String, AttrValue>, Vec<FormatError>), Error> {
with_bytes!(&self.file.data, |d| read_attrs(
self.file.retry_attrs(|| {
let (attrs, errors, read_errors) = with_bytes!(&self.file.data, |d| {
read_attrs_reporting(
d,
&self.header,
self.file.offset_size(),
self.file.length_size(),
))
)
})?;
let transient = errors.iter().chain(&read_errors).cloned().collect();
Ok(((attrs, errors), transient))
})
}
/// The attribute called `name`, or `None` if it has none by that name
@@ -1458,13 +1834,15 @@ impl<'f> Dataset<'f> {
/// that name, found without reading the other attributes when they are
/// stored densely.
pub fn attr(&self, name: &str) -> Result<Option<AttrValue>, Error> {
with_bytes!(&self.file.data, |d| read_attr(
self.file.retry_attrs(|| {
with_bytes!(&self.file.data, |d| read_attr_reporting(
d,
&self.header,
name,
self.file.offset_size(),
self.file.length_size(),
))
})
}
/// Verify this dataset's content against its stored provenance hash
@@ -1484,12 +1862,14 @@ impl<'f> Dataset<'f> {
/// result is not a tamper-evidence or authenticity guarantee.
#[cfg(feature = "provenance")]
pub fn verify_provenance(&self) -> Result<clawhdf5_format::provenance::VerifyResult, Error> {
self.file.retry(|| {
Ok(clawhdf5_format::provenance::verify_dataset_in(
&self.file.data,
&self.header,
self.file.offset_size(),
self.file.length_size(),
)?)
})
}
/// A header message's payload, resolved through the shared-message
@@ -1499,11 +1879,10 @@ impl<'f> Dataset<'f> {
&self,
msg_type: MessageType,
) -> Result<Option<std::borrow::Cow<'_, [u8]>>, Error> {
self.header
.messages
.iter()
.find(|m| m.msg_type == msg_type)
.map(|msg| {
let msg = self.header.messages.iter().find(|m| m.msg_type == msg_type);
// A shared message is read from the header it lives in.
self.file.retry(|| {
msg.map(|msg| {
with_bytes!(&self.file.data, |d| {
clawhdf5_format::shared_message::message_data_in(
d,
@@ -1515,6 +1894,7 @@ impl<'f> Dataset<'f> {
.map_err(Error::Format)
})
.transpose()
})
}
fn required_payload(&self, msg_type: MessageType) -> Result<std::borrow::Cow<'_, [u8]>, Error> {
@@ -1587,6 +1967,8 @@ impl<'f> Dataset<'f> {
}
let ds = self.dataspace()?;
let pipeline = self.filter_pipeline()?;
self.file.retry(|| {
self.file.with_chunk_cache(|cache| {
Ok(data_read::read_chunked_native_in::<T, _>(
&self.header.messages,
&self.file.data,
@@ -1596,11 +1978,17 @@ impl<'f> Dataset<'f> {
pipeline.as_ref(),
self.file.offset_size(),
self.file.length_size(),
Some(&self.file.chunk_cache),
Some(cache),
)?)
})
})
}
fn read_raw(&self) -> Result<Vec<u8>, Error> {
self.file.retry(|| self.read_raw_once())
}
fn read_raw_once(&self) -> Result<Vec<u8>, Error> {
let dt = self.datatype()?;
let ds = self.dataspace()?;
let dl = self.data_layout()?;
@@ -1622,6 +2010,7 @@ impl<'f> Dataset<'f> {
self.file.offset_size(),
self.file.length_size(),
|| {
self.file.with_chunk_cache(|cache| {
Ok(data_read::read_raw_data_cached_in(
&self.file.data,
&dl,
@@ -1630,8 +2019,9 @@ impl<'f> Dataset<'f> {
pipeline.as_ref(),
self.file.offset_size(),
self.file.length_size(),
&self.file.chunk_cache,
cache,
)?)
})
},
)
}
+362
View File
@@ -0,0 +1,362 @@
//! Reading files a SWMR writer is still appending to (see
//! `docs/design/swmr.md` and [`File::open_swmr`](crate::File::open_swmr)):
//! a [`FileStorage`] whose length grows with the file, and the bounded
//! retries of operations a concurrent write can make fail.
use std::borrow::Cow;
use std::cell::Cell;
use std::path::Path;
use std::sync::atomic::{AtomicU64, Ordering};
use std::time::Duration;
use clawhdf5_format::error::FormatError;
use clawhdf5_format::storage::Storage;
use crate::error::Error;
/// How many times a live file ([`File::open_swmr`](crate::File::open_swmr))
/// tries an operation that fails with an error a concurrent write can cause
/// before returning the error: 100, libhdf5's default number of metadata
/// read attempts for SWMR access (`H5Pset_metadata_read_attempts`).
pub const SWMR_READ_ATTEMPTS: u32 = 100;
/// Longest pause between two attempts.
const MAX_PAUSE: Duration = Duration::from_millis(10);
/// A local file read with positioned reads (`pread` on Unix, `seek_read` on
/// Windows), never mapped, whose [`Storage::len`] is the file's length at
/// the time of the call: a [`Storage`] for a file that another process is
/// appending to. A read past the end is short, as the trait allows.
#[derive(Debug)]
pub struct FileStorage {
file: std::fs::File,
#[cfg(not(any(unix, windows)))]
lock: std::sync::Mutex<()>,
}
impl FileStorage {
/// Open the file at `path` for reading.
pub fn open<P: AsRef<Path>>(path: P) -> std::io::Result<Self> {
Ok(Self::new(std::fs::File::open(path)?))
}
/// A storage over an open file.
pub fn new(file: std::fs::File) -> Self {
Self {
file,
#[cfg(not(any(unix, windows)))]
lock: std::sync::Mutex::new(()),
}
}
/// Up to `buf.len()` bytes at `offset`; fewer only at the end of file.
fn read_into(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
let mut got = 0;
while got < buf.len() {
match self.read_once(offset + got as u64, &mut buf[got..]) {
Ok(0) => break,
Ok(n) => got += n,
Err(e) if e.kind() == std::io::ErrorKind::Interrupted => {}
Err(e) => return Err(e),
}
}
Ok(got)
}
#[cfg(unix)]
fn read_once(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
std::os::unix::fs::FileExt::read_at(&self.file, buf, offset)
}
#[cfg(windows)]
fn read_once(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
std::os::windows::fs::FileExt::seek_read(&self.file, buf, offset)
}
#[cfg(not(any(unix, windows)))]
fn read_once(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
use std::io::{Read, Seek, SeekFrom};
let _guard = self.lock.lock().unwrap_or_else(|p| p.into_inner());
let mut f = &self.file;
f.seek(SeekFrom::Start(offset))?;
f.read(buf)
}
}
impl Storage for FileStorage {
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
// Never allocate more than the file holds, whatever a (possibly
// hostile) size field asked for.
let avail = self.len().saturating_sub(offset);
let len = usize::try_from(avail).map_or(len, |a| a.min(len));
let mut buf = vec![0u8; len];
let got = self
.read_into(offset, &mut buf)
.map_err(|e| FormatError::Storage(format!("read of {len} bytes at {offset}: {e}")))?;
buf.truncate(got);
Ok(Cow::Owned(buf))
}
fn len(&self) -> u64 {
self.file.metadata().map_or(0, |m| m.len())
}
}
/// Whether `e` can be caused by reading a structure while a SWMR writer
/// rewrites it or has not yet written it, so that the operation is worth
/// running again (see [`is_transient_format`]).
pub(crate) fn is_transient(e: &Error) -> bool {
match e {
Error::Format(f) => is_transient_format(f),
Error::Io(io) => io.kind() == std::io::ErrorKind::UnexpectedEof,
_ => false,
}
}
/// The failures libhdf5's SWMR reader retries, and nothing else. libhdf5
/// (`H5C__load_entry`) reads a metadata structure again when its checksum
/// fails, and when the prefix it decodes before the checksum to learn the
/// structure's size does not decode (for an object header, its signature
/// and version: a header whose every byte is garbled fails there). A read
/// past the file's current end is short here; libhdf5 reads zeros there,
/// which then fail the checksum. So:
///
/// - [`FormatError::ChecksumMismatch`], of any checksummed structure;
/// - [`FormatError::UnexpectedEof`], a read past the current end;
/// - [`FormatError::InvalidObjectHeaderSignature`] and
/// [`FormatError::InvalidObjectHeaderVersion`], the object header prefix.
///
/// Every other error — an unsupported version or message, a file that is not
/// HDF5, a structure that is corrupt behind a valid checksum — is returned
/// at once: a concurrent write does not cause it, and retrying it only
/// costs up to a second of pauses.
pub(crate) fn is_transient_format(e: &FormatError) -> bool {
use FormatError as F;
matches!(
e,
F::ChecksumMismatch { .. }
| F::UnexpectedEof { .. }
| F::InvalidObjectHeaderSignature
| F::InvalidObjectHeaderVersion(_)
)
}
thread_local! {
/// Set while this thread runs an operation under [`retry`], so an
/// operation made of retried operations retries as a whole, not each
/// part up to the limit.
static RETRYING: Cell<bool> = const { Cell::new(false) };
}
/// A live file's retry policy: attempts per operation, and a count of the
/// retries made.
#[derive(Debug)]
pub(crate) struct Retries {
pub(crate) attempts: u32,
retried: AtomicU64,
}
impl Default for Retries {
fn default() -> Self {
Self {
attempts: SWMR_READ_ATTEMPTS,
retried: AtomicU64::new(0),
}
}
}
impl Retries {
/// [`retry`] with this policy, counting the retries.
pub(crate) fn retry<T>(&self, op: impl FnMut() -> Result<T, Error>) -> Result<T, Error> {
retry(self.attempts, &self.retried, op)
}
/// Retries made so far.
pub(crate) fn retries(&self) -> u64 {
self.retried.load(Ordering::Relaxed)
}
}
/// Run `op` up to `attempts` times while it fails with a transient error
/// ([`is_transient`]), pausing 1 µs, 2 µs, 4 µs, … up to 10 ms between
/// attempts, and adding each retry to `retried`; return its first success
/// or last error. Inside another `retry` on the same thread, `op` runs
/// once.
fn retry<T>(
attempts: u32,
retried: &AtomicU64,
mut op: impl FnMut() -> Result<T, Error>,
) -> Result<T, Error> {
if RETRYING.with(Cell::get) {
return op();
}
struct Reset;
impl Drop for Reset {
fn drop(&mut self) {
RETRYING.with(|r| r.set(false));
}
}
RETRYING.with(|r| r.set(true));
let _reset = Reset;
let mut pause = Duration::from_micros(1);
let mut attempt = 1;
loop {
match op() {
Err(e) if attempt < attempts && is_transient(&e) => {
retried.fetch_add(1, Ordering::Relaxed);
std::thread::sleep(pause);
pause = (pause * 2).min(MAX_PAUSE);
attempt += 1;
}
result => return result,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn file_storage_reads_what_the_file_holds_now() {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("grow.bin");
std::fs::write(&path, b"hello").unwrap();
let s = FileStorage::open(&path).unwrap();
assert_eq!(s.len(), 5);
assert_eq!(&*s.read_at(1, 3).unwrap(), b"ell");
assert_eq!(&*s.read_at(3, 10).unwrap(), b"lo");
assert!(s.read_at(9, 4).unwrap().is_empty());
// The file grows after the storage was opened.
use std::io::Write;
std::fs::OpenOptions::new()
.append(true)
.open(&path)
.unwrap()
.write_all(b", world")
.unwrap();
assert_eq!(s.len(), 12);
assert_eq!(&*s.read_at(3, 100).unwrap(), b"lo, world");
}
#[test]
fn permanent_errors_are_returned_at_once() {
// Errors a concurrent write does not cause: none is retried, so
// none costs the pauses (about 0.9 s for 100 attempts).
let permanent = [
FormatError::SignatureNotFound,
FormatError::UnsupportedVersion(9),
FormatError::UnsupportedMessage(0x99),
FormatError::TruncatedFile {
stored_eof: 10,
actual_len: 5,
},
FormatError::InvalidDatatypeClass(15),
FormatError::InvalidLayoutVersion(9),
FormatError::InvalidBTreeSignature,
FormatError::ChunkedReadError("x".into()),
FormatError::DecompressionError("x".into()),
FormatError::Fletcher32Mismatch {
expected: 1,
computed: 2,
},
];
let n = AtomicU64::new(0);
let started = std::time::Instant::now();
for e in permanent {
assert!(!is_transient_format(&e), "{e:?}");
let mut calls = 0;
let r: Result<(), Error> = retry(SWMR_READ_ATTEMPTS, &n, || {
calls += 1;
Err(Error::Format(e.clone()))
});
assert!(r.is_err());
assert_eq!(calls, 1, "{e:?}");
}
assert_eq!(n.load(Ordering::Relaxed), 0);
assert!(started.elapsed() < Duration::from_millis(50));
assert!(!is_transient(&Error::Io(std::io::Error::other("x"))));
// The transient ones are retried to the limit.
for e in [
FormatError::ChecksumMismatch {
expected: 1,
computed: 2,
},
FormatError::UnexpectedEof {
expected: 8,
available: 0,
},
FormatError::InvalidObjectHeaderSignature,
FormatError::InvalidObjectHeaderVersion(0x4f),
] {
let mut calls = 0;
let _: Result<(), Error> = retry(3, &n, || {
calls += 1;
Err(Error::Format(e.clone()))
});
assert_eq!(calls, 3, "{e:?}");
}
assert!(is_transient(&Error::Io(
std::io::ErrorKind::UnexpectedEof.into()
)));
}
#[test]
fn retry_runs_again_only_for_transient_errors() {
let n = AtomicU64::new(0);
let mut calls = 0;
let r: Result<u32, Error> = retry(5, &n, || {
calls += 1;
if calls < 3 {
Err(Error::Format(FormatError::ChecksumMismatch {
expected: 1,
computed: 2,
}))
} else {
Ok(7)
}
});
assert_eq!(r.unwrap(), 7);
assert_eq!(calls, 3);
assert_eq!(n.load(Ordering::Relaxed), 2);
// Gives up after `attempts`.
let mut calls = 0;
let r: Result<(), Error> = retry(4, &n, || {
calls += 1;
Err(Error::Format(FormatError::UnexpectedEof {
expected: 8,
available: 0,
}))
});
assert!(r.is_err());
assert_eq!(calls, 4);
assert_eq!(n.load(Ordering::Relaxed), 5);
// A permanent error is returned at once.
let mut calls = 0;
let r: Result<(), Error> = retry(4, &n, || {
calls += 1;
Err(Error::Format(FormatError::UnsupportedFilter(999)))
});
assert!(r.is_err());
assert_eq!(calls, 1);
assert_eq!(n.load(Ordering::Relaxed), 5);
// Nested: the inner operation runs once per outer attempt.
let mut inner = 0;
let mut outer = 0;
let r: Result<(), Error> = retry(3, &n, || {
outer += 1;
retry(3, &n, || {
inner += 1;
Err(Error::Format(FormatError::InvalidObjectHeaderVersion(0x76)))
})
});
assert!(r.is_err());
assert_eq!((outer, inner), (3, 3));
// The flag is reset afterwards.
assert!(!RETRYING.with(Cell::get));
}
}
+63 -18
View File
@@ -170,16 +170,38 @@ pub(crate) fn read_attrs<S: clawhdf5_format::storage::Storage + ?Sized>(
),
crate::Error,
> {
read_attrs_reporting(file_data, header, offset_size, length_size)
.map(|(attrs, errors, _)| (attrs, errors))
}
/// What [`read_attrs_reporting`] returns: the attributes, the errors of
/// those left out, and the errors behind values returned as
/// [`AttrValue::Raw`].
pub(crate) type AttrsReport = (
HashMap<String, AttrValue>,
Vec<clawhdf5_format::error::FormatError>,
Vec<clawhdf5_format::error::FormatError>,
);
/// [`read_attrs`], also returning the errors of reading the file for a
/// value (a variable-length string's global heap) that was returned as
/// [`AttrValue::Raw`] instead: a live file retries on them (see
/// `File::open_swmr`).
pub(crate) fn read_attrs_reporting<S: clawhdf5_format::storage::Storage + ?Sized>(
file_data: &S,
header: &clawhdf5_format::object_header::ObjectHeader,
offset_size: u8,
length_size: u8,
) -> Result<AttrsReport, crate::Error> {
let (msgs, errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
file_data,
header,
offset_size,
length_size,
)?;
Ok((
attrs_to_map(&msgs, file_data, offset_size, length_size),
errors,
))
let mut read_errors = Vec::new();
let map = attrs_to_map_reporting(&msgs, file_data, offset_size, length_size, &mut read_errors);
Ok((map, errors, read_errors))
}
/// The attribute called `name` on the object with header `header`, decoded
@@ -192,30 +214,50 @@ pub(crate) fn read_attr<S: clawhdf5_format::storage::Storage + ?Sized>(
offset_size: u8,
length_size: u8,
) -> Result<Option<AttrValue>, crate::Error> {
let Some(msg) = clawhdf5_format::attribute::find_attribute_in(
read_attr_reporting(file_data, header, name, offset_size, length_size).map(|(v, _)| v)
}
/// [`read_attr`], also returning the errors of the attributes it could not
/// read on the way (the one asked for may be among them) and of reading the
/// file for its value, if one made it [`AttrValue::Raw`] (see
/// [`read_attrs_reporting`]).
pub(crate) fn read_attr_reporting<S: clawhdf5_format::storage::Storage + ?Sized>(
file_data: &S,
header: &clawhdf5_format::object_header::ObjectHeader,
name: &str,
offset_size: u8,
length_size: u8,
) -> Result<(Option<AttrValue>, Vec<clawhdf5_format::error::FormatError>), crate::Error> {
let (found, mut read_errors) = clawhdf5_format::attribute::find_attribute_reporting_in(
file_data,
header,
name,
offset_size,
length_size,
)?
else {
return Ok(None);
)?;
let Some(msg) = found else {
return Ok((None, read_errors));
};
Ok(attrs_to_map(
let value = attrs_to_map_reporting(
std::slice::from_ref(&msg),
file_data,
offset_size,
length_size,
&mut read_errors,
)
.remove(name))
.remove(name);
Ok((value, read_errors))
}
pub(crate) fn attrs_to_map<S: clawhdf5_format::storage::Storage + ?Sized>(
/// The attributes `attrs` by name, decoded; an error reading the file for
/// a value that is returned as [`AttrValue::Raw`] because of it is added
/// to `read_errors`.
pub(crate) fn attrs_to_map_reporting<S: clawhdf5_format::storage::Storage + ?Sized>(
attrs: &[clawhdf5_format::attribute::AttributeMessage],
file_data: &S,
offset_size: u8,
length_size: u8,
read_errors: &mut Vec<clawhdf5_format::error::FormatError>,
) -> HashMap<String, AttrValue> {
let mut map = HashMap::new();
for attr in attrs {
@@ -224,13 +266,11 @@ pub(crate) fn attrs_to_map<S: clawhdf5_format::storage::Storage + ?Sized>(
// verbatim as `AttrValue::Raw` rather than dropped — a partial
// attribute list with no indication anything is missing is worse than
// an undecoded value.
let val =
decode_attr_value(attr, file_data, offset_size, length_size).unwrap_or_else(|| {
AttrValue::Raw {
let val = decode_attr_value(attr, file_data, offset_size, length_size, read_errors)
.unwrap_or_else(|| AttrValue::Raw {
datatype: attr.datatype.clone(),
shape: attr.dataspace.dimensions.clone(),
data: attr.raw_data.clone(),
}
});
map.insert(attr.name.clone(), val);
}
@@ -271,6 +311,7 @@ fn decode_attr_value<S: clawhdf5_format::storage::Storage + ?Sized>(
file_data: &S,
offset_size: u8,
length_size: u8,
read_errors: &mut Vec<clawhdf5_format::error::FormatError>,
) -> Option<AttrValue> {
use clawhdf5_format::datatype::Datatype;
@@ -310,9 +351,13 @@ fn decode_attr_value<S: clawhdf5_format::storage::Storage + ?Sized>(
Datatype::VariableLength {
is_string: true, ..
} => {
let strings = attr
.read_vl_strings_in(file_data, offset_size, length_size)
.ok()?;
let strings = match attr.read_vl_strings_in(file_data, offset_size, length_size) {
Ok(strings) => strings,
Err(e) => {
read_errors.push(e);
return None;
}
};
if strings.len() == 1 {
Some(AttrValue::String(strings[0].clone()))
} else {
@@ -0,0 +1,287 @@
//! `FileEditor::resize` on chunked datasets whose dataspace records no
//! maximum dimensions, as clawhdf5's writer stored a dataset created without
//! a `maxshape` up to 2.7.0 (`fixtures/chunked_no_maxshape_v2_7_0.h5`). libhdf5 never writes such a dataspace (`H5S_set_extent_simple`
//! always records the maximum, equal to the dimensions when none is given),
//! and its Fixed Array chunk index linearises chunks by the maximum
//! dimensions. The editor therefore records the maximum libhdf5 would have
//! written (the dimensions the index was built with) before it changes the
//! current ones, so existing chunks stay where the index put them and the
//! dataset can grow back to its original extent.
//!
//! The writer now records the maximum too, so h5py can resize what it writes.
//!
//! Checked against a model of the expected values, with our reader and with
//! h5py (`CLAWHDF5_PYTHON`; skipped without it unless
//! `CLAWHDF5_REQUIRE_INTEROP=1`), on files clawhdf5 (old and new) and h5py
//! wrote.
use std::path::{Path, PathBuf};
use std::process::Command;
use clawhdf5::{Error, File, FileBuilder, FileEditor};
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn h5py_ok() -> bool {
let ok = Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.is_ok_and(|o| o.status.success());
if !ok {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but h5py/numpy is not available"
);
eprintln!("SKIP (h5py part): h5py/numpy not available");
}
ok
}
fn py(script: &str) -> String {
let o = Command::new(python())
.args(["-c", script])
.output()
.expect("run python");
assert!(
o.status.success(),
"python failed:\n{script}\nSTDERR: {}",
String::from_utf8_lossy(&o.stderr)
);
String::from_utf8_lossy(&o.stdout).trim().to_string()
}
/// Row-major values of a 2-D model after resizing `data` (shape `old`) to
/// `new`: kept elements keep their values, new ones are 0 (the fill value).
fn resized(data: &[f32], old: [u64; 2], new: [u64; 2]) -> Vec<f32> {
let mut out = vec![0f32; (new[0] * new[1]) as usize];
for r in 0..old[0].min(new[0]) {
for c in 0..old[1].min(new[1]) {
out[(r * new[1] + c) as usize] = data[(r * old[1] + c) as usize];
}
}
out
}
/// Our reader and (when available) h5py read `expect` at `shape`.
fn check(path: &Path, name: &str, shape: [u64; 2], expect: &[f32], with_h5py: bool) {
let f = File::open(path).unwrap();
let d = f.dataset(name).unwrap();
assert_eq!(d.shape().unwrap(), shape);
assert_eq!(
d.read_f32().unwrap(),
expect,
"{name}: our reader at {shape:?}"
);
if with_h5py {
let got = py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'r') as f:\n\
\x20 d = f[{name:?}][()]\n\
print(d.shape, ','.join(repr(float(x)) for x in d.ravel()))",
p = path.to_str().unwrap()
));
let want = format!(
"({}, {}) {}",
shape[0],
shape[1],
expect
.iter()
.map(|x| format!("{:?}", f64::from(*x)))
.collect::<Vec<_>>()
.join(",")
);
assert_eq!(got, want.trim(), "{name}: h5py at {shape:?}");
}
}
/// Resize `name` (20 x 20, values 0..400) through a sequence of shrinks,
/// zero extents and growth back, checking every step.
fn run(path: &Path, name: &str, with_h5py: bool) {
let orig: Vec<f32> = (0..400).map(|i| i as f32).collect();
let mut shape = [20u64, 20];
let mut data = orig.clone();
check(path, name, shape, &data, with_h5py);
for next in [
[15, 15],
[3, 2],
[20, 20],
[1, 1],
[1, 0],
[0, 0],
[7, 20],
[20, 13],
[20, 20],
] {
let mut ed = FileEditor::open(path).unwrap();
ed.resize(name, &next).unwrap();
drop(ed);
data = resized(&data, shape, next);
shape = next;
check(path, name, shape, &data, with_h5py);
}
// The maximum is the extent the dataset was created with.
let mut ed = FileEditor::open(path).unwrap();
assert!(matches!(
ed.resize(name, &[21, 20]),
Err(Error::InvalidArgument(_))
));
drop(ed);
let f = File::open(path).unwrap();
assert_eq!(
f.dataset(name).unwrap().max_dimensions().unwrap(),
Some(vec![20, 20])
);
drop(f);
// A shrink keeps the values it keeps.
let mut ed = FileEditor::open(path).unwrap();
let vals: Vec<f32> = orig.iter().map(|v| v + 0.5).collect();
ed.write_values(name, &clawhdf5::Selection::All, &vals)
.unwrap();
ed.resize(name, &[15, 15]).unwrap();
drop(ed);
check(
path,
name,
[15, 15],
&resized(&vals, [20, 20], [15, 15]),
with_h5py,
);
}
fn fixture(dir: &Path, name: &str) -> PathBuf {
let path = dir.join(name);
std::fs::copy(
Path::new(env!("CARGO_MANIFEST_DIR"))
.join("tests/fixtures")
.join(name),
&path,
)
.unwrap();
path
}
/// Files clawhdf5 2.7.0 wrote: a Fixed Array index (Single Chunk for `s`)
/// and a dataspace with no maximum. Shrinking scrambled the values
/// (released in 2.7.0's `FileEditor`, PR #18).
#[test]
fn resize_without_stored_maxshape_keeps_values() {
let with_h5py = h5py_ok();
for name in ["d", "z"] {
let dir = tempfile::tempdir().unwrap();
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
run(&path, name, with_h5py);
}
// A single-chunk dataset and an empty one keep their extents as maxima.
let dir = tempfile::tempdir().unwrap();
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
let mut ed = FileEditor::open(&path).unwrap();
ed.resize("s", &[2, 4]).unwrap();
ed.resize("s", &[3, 4]).unwrap();
assert!(matches!(
ed.resize("s", &[4, 4]),
Err(Error::InvalidArgument(_))
));
assert!(matches!(
ed.resize("e", &[1, 5]),
Err(Error::InvalidArgument(_))
));
ed.resize("e", &[0, 3]).unwrap();
drop(ed);
let f = File::open(&path).unwrap();
let s = f.dataset("s").unwrap();
assert_eq!(s.max_dimensions().unwrap(), Some(vec![3, 4]));
let mut want: Vec<i32> = (0..12).collect();
want[8..].fill(0);
assert_eq!(s.read_i32().unwrap(), want);
assert_eq!(
f.dataset("e").unwrap().max_dimensions().unwrap(),
Some(vec![0, 5])
);
}
fn written(path: &Path, deflate: bool) {
let data: Vec<f32> = (0..400).map(|i| i as f32).collect();
let mut b = FileBuilder::new();
let d = b
.create_dataset("d")
.with_f32_data(&data)
.with_shape(&[20, 20])
.with_chunks(&[6, 6]);
if deflate {
d.with_deflate(4);
}
b.write(path).unwrap();
}
/// clawhdf5's writer now records the maximum, as libhdf5 does.
#[test]
fn resize_file_written_without_maxshape_keeps_values() {
let with_h5py = h5py_ok();
for deflate in [false, true] {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("cw.h5");
written(&path, deflate);
let f = File::open(&path).unwrap();
assert_eq!(
f.dataset("d").unwrap().max_dimensions().unwrap(),
Some(vec![20, 20])
);
drop(f);
run(&path, "d", with_h5py);
}
}
/// h5py resizing a file clawhdf5 wrote without a maxshape keeps its values
/// (it scrambled them while the writer recorded no maximum).
#[test]
fn h5py_resizes_what_clawhdf5_writes() {
if !h5py_ok() {
return;
}
for deflate in [false, true] {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("cw.h5");
written(&path, deflate);
let out = py(&format!(
"import h5py, numpy as np\n\
exp = np.arange(400, dtype='f4').reshape(20, 20)\n\
with h5py.File({p:?}, 'r+') as f:\n\
\x20 f['d'].resize((15, 15))\n\
\x20 ok = np.array_equal(f['d'][()], exp[:15, :15])\n\
\x20 f['d'].resize((20, 20))\n\
\x20 back = f['d'][()]\n\
want = np.zeros((20, 20), 'f4'); want[:15, :15] = exp[:15, :15]\n\
print(ok and np.array_equal(back, want))",
p = path.to_str().unwrap()
));
assert_eq!(out, "True");
let mut want = vec![0f32; 400];
for r in 0..15 {
for c in 0..15 {
want[r * 20 + c] = (r * 20 + c) as f32;
}
}
check(&path, "d", [20, 20], &want, false);
}
}
/// h5py's files record the maximum; the same sequence must hold.
#[test]
fn resize_h5py_file_without_maxshape_keeps_values() {
if !h5py_ok() {
return;
}
for libver in ["earliest", "v110", "latest"] {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hp.h5");
py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'w', libver=({libver:?}, 'latest')) as f:\n\
\x20 f.create_dataset('d', data=np.arange(400, dtype='f4').reshape(20, 20), chunks=(6, 6))",
p = path.to_str().unwrap()
));
run(&path, "d", true);
}
}
+56
View File
@@ -162,3 +162,59 @@ fn shrink_then_grow_reads_fill() {
raw[..4].fill(0.5);
assert_eq!(f.dataset("raw").unwrap().read_f64().unwrap(), raw);
}
/// The editor plans every edit from the file it holds open, never by
/// re-opening its path: with the path renamed away and another file put in
/// its place, edits go to the held file, planned from its own metadata,
/// and the file now at the path is untouched (planning from it and writing
/// into the held file corrupted the held one).
#[test]
fn edits_go_to_the_file_held_not_the_path() {
let dir = tempfile::tempdir().unwrap();
let a = dir.path().join("a.h5");
let b = dir.path().join("b.h5");
let mut fb = FileBuilder::new();
fb.create_dataset("x")
.with_i32_data(&[0; 10])
.with_shape(&[10]);
fb.create_dataset("big")
.with_f64_data(&[1.5; 5000])
.with_shape(&[5000]);
fb.set_attr("title", AttrValue::String("a".into()));
fb.write(&a).unwrap();
let mut fb = FileBuilder::new();
fb.create_dataset("pad")
.with_f64_data(&[2.5; 3000])
.with_shape(&[3000]);
fb.create_dataset("x")
.with_i32_data(&[500; 10])
.with_shape(&[10]);
fb.write(&b).unwrap();
let mut ed = FileEditor::open(&a).unwrap();
assert!(ed.path().is_absolute());
let moved = dir.path().join("moved.h5");
std::fs::rename(&a, &moved).unwrap();
std::fs::rename(&b, &a).unwrap();
let other = std::fs::read(&a).unwrap();
ed.write_values("x", &Selection::All, &[7i32; 10]).unwrap();
let vals: Vec<f64> = (0..50).map(f64::from).collect();
ed.set_attr("/", "note", &AttrValue::F64Array(vals.clone()))
.unwrap();
// The editor's own reader sees the held file.
let r = ed.reader().unwrap();
assert_eq!(r.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
assert!(r.dataset("pad").is_err());
drop(r);
drop(ed);
assert!(
std::fs::read(&a).unwrap() == other,
"the file at the path changed"
);
let f = File::open(&moved).unwrap();
assert_eq!(f.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
assert_eq!(f.dataset("big").unwrap().read_f64().unwrap(), [1.5; 5000]);
assert!(matches!(f.root().attr("note").unwrap(), Some(AttrValue::F64Array(v)) if v == vals));
}
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -455,3 +455,70 @@ with h5py.File("{p}", "w") as f:
assert!(lazy.dataset("line").unwrap().read_i32().is_err());
assert!(lazy.dataset("grid").unwrap().read_f64().is_err());
}
// ---------------------------------------------------------------------------
// Several damaged chunks: which one is reported
// ---------------------------------------------------------------------------
/// With more than one damaged chunk, the error names the first damaged chunk
/// in the chunk index's order, every time and on every read path. The cached
/// reader used to walk the chunks in hash-map order, so two opens of the same
/// file could report different chunks (`cve-2025-2310.h5`).
#[test]
fn several_damaged_chunks_report_the_same_chunk_every_time() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("damaged.h5");
let p = path.display().to_string();
run_python(&format!(
r#"
import h5py, numpy as np, zlib
with h5py.File("{p}", "w") as f:
ds = f.create_dataset("d", shape=(512,), chunks=(8,), dtype="<f8", compression="gzip")
ds[...] = np.arange(512.0)
# Each damaged chunk inflates to a different short length, so each
# error names its own chunk.
for k, off in enumerate((40, 136, 320, 488)):
ds.id.write_direct_chunk((off,), zlib.compress(bytes(8 * (k + 1))))
"#
));
let first = File::open(&path)
.unwrap()
.dataset("d")
.unwrap()
.read_f64()
.expect_err("damaged chunks must fail")
.to_string();
// A whole-dataset selection goes through the file's chunk cache, whose index is built
// afresh (a new hash map) for every open.
let first_raw = File::open(&path)
.unwrap()
.dataset("d")
.unwrap()
.read_selection(&Selection::All)
.expect_err("damaged chunks must fail")
.to_string();
for _ in 0..40 {
let f = File::open(&path).unwrap();
let ds = f.dataset("d").unwrap();
assert_eq!(ds.read_f64().unwrap_err().to_string(), first);
assert_eq!(
ds.read_selection(&Selection::All).unwrap_err().to_string(),
first_raw
);
// A second read goes through the cached chunk index.
assert_eq!(
ds.read_selection(&Selection::All).unwrap_err().to_string(),
first_raw
);
}
let lazy = clawhdf5::LazyFile::open_mmap(&path).unwrap();
assert_eq!(
lazy.dataset("d")
.unwrap()
.read_f64()
.unwrap_err()
.to_string(),
first
);
}
@@ -966,7 +966,10 @@ fn skipped_optional_filters_are_masked_as_libhdf5_masks_them() {
/// Files whose chunks all compress are written exactly as before optional
/// filters could be skipped: every mask is 0 and nothing else changed. The
/// hashes are of the files the writer produced before that change.
/// hashes are of the files the writer produced before that change, except
/// that a chunked dataset without a maxshape now records its maximum
/// dimensions (8 bytes per dimension; `lzf_fixed`, `lzf_single`, and
/// `blosc_fixed`).
#[cfg(feature = "lzf")]
#[test]
fn files_whose_chunks_all_compress_are_unchanged() {
@@ -985,7 +988,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
.with_chunks(&[500])
.with_lzf();
},
(3965, 449169442),
(3973, 452644487),
),
(
"lzf_ea_noshuffle",
@@ -1017,7 +1020,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
.with_chunks(&[3000])
.with_lzf();
},
(546, 690805477),
(554, 1394027497),
),
];
#[cfg(feature = "blosc")]
@@ -1029,7 +1032,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
.with_chunks(&[1024])
.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
},
(2776, 4278611376),
(2784, 1180133244),
));
for (name, build, want) in &cases {
let mut fb = clawhdf5::FileBuilder::new();
+847
View File
@@ -0,0 +1,847 @@
//! Files a libhdf5 SWMR writer (h5py `f.swmr_mode = True`) has open: a copy
//! taken mid-write (fixture), and a live file appended to by an h5py writer
//! process while clawhdf5 and h5py's own SWMR reader read it (see
//! `docs/design/swmr.md`).
//!
//! The live tests need python3 with h5py; they are skipped without it,
//! unless `CLAWHDF5_REQUIRE_INTEROP=1`.
use std::io::{BufRead, BufReader, Read};
use std::path::{Path, PathBuf};
use std::process::{Command, Stdio};
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
use clawhdf5::{File, Selection};
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn interop_required() -> bool {
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
}
fn python_available() -> bool {
Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
macro_rules! skip_if_no_python {
() => {
if !python_available() {
assert!(
!interop_required(),
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
);
eprintln!("SKIP: python3 with h5py not available");
return;
}
};
}
fn fixture(name: &str) -> PathBuf {
Path::new(env!("CARGO_MANIFEST_DIR"))
.join("tests/fixtures")
.join(name)
}
/// `swmr_mid_write.h5`: a copy h5py 3.16 (HDF5 2.0) made of its own file
/// while writing it in SWMR mode, after 4 appends of 37 rows to
/// `/a` (int64, chunks of 100, no filter: `a[i] = i + 1`) and `/b` (float64
/// `(n, 4)`, chunks of 16 x 4, gzip: `b[i, j] = 10 i + j + 1`), each
/// followed by a flush. Its superblock (v3) still has the SWMR-write flag
/// set and records an end of file of 715 in a 6 030-byte file.
fn check_mid_write_copy(f: &File) {
let sb = f.superblock();
assert_eq!(sb.version, 3);
assert!(sb.is_swmr_write());
let a = f.dataset("a").unwrap();
assert_eq!(a.shape().unwrap(), vec![148]);
let want_a: Vec<i64> = (1..=148).collect();
assert_eq!(a.read_i64().unwrap(), want_a);
let b = f.dataset("b").unwrap();
assert_eq!(b.shape().unwrap(), vec![148, 4]);
let want_b: Vec<f64> = (0..148)
.flat_map(|i| (0..4).map(move |j| (10 * i + j + 1) as f64))
.collect();
assert_eq!(b.read_f64().unwrap(), want_b);
}
#[test]
fn a_copy_taken_mid_write_reads_past_its_recorded_end_of_file() {
let path = fixture("swmr_mid_write.h5");
// The recorded end of file (715) is far below the file's length; the
// chunk indexes and chunks lie past it. libhdf5's SWMR reader does not
// bound reads by it, and neither does any open path here.
check_mid_write_copy(&File::open(&path).unwrap());
check_mid_write_copy(&File::open_buffered(&path).unwrap());
check_mid_write_copy(&File::from_bytes(std::fs::read(&path).unwrap()).unwrap());
let want_a: Vec<i64> = (1..=148).collect();
let want_b: Vec<f64> = (0..148)
.flat_map(|i| (0..4).map(move |j| (10 * i + j + 1) as f64))
.collect();
let mm = clawhdf5::MmapFile::open(&path).unwrap();
assert_eq!(mm.dataset("a").unwrap().read_i64().unwrap(), want_a);
assert_eq!(mm.dataset("b").unwrap().read_f64().unwrap(), want_b);
let lazy = clawhdf5::LazyFile::open_mmap(&path).unwrap();
assert_eq!(lazy.dataset("a").unwrap().read_i64().unwrap(), want_a);
assert_eq!(lazy.dataset("b").unwrap().read_f64().unwrap(), want_b);
let storage = File::open_storage(Arc::new(std::fs::read(&path).unwrap())).unwrap();
check_mid_write_copy(&storage);
}
#[test]
fn a_copy_taken_mid_write_reads_as_h5py_swmr_reader_reads_it() {
skip_if_no_python!();
let path = fixture("swmr_mid_write.h5");
// libhdf5 refuses a non-SWMR open of this file ("file is already open
// for write"); its SWMR reader reads the values checked above.
let script = format!(
r#"
import h5py, numpy as np
with h5py.File("{p}", "r", swmr=True, locking=False) as f:
a = f["a"][()]
b = f["b"][()]
assert np.array_equal(a, np.arange(148) + 1), a
assert np.array_equal(b, (np.arange(148)[:, None] * 10 + np.arange(4) + 1).astype("f8")), b
print("ok")
"#,
p = path.display()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.unwrap();
assert!(
out.status.success(),
"h5py failed:\n{}",
String::from_utf8_lossy(&out.stderr)
);
check_mid_write_copy(&File::open(&path).unwrap());
}
#[test]
fn open_swmr_reads_the_mid_write_copy() {
let f = File::open_swmr(fixture("swmr_mid_write.h5")).unwrap();
assert!(f.is_swmr_read());
assert_eq!(f.swmr_read_attempts(), clawhdf5::SWMR_READ_ATTEMPTS);
// The copy still has the SWMR-write flag: as far as it says, its writer
// is still writing.
assert!(f.swmr_writer_active().unwrap());
check_mid_write_copy(&f);
let mut a = f.dataset("a").unwrap();
a.refresh().unwrap();
assert_eq!(a.shape().unwrap(), vec![148]);
assert_eq!(f.swmr_retries(), 0);
assert!(
!File::open(fixture("swmr_mid_write.h5"))
.unwrap()
.is_swmr_read()
);
}
/// The mid-write copy with its superblock's flags set to `flags` (and the
/// superblock checksum updated).
fn mid_write_copy_with_flags(flags: u8) -> Vec<u8> {
let mut bytes = std::fs::read(fixture("swmr_mid_write.h5")).unwrap();
// Superblock v3, 8-byte offsets: 12 bytes, 4 addresses, checksum.
bytes[11] = flags;
let sum = clawhdf5_format::checksum::jenkins_lookup3(&bytes[..44]);
bytes[44..48].copy_from_slice(&sum.to_le_bytes());
bytes
}
#[test]
fn open_swmr_refuses_a_file_open_for_writing_without_swmr() {
// libhdf5 refuses it too: "file is already open for write".
let bytes = mid_write_copy_with_flags(0x01);
let err = File::open_storage_swmr(Arc::new(bytes)).unwrap_err();
assert!(matches!(err, clawhdf5::Error::Locked(_)), "{err}");
// Closed (flags clear): opens, and no writer is active.
let f = File::open_storage_swmr(Arc::new(mid_write_copy_with_flags(0))).unwrap();
assert!(!f.swmr_writer_active().unwrap());
}
/// The outcome of reading every dataset of `f` in full, as text: the
/// values, or the error.
fn read_all(f: &File) -> Vec<String> {
["a", "b"]
.iter()
.map(|n| match f.dataset(n).and_then(|d| d.read_f64()) {
Ok(v) => format!("{n}: {} values", v.len()),
Err(e) => format!("{n}: error {e}"),
})
.collect()
}
#[test]
fn open_swmr_reads_a_file_without_the_swmr_write_flag_as_file_open_does() {
// The mid-write copy with its flags cleared: a closed file that
// records an end of file of 715 in 6 030 bytes, its chunk indexes past
// that end. File::open (and libhdf5's plain reader) refuse what lies
// past the recorded end; open_swmr must too, and must not retry (the
// file is not live). See the h5py test below for libhdf5's SWMR reader.
let bytes = mid_write_copy_with_flags(0);
let dir = tempfile::tempdir_in(env!("CARGO_TARGET_TMPDIR")).unwrap();
let path = dir.path().join("closed_short_eof.h5");
std::fs::write(&path, &bytes).unwrap();
let plain = File::open(&path).unwrap();
let swmr = File::open_swmr(&path).unwrap();
let swmr_storage = File::open_storage_swmr(Arc::new(bytes)).unwrap();
let want = read_all(&plain);
assert!(
want.iter().all(|r| r.contains("error")),
"File::open reads past the recorded end: {want:?}"
);
for f in [&swmr, &swmr_storage] {
assert!(!f.is_swmr_read());
assert!(!f.swmr_writer_active().unwrap());
let started = std::time::Instant::now();
assert_eq!(read_all(f), want);
assert!(started.elapsed() < std::time::Duration::from_millis(100));
assert_eq!(f.swmr_retries(), 0);
}
// The same bytes with the SWMR-write flag (and write access) set are
// read live, past the recorded end.
let live = File::open_storage_swmr(Arc::new(mid_write_copy_with_flags(0x05))).unwrap();
assert!(live.is_swmr_read());
check_mid_write_copy(&live);
}
#[test]
fn how_h5py_reads_the_closed_copy_past_its_recorded_end_of_file() {
skip_if_no_python!();
let dir = tempfile::tempdir_in(env!("CARGO_TARGET_TMPDIR")).unwrap();
let path = dir.path().join("closed_short_eof.h5");
std::fs::write(&path, mid_write_copy_with_flags(0)).unwrap();
// libhdf5's plain reader refuses `a` (its chunk index is past the
// recorded end of file), as File::open and open_swmr do above.
// libhdf5's SWMR reader skips its end-of-allocation check for every
// file it opens (H5FD_read), and reads it; it still refuses an object
// *header* past the end ("address of object past end of allocation",
// H5O_protect). open_swmr does not copy that half-way rule for files
// without the SWMR-write flag: they read as File::open reads them.
let script = format!(
r#"
import h5py
for swmr in (False, True):
try:
with h5py.File("{p}", "r", swmr=swmr, locking=False) as f:
f["a"][()]
except Exception as e:
print("refused", swmr)
else:
print("read", swmr)
"#,
p = path.display()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.unwrap();
let stdout = String::from_utf8_lossy(&out.stdout);
assert_eq!(
stdout.split_whitespace().collect::<Vec<_>>(),
["refused", "False", "read", "True"],
"{}",
String::from_utf8_lossy(&out.stderr)
);
}
#[test]
fn open_swmr_returns_a_permanent_error_at_once() {
// Not HDF5 at all: SignatureNotFound is not something a writer
// causes, so it is not retried (100 attempts would pause about 0.9 s).
let started = std::time::Instant::now();
let err = File::open_storage_swmr(Arc::new(vec![7u8; 4096])).unwrap_err();
assert!(
matches!(
err,
clawhdf5::Error::Format(clawhdf5_format::error::FormatError::SignatureNotFound)
),
"{err}"
);
assert!(started.elapsed() < std::time::Duration::from_millis(100));
// A live file: a lookup of a name it does not have fails at once, and
// a torn object header read (below) is retried.
let f = File::open_storage_swmr(Arc::new(mid_write_copy_with_flags(0x05))).unwrap();
assert!(f.is_swmr_read());
let started = std::time::Instant::now();
assert!(f.dataset("no_such").is_err());
assert!(started.elapsed() < std::time::Duration::from_millis(100));
assert_eq!(f.swmr_retries(), 0);
}
/// A storage whose next `torn` reads come back garbled, as a read racing a
/// rewrite of the structure can see them: the middle byte changed, or with
/// `invert` every byte (signatures included).
struct Torn {
bytes: Vec<u8>,
torn: AtomicU32,
invert: AtomicBool,
}
impl clawhdf5::Storage for Torn {
fn read_at(
&self,
offset: u64,
len: usize,
) -> Result<std::borrow::Cow<'_, [u8]>, clawhdf5_format::error::FormatError> {
let got = self.bytes.as_slice().read_at(offset, len)?;
let tear = self
.torn
.fetch_update(Ordering::SeqCst, Ordering::SeqCst, |n| n.checked_sub(1))
.is_ok();
if tear && !got.is_empty() {
let mut v = got.into_owned();
if self.invert.load(Ordering::SeqCst) {
v.iter_mut().for_each(|b| *b = !*b);
} else {
let mid = v.len() / 2;
v[mid] ^= 0x5a;
}
return Ok(v.into());
}
Ok(got)
}
fn len(&self) -> u64 {
self.bytes.len() as u64
}
}
#[test]
fn a_torn_metadata_read_is_retried_and_never_returned() {
use clawhdf5::Storage as _;
let torn = Arc::new(Torn {
bytes: std::fs::read(fixture("swmr_mid_write.h5")).unwrap(),
torn: AtomicU32::new(0),
invert: AtomicBool::new(false),
});
let f = File::open_storage_swmr(torn.clone()).unwrap();
assert_eq!(f.swmr_retries(), 0);
assert!(torn.len() > 0);
// The next 3 reads (object headers) come back garbled: the lookup
// fails its checks, is run again, and returns the right dataset. The
// same with every byte garbled, signatures included.
for invert in [false, true] {
torn.invert.store(invert, Ordering::SeqCst);
let before = f.swmr_retries();
torn.torn.store(3, Ordering::SeqCst);
let a = f.dataset("a").unwrap();
assert!(f.swmr_retries() > before, "invert {invert}");
assert_eq!(a.read_i64().unwrap(), (1..=148).collect::<Vec<i64>>());
// Two garbled reads: the prefix of the root group's header (whose
// garbled byte may go unused, as it is read again whole) and the
// whole header, which fails its checksum.
let before = f.swmr_retries();
torn.torn.store(2, Ordering::SeqCst);
let b = f.dataset("b").unwrap();
assert_eq!(b.read_f64().unwrap().len(), 148 * 4);
assert!(f.swmr_retries() > before, "invert {invert}");
}
torn.invert.store(false, Ordering::SeqCst);
// With one attempt the error is returned instead (both reads of the
// header garbled, as above).
let mut f = File::open_storage_swmr(torn.clone()).unwrap();
f.set_swmr_read_attempts(1);
torn.torn.store(2, Ordering::SeqCst);
assert!(f.dataset("a").is_err());
assert_eq!(f.swmr_retries(), 0);
torn.torn.store(0, Ordering::SeqCst);
// A refresh that keeps failing gives up after the attempts, and the
// handle keeps its extent.
f.set_swmr_read_attempts(5);
let mut b = f.dataset("b").unwrap();
torn.torn.store(1000, Ordering::SeqCst);
assert!(b.refresh().is_err());
assert_eq!(f.swmr_retries(), 4);
torn.torn.store(0, Ordering::SeqCst);
assert_eq!(b.shape().unwrap(), vec![148, 4]);
b.refresh().unwrap();
assert_eq!(b.read_f64().unwrap().len(), 148 * 4);
// A file not opened for SWMR reading does not retry.
let plain = File::open_storage(torn.clone()).unwrap();
torn.torn.store(2, Ordering::SeqCst);
assert!(plain.dataset("a").is_err());
assert_eq!(plain.swmr_retries(), 0);
}
// ---------------------------------------------------------------------------
// A live file: an h5py SWMR writer appends while clawhdf5 and h5py read.
// ---------------------------------------------------------------------------
/// The h5py SWMR writer. It creates three chunked datasets whose values are
/// a function of their position, so a reader can check every value it
/// reads without knowing when it was written:
///
/// - `a`: int64 `(n,)`, chunks of 100, no filter — `a[i] = i + 1`
/// (Extensible Array index);
/// - `b`: float64 `(n, 4)`, chunks of 16 x 4, gzip — `b[i, j] = 10 i + j + 1`
/// (Extensible Array index);
/// - `c`: int32 `(r, k)`, both unlimited, chunks of 8 x 8, no filter —
/// `c[i, j] = 1000 i + j + 1` (version-2 B-tree index).
///
/// It switches to SWMR mode, prints `ready`, then for `steps` steps appends a
/// random number of rows to each dataset (and every 7th step a column to
/// `c`), resizing before writing as SWMR writers must, flushing each
/// dataset and pausing 2 ms, and finally closes the file and writes the
/// final contents to `<path>.<name>.bin`. `CLAWHDF5_SWMR_STEPS` sets the
/// number of steps (default 2500).
const WRITER: &str = r#"
import sys, time, random, h5py, numpy as np
path, steps = sys.argv[1], int(sys.argv[2])
rng = random.Random(1234)
f = h5py.File(path, "w", libver="latest")
a = f.create_dataset("a", shape=(0,), maxshape=(None,), chunks=(100,), dtype="i8")
b = f.create_dataset("b", shape=(0, 4), maxshape=(None, 4), chunks=(16, 4), dtype="f8",
compression="gzip")
c = f.create_dataset("c", shape=(0, 1), maxshape=(None, None), chunks=(8, 8), dtype="i4")
f.swmr_mode = True
print("ready", flush=True)
na = nb = rc = 0
kc = 1
for step in range(steps):
k = rng.randint(1, 60)
a.resize((na + k,)); a[na:na + k] = np.arange(na, na + k) + 1; na += k
a.flush()
k = rng.randint(1, 20)
rows = np.arange(nb, nb + k)[:, None] * 10 + np.arange(4) + 1
b.resize((nb + k, 4)); b[nb:nb + k] = rows; nb += k
b.flush()
if step % 7 == 6:
c.resize((rc, kc + 1))
if rc:
c[:, kc] = np.arange(rc) * 1000 + kc + 1
kc += 1
k = rng.randint(1, 3)
c.resize((rc + k, kc))
c[rc:rc + k] = np.arange(rc, rc + k)[:, None] * 1000 + np.arange(kc) + 1
rc += k
c.flush()
time.sleep(0.002)
f.close()
with h5py.File(path, "r") as f:
for name in "abc":
open(f"{path}.{name}.bin", "wb").write(f[name][()].tobytes())
print(name, *f[name].shape, flush=True)
"#;
/// h5py's own SWMR reader, run beside ours as the reference: the same
/// checks, until `<path>.stop` exists. Prints its iteration count.
const H5PY_READER: &str = r#"
import os, sys, h5py, numpy as np
path = sys.argv[1]
f = h5py.File(path, "r", swmr=True)
ds = {n: f[n] for n in "abc"}
last = {n: (0,) * ds[n].ndim for n in "abc"}
its = 0
while True:
stop = os.path.exists(path + ".stop")
for n, d in ds.items():
d.refresh()
shape = d.shape
assert all(s >= l for s, l in zip(shape, last[n])), (n, shape, last[n])
last[n] = shape
v = d[()]
if n == "a":
want = np.arange(shape[0]) + 1
elif n == "b":
want = np.arange(shape[0])[:, None] * 10 + np.arange(4) + 1
else:
want = np.arange(shape[0])[:, None] * 1000 + np.arange(shape[1]) + 1
bad = np.argwhere(v != want)
assert len(bad) == 0, (n, shape, bad[:5], v[tuple(bad[0])], want[tuple(bad[0])])
its += 1
if stop:
break
print("iterations", its, *last["a"], *last["b"], *last["c"], flush=True)
"#;
fn want_a(n: u64) -> Vec<i64> {
(1..=n as i64).collect()
}
fn want_b(rows: std::ops::Range<u64>) -> Vec<f64> {
rows.flat_map(|i| (0..4).map(move |j| (10 * i + j + 1) as f64))
.collect()
}
fn want_c(rows: std::ops::Range<u64>, cols: u64) -> Vec<i32> {
rows.flat_map(|i| (0..cols).map(move |j| (1000 * i + j + 1) as i32))
.collect()
}
/// One pass of the Rust reader: refresh every dataset, check its extent
/// did not shrink, and check what it reads (the last rows of each, and all
/// of them every `full_every` passes).
fn check_pass(
datasets: &mut [clawhdf5::Dataset<'_>; 3],
last: &mut [Vec<u64>; 3],
pass: u64,
full_every: u64,
) {
for (k, d) in datasets.iter_mut().enumerate() {
d.refresh().unwrap_or_else(|e| panic!("refresh {k}: {e}"));
let shape = d.shape().unwrap();
assert!(
shape.iter().zip(&last[k]).all(|(s, l)| s >= l),
"dataset {k} shrank: {shape:?} after {:?}",
last[k]
);
last[k] = shape;
}
let full = pass.is_multiple_of(full_every);
let na = last[0][0];
let from = if full { 0 } else { na.saturating_sub(200) };
let got = if full {
datasets[0].read_i64()
} else {
datasets[0].read_i64_selection(&Selection::slice(std::slice::from_ref(&(from..na))))
}
.unwrap_or_else(|e| panic!("read a: {e}"));
let want: Vec<i64> = want_a(na).split_off(from as usize);
assert!(got == want, "a {from}..{na}: {:?}", first_diff(&got, &want));
let nb = last[1][0];
let from = if full { 0 } else { nb.saturating_sub(50) };
let sel = Selection::slice(&[from..nb, 0..4]);
let got = datasets[1]
.read_f64_selection(&sel)
.unwrap_or_else(|e| panic!("read b: {e}"));
let want = want_b(from..nb);
assert!(got == want, "b {from}..{nb}: {:?}", first_diff(&got, &want));
let (rc, kc) = (last[2][0], last[2][1]);
let from = if full { 0 } else { rc.saturating_sub(9) };
let got = if full {
datasets[2].read_i32()
} else {
datasets[2].read_i32_selection(&Selection::slice(&[from..rc, 0..kc]))
}
.unwrap_or_else(|e| panic!("read c: {e}"));
let want = want_c(from..rc, kc);
assert!(
got == want,
"c {from}..{rc} x {kc}: {:?}",
first_diff(&got, &want)
);
}
fn first_diff<T: PartialEq + std::fmt::Debug>(got: &[T], want: &[T]) -> String {
if got.len() != want.len() {
return format!("{} values, want {}", got.len(), want.len());
}
let i = got.iter().zip(want).position(|(g, w)| g != w).unwrap();
format!("first difference at {i}: {:?}, want {:?}", got[i], want[i])
}
#[test]
fn a_live_file_reads_consistently_while_an_h5py_swmr_writer_appends() {
skip_if_no_python!();
let steps: u32 = std::env::var("CLAWHDF5_SWMR_STEPS")
.ok()
.and_then(|s| s.parse().ok())
.unwrap_or(2500);
let dir = tempfile::tempdir_in(env!("CARGO_TARGET_TMPDIR")).unwrap();
let path = dir.path().join("live.h5");
let mut writer = Command::new(python())
.args(["-c", WRITER, path.to_str().unwrap(), &steps.to_string()])
.stdout(Stdio::piped())
.spawn()
.unwrap();
let mut out = BufReader::new(writer.stdout.take().unwrap());
let mut line = String::new();
out.read_line(&mut line).unwrap();
assert_eq!(line.trim(), "ready", "writer did not start");
// h5py's SWMR reader on the same file, as the reference.
let h5py_reader = Command::new(python())
.args(["-c", H5PY_READER, path.to_str().unwrap()])
.stdout(Stdio::piped())
.stderr(Stdio::piped())
.spawn()
.unwrap();
let file = File::open_swmr(&path).unwrap();
assert!(file.is_swmr_read());
assert!(file.swmr_writer_active().unwrap());
// Two reader threads share the file, each with its own handles; each
// makes passes until it has seen the writer exit, then one more.
let writer_done = AtomicBool::new(false);
let reader = || {
let mut datasets = ["a", "b", "c"].map(|n| file.dataset(n).unwrap());
let mut last = [vec![0], vec![0, 4], vec![0, 1]];
let (mut passes, mut grew) = (0u64, 0u64);
loop {
let done = writer_done.load(Ordering::Acquire);
let before = last[0][0];
check_pass(&mut datasets, &mut last, passes, 50);
passes += 1;
grew += u64::from(last[0][0] > before);
if done {
break;
}
}
// The writer has closed the file: one more pass reads the final
// extents, all of them.
check_pass(&mut datasets, &mut last, 0, 50);
(passes, grew, last)
};
let (status, results) = std::thread::scope(|s| {
let readers = [s.spawn(reader), s.spawn(reader)];
let status = writer.wait().unwrap();
writer_done.store(true, Ordering::Release);
(status, readers.map(|r| r.join().unwrap()))
});
assert!(status.success(), "writer failed");
assert!(!file.swmr_writer_active().unwrap());
let last = results[0].2.clone();
assert_eq!(results[1].2, last);
let passes: Vec<u64> = results.iter().map(|r| r.0).collect();
let grew: u64 = results.iter().map(|r| r.1).min().unwrap();
std::fs::write(path.with_extension("h5.stop"), b"").unwrap();
let reference = h5py_reader.wait_with_output().unwrap();
assert!(
reference.status.success(),
"h5py's SWMR reader failed:\n{}",
String::from_utf8_lossy(&reference.stderr)
);
let reference = String::from_utf8_lossy(&reference.stdout).into_owned();
// The writer's final contents, as h5py reads the closed file.
let mut finals = String::new();
out.read_to_string(&mut finals).unwrap();
let shapes: std::collections::HashMap<&str, Vec<u64>> = finals
.lines()
.map(|l| {
let mut w = l.split_whitespace();
let name = w.next().unwrap();
(name, w.map(|v| v.parse().unwrap()).collect())
})
.collect();
assert_eq!(shapes["a"], last[0]);
assert_eq!(shapes["b"], last[1]);
assert_eq!(shapes["c"], last[2]);
let bin = |name: &str| std::fs::read(format!("{}.{name}.bin", path.display())).unwrap();
let le = |b: &[u8], n: usize| -> Vec<[u8; 8]> {
b.chunks(n)
.map(|c| {
let mut v = [0u8; 8];
v[..n].copy_from_slice(c);
v
})
.collect()
};
// The live handle and a fresh non-SWMR open of the closed file both
// read exactly what h5py reads.
for f in [&file, &File::open(&path).unwrap()] {
let a: Vec<[u8; 8]> = f
.dataset("a")
.unwrap()
.read_i64()
.unwrap()
.iter()
.map(|v| v.to_le_bytes())
.collect();
assert_eq!(a, le(&bin("a"), 8));
let b: Vec<[u8; 8]> = f
.dataset("b")
.unwrap()
.read_f64()
.unwrap()
.iter()
.map(|v| v.to_le_bytes())
.collect();
assert_eq!(b, le(&bin("b"), 8));
let c: Vec<[u8; 8]> = f
.dataset("c")
.unwrap()
.read_i32()
.unwrap()
.iter()
.map(|v| {
let mut x = [0u8; 8];
x[..4].copy_from_slice(&v.to_le_bytes());
x
})
.collect();
assert_eq!(c, le(&bin("c"), 4));
}
eprintln!(
"SWMR: {steps} writer steps; clawhdf5 readers made {passes:?} passes (each saw `a` \
grow at least {grew} times), {} retries; h5py reader: {}",
file.swmr_retries(),
reference.trim()
);
assert!(
grew >= 2,
"a reader never saw the file grow ({passes:?} passes)"
);
}
// ---------------------------------------------------------------------------
// Every read path retries: strings, variable-length data, attributes.
// ---------------------------------------------------------------------------
/// A storage over `bytes` whose read number `fail_at` (counted from the
/// last [`Flaky::arm`]) fails with a checksum mismatch, as a read racing a
/// SWMR writer's flush can; it counts the reads.
struct Flaky {
bytes: Vec<u8>,
reads: AtomicU32,
fail_at: AtomicU32,
}
impl Flaky {
fn arm(&self, fail_at: u32) {
self.reads.store(0, Ordering::SeqCst);
self.fail_at.store(fail_at, Ordering::SeqCst);
}
}
impl clawhdf5::Storage for Flaky {
fn read_at(
&self,
offset: u64,
len: usize,
) -> Result<std::borrow::Cow<'_, [u8]>, clawhdf5_format::error::FormatError> {
let n = self.reads.fetch_add(1, Ordering::SeqCst);
if n == self.fail_at.load(Ordering::SeqCst) {
return Err(clawhdf5_format::error::FormatError::ChecksumMismatch {
expected: 0,
computed: 1,
});
}
self.bytes.as_slice().read_at(offset, len)
}
fn len(&self) -> u64 {
self.bytes.len() as u64
}
}
/// `swmr_strings_attrs.h5`: a copy h5py 3.16 (HDF5 2.0) made of its own
/// file right after switching it to SWMR mode (the SWMR-write flag is set),
/// holding a variable-length string dataset `s` (`alpha beta gamma
/// delta`), a variable-length int32 dataset `v` (`[1] [2 3] [4 5 6]`), an
/// int64 dataset `d` (`0..10`) with 12 string attributes `k00`..`k11`
/// (`value 0`..; dense storage: a fractal heap and a v2 B-tree), a group
/// `g` with attribute `note`, and root attributes `title` (a
/// variable-length string) and `n`.
#[test]
fn every_read_path_of_a_live_file_retries_a_failed_read() {
let flaky = Arc::new(Flaky {
bytes: std::fs::read(fixture("swmr_strings_attrs.h5")).unwrap(),
reads: AtomicU32::new(0),
fail_at: AtomicU32::new(u32::MAX),
});
let f = File::open_storage_swmr(flaky.clone()).unwrap();
assert!(f.is_swmr_read());
let root = f.root();
let g = f.group("g").unwrap();
let s = f.dataset("s").unwrap();
let v = f.dataset("v").unwrap();
let d = f.dataset("d").unwrap();
let sorted = |m: std::collections::HashMap<String, clawhdf5::AttrValue>| {
format!(
"{:?}",
m.into_iter().collect::<std::collections::BTreeMap<_, _>>()
)
};
type Op<'a> = Box<dyn Fn() -> Result<String, clawhdf5::Error> + 'a>;
let ops: Vec<(&str, Op<'_>)> = vec![
("root.attrs", Box::new(|| root.attrs().map(sorted))),
(
"root.attr",
Box::new(|| root.attr("title").map(|a| format!("{a:?}"))),
),
("g.attrs", Box::new(|| g.attrs().map(sorted))),
("d.attrs", Box::new(|| d.attrs().map(sorted))),
(
"d.attr",
Box::new(|| d.attr("k05").map(|a| format!("{a:?}"))),
),
(
"root.datasets",
Box::new(|| root.datasets().map(|n| format!("{n:?}"))),
),
(
"root.groups",
Box::new(|| root.groups().map(|n| format!("{n:?}"))),
),
("d.shape", Box::new(|| d.shape().map(|n| format!("{n:?}")))),
(
"d.read_i64",
Box::new(|| d.read_i64().map(|n| format!("{n:?}"))),
),
(
"s.read_string",
Box::new(|| s.read_string().map(|n| format!("{n:?}"))),
),
(
"s.read_string_bytes",
Box::new(|| s.read_string_bytes().map(|n| format!("{n:?}"))),
),
(
"s.read_string_selection",
Box::new(|| {
s.read_string_selection(&Selection::slice(std::slice::from_ref(&(1..3))))
.map(|n| format!("{n:?}"))
}),
),
(
"v.read_vlen",
Box::new(|| v.read_vlen::<i32>().map(|n| format!("{n:?}"))),
),
(
"v.read_vlen_selection",
Box::new(|| {
v.read_vlen_selection::<i32>(&Selection::slice(std::slice::from_ref(&(1..3))))
.map(|n| format!("{n:?}"))
}),
),
];
for (name, op) in &ops {
flaky.arm(u32::MAX);
let want = op().unwrap_or_else(|e| panic!("{name}: {e}"));
let reads = flaky.reads.load(Ordering::SeqCst);
// Each read the operation makes fails once in turn: the operation
// is run again and returns the same result.
for k in 0..reads {
flaky.arm(k);
let before = f.swmr_retries();
let got = op().unwrap_or_else(|e| panic!("{name}, read {k} of {reads} failing: {e}"));
assert_eq!(got, want, "{name}, read {k} of {reads} failing");
assert_eq!(f.swmr_retries(), before + 1, "{name}, read {k} of {reads}");
}
}
flaky.arm(u32::MAX);
assert!(ops[0].1().unwrap().contains("live strings"));
assert!(ops[3].1().unwrap().contains("value 11"));
assert_eq!(
ops[9].1().unwrap(),
r#"["alpha", "beta", "gamma", "delta"]"#
);
assert_eq!(ops[12].1().unwrap(), "[[1], [2, 3], [4, 5, 6]]");
// With one attempt the failure reaches the caller.
let mut f = File::open_storage_swmr(flaky.clone()).unwrap();
f.set_swmr_read_attempts(1);
let s = f.dataset("s").unwrap();
flaky.arm(0);
assert!(s.read_string().is_err());
}
+106 -5
View File
@@ -7,20 +7,33 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
`File::open_storage` gives the facade's read API over any `Storage` (see
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is
object stores) and URLs in `h5rs` (see the M3 status below). M5 (SWMR) is
done on branch `feat/p3-m5-swmr-reader`, with its own design in
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
next. Every count in §1–§2 was
object stores) and URLs in `h5rs` (see the M3 status below); the Python
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet.
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
reader, through the restartable `NeedBytes` mode (see the M4 status
below). M5 (SWMR) is not started.
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
converted code without changing the plan: object-header continuation chunks
are followed without recursion (still one bounded `read_at` per chunk), and
implicit chunk indexes are addressed over the maximum chunk grid (in
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
working on the whole file in memory; it is not part of this design. Every count below was
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
next to it. No timing numbers appear here on purpose: the machine was shared
working on the whole file in memory; it is not part of this design. Every
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
every count in a milestone's status on the date it gives, with the
commands given next to it. No timing numbers appear here on purpose: the machine was shared
with other build jobs when this was written.
## The problem
@@ -475,7 +488,7 @@ fast path within benchmark noise.
a request counter exposed for tests and users.
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
Python bindings, with these choices:
Python bindings (done 2026-09-27, below), with these choices:
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
sits below the facade.
@@ -512,6 +525,17 @@ fast path within benchmark noise.
same work without a cache: 141 936 requests. Per file: A lists in 2
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
6.4 MB: 35 001 object headers spread over the file).
- *Status 2026-09-27, Python bindings:* done on branch
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
`File.open_url(url, **options)` (cache and HTTP options) go through
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
parsing (path lookups, headers, attributes, listings, the global heap)
moved from `File::as_bytes` to `File::storage()` and the `*_in`
functions, and every file access, metadata included, runs with the GIL
released. Checked by running the whole read-vs-h5py suite over an
in-process range server (1 MiB and 1 KiB blocks) and by request counts
in `crates/clawhdf5-py/tests/test_remote.py`.
**M4 — wasm lazy loading (1–2 weeks).**
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
@@ -519,10 +543,87 @@ fast path within benchmark noise.
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
to a whole download when the server does not answer 206.
- `examples/wasm-viewer`: open by URL.
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
with these choices:
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
`Storage` over the blocks fetched so far. A call (open, list, read)
runs as a pass; a read that misses records its blocks and fails with
a storage error. The error itself is not the signal: parsers catch
errors and carry on (a listing leaves out a link it cannot resolve),
so *any* pass that recorded a miss is thrown away, whatever it
returned, and re-run once the missing ranges have been fetched
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
synchronous XHR would have kept the parsers' single pass, but it
needs the page to run the reader off the main thread (a second
module, messages for every call and every typed array copied back),
since synchronous XHR can return binary data only there; the
restartable loop costs a
re-parse per wave of misses instead, which is CPU, not network.
- **Progress:** no block is evicted while a call is in flight
(`LazyStorage::operation`), so each pass that does not finish asks for
at least one new block, and a call ends after at most one pass per
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
evicted before metadata. The price: a call holds everything it reads
until it finishes.
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
backend by blocking, evicts during a read, and does not keep reads
that miss more than half its budget; a restartable pass needs every
block it has read to be there when it is re-run. The block
arithmetic and coalescing (runs of consecutive blocks, a one-block
hole filled, at most 8 MiB per request) follow it; the cache is
about 450 lines (with its documentation) in the wasm crate, with no
in-flight tracking or HTTP. (A hole already cached is not filled:
it would be fetched again.)
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
requests at a time, every answer checked (206, `Content-Range` when
visible, body length, the ETag/Last-Modified and length of the
first answer). A `200` to the first request is kept as the whole
file up to `maxDownload` (512 MiB), or refused with
`fallback: "error"`. The first block comes with the request that
learns the length, as in M3.
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
--test lazy` compares what the viewer shows of each file (kinds,
listings, attributes, info, whole reads, hyperslabs) read lazily at
512 B to 1 MiB blocks with the in-memory reader and the facade's
range-storage path, for built files, the h5py/netCDF4 fixture and,
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
files. The built package, under Node and headless Chromium
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
`test/serve.py`, which counts requests): every fixture check over
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
attributes and a 10-value window of the 25-million-value dataset
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
with `open(bytes)`.
- After review (2026-09-27): a call may fetch at most `maxFetch`
(512 MiB, at most 1 GiB) and a longer single read is refused before
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
more is refused at open on wasm32 (offsets become `usize`); every
response body is cut off at the length asked for. Parsers that walk
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
the wasm listing's child headers) read the siblings after the first
failure before returning it, so a pass asks for a whole level's
missing blocks: listing 3000 datasets went from 185 passes to 6.
The 32-bit risk below is covered by a Node test that reads data at
3 GiB from a mock server and is refused a 4 GiB file.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
blocks past the old end. Needs libhdf5 SWMR semantics research first.
- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader`; design and
libhdf5 research in [`swmr.md`](swmr.md). Differences from the sketch
above: the refresh is per dataset (`Dataset::refresh`, as libhdf5's
`H5Drefresh`), not per file — a SWMR writer only grows datasets, and the
superblock's EOF is not kept up to date by it, so there is nothing to
re-read there (`File::swmr_writer_active` re-reads its flags). A live
file (`File::open_swmr`) reads through `FileStorage` (positioned reads,
`len()` the current length) with no block or chunk cache, rather than
invalidating cached blocks: `BlockCache`/`HttpStorage` stay snapshot
readers. Operations that fail with an error a racing write can cause are
retried up to 100 times (libhdf5's default for SWMR readers).
Tested against a live h5py writer and h5py's SWMR reader
(`crates/clawhdf5/tests/swmr_interop.rs`).
Total: roughly 6–10 engineer-weeks for M0–M4 (estimate, not measured).
+178
View File
@@ -0,0 +1,178 @@
# Design: reading files a SWMR writer is still appending to (range-read M5)
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
(see "Status" at the end). This is milestone M5 of
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
## What libhdf5 does
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
opened a file with `libver='latest'` and switched to SWMR mode (h5py
`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file
with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the
writer keeps appending. What the format and the library guarantee:
- **Superblock v3, flags set.** The writer sets the superblock's
file-consistency flags to write access + SWMR write (`0x05`) and clears
them on close. The superblock's end-of-file address is *not* kept up to
date while writing: a copy of a file taken mid-write records an EOF of a
few hundred bytes while the file is tens of kilobytes (checked on tank,
2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR
reader therefore skips libhdf5's end-of-allocation check for every read
(`H5FD_read`: "allow access to data past the end of the allocated space
… for SWMR read access"), and bounds reads by the file's real length.
- **A non-SWMR open of such a file fails** in libhdf5: "file is already
open for write (may use <h5clear file> to clear file consistency flags)".
- **The writer only appends.** New objects and attributes cannot be created
in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked
datasets with one unlimited dimension use an Extensible Array index, with
more than one a version-2 B-tree; both are updated in a SWMR-safe way.
Fixed Array and single-chunk indexes are for datasets that cannot grow.
- **Flush ordering.** Every metadata structure the writer uses is
checksummed, and flush dependencies order the writes: a chunk's data is
written before the index entry that points to it, and index blocks before
the object header whose dataspace announces the new extent. A reader
that reads the object header first and the index after sees an index at
least as new as the extent, so every chunk inside the extent it read is
either in the index or never written (then it reads as the fill value, as
it does for libhdf5's reader).
- **Refresh.** A reader sees a dataset's new extent only when it refreshes
it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's
cached metadata and reads the object header again.
- **Retries.** Reads are not atomic against writes on every system, so a
checksum can fail when a structure is read while the writer rewrites it.
A SWMR reader reads checksummed metadata up to 100 times before failing
(`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1
otherwise — `H5Ppublic.h` of HDF5 1.14.6).
## What clawhdf5 did before
- `File::open` of a SWMR-flagged file bounded every read by the recorded
EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the
file). A file copied or read mid-write therefore listed, but every
chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and
`h5rs check` reported chunk indexes "past the end of the file".
- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a
per-file chunk cache that keeps each dataset's chunk index and decoded
chunks for the life of the `File`. A reader could not see growth at all,
and a mapping of a file that is being rewritten can change under a read.
## Design
1. **Bound SWMR files by their length.** `Superblock::data_end` returns the
file's length for a version-3 superblock with the SWMR-write flag,
whatever EOF it records, as libhdf5's SWMR reader does. Every existing
open path (`File::open`, `open_storage`, `h5rs`) then reads a finished
copy of a live file. We keep opening such files without a SWMR flag
(libhdf5 refuses): it is read-only and the alternative is an error.
2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.**
- The file is read with positioned reads (`FileStorage`, `pread` on
Unix, `seek_read` on Windows), never mapped, and `len()` is the file's
current length, so reads past the length seen at open work.
- The facade's view of the file (`FileData`) is *live*: reads are not
clamped to an end fixed at open, only by the storage's current length.
- The chunk cache is not used: every read reads the chunk index and the
chunks it needs again. A cached index would hide new chunks, and a
cached partial edge chunk would read as fill where the writer has since
written data (unfiltered edge chunks are rewritten in place).
- Only a file whose superblock has the SWMR-write flag when it is opened
is read live. Any other file is read exactly as `File::open` reads it
(bounded by its recorded EOF, through the chunk cache, no retries), and
`is_swmr_read()` is `false`. libhdf5's SWMR reader is looser: it skips
the end-of-allocation check in `H5FD_read` for every file it opens,
flagged or not, yet still refuses an object header past the EOF
(`H5O_protect`, "address of object past end of allocation"). On a
closed file whose EOF is below its length (checked 2026-09-27, h5py
3.16 / HDF5 2.0, the mid-write fixture with its flags cleared) it
therefore reads a dataset whose chunk index lies past the EOF, which
its plain reader refuses. We follow the plain reader there: the EOF of
a file no SWMR writer has open is what the file says it is.
- A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would
serve stale bytes; `HttpStorage` also pins a file by ETag and length and
refuses a changed file. Remote SWMR is out of scope.
3. **`Dataset::refresh()`** reads the dataset's object header again (same
address) and replaces the handle's copy, so `shape()` and every later
read use the new extent. Like h5py, a handle that is not refreshed keeps
its extent; its reads still read the index as it is now, and only return
elements inside that extent.
4. **Bounded retries.** In a live file, an operation (open, lookups and
listings, refresh, every dataset read — typed, raw, selections, strings
and variable-length data with their global-heap decoding — attributes,
`File::decode_*`, `verify_provenance`) that fails with an error a
concurrent write can
cause is run again from the start, up to `File::swmr_read_attempts()`
times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
in all). libhdf5 retries the one structure whose checksum failed; we
retry the whole operation, because the parsers are pure functions of the
bytes they read. Which errors: those libhdf5 retries (`H5C__load_entry`)
— a checksum mismatch, and a failure to decode the prefix it reads
before the checksum to size the structure (for an object header its
signature and version: a test that garbles every byte of a header gets
`InvalidObjectHeaderVersion`) — and a read past the file's current end
(short here; libhdf5 reads zeros there, which fail the checksum). Every
other error (an unsupported version or message, a file that is not HDF5,
a structure corrupt behind a valid checksum) is returned at once: an
earlier version retried nearly every format error, so a permanent one
cost about 0.9 s of pauses. Non-live files never retry.
Results are only returned from a run where every structure verified, so
a torn metadata read is an error, never data. `File::swmr_retries()`
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
Attribute reads leave out an attribute they cannot read (or return a
variable-length string one as `AttrValue::Raw`) instead of failing;
on a live file such an error of a retried kind runs the read again too,
and after the last attempt the last result is returned as before. The
zero-copy reads (`read_raw_ref`, `read_as_slice`, `read_*_zerocopy`)
need the file in memory, which a live file never is: they report
`None` / `ContiguousStorageRequired` without reading data.
Global heap collections (variable-length data) have no checksum in
HDF5, like raw data, so a torn read of one is only caught when it fails
a check (its signature, a bound).
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
as here: correctness rests on the writer's ordering (chunk data before
the index entry, only appends), as for libhdf5's reader.
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
flags again, so a reader can tell when the writer has closed the file.
A writer that crashed or was killed never clears the flag, so a reader
that follows a file needs a second stop condition (the README example
stops after a minute without growth).
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
writer cannot add them), and `MmapFile`/`LazyFile`.
## Tests
- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose
EOF is below the file length.
- `crates/clawhdf5/tests/swmr_interop.rs`:
- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
- the same bytes with the flags cleared read as `File::open` reads them
(and how h5py's two readers read them);
- a permanent error returns at once; every read path (listings,
attributes, strings, variable-length data) with each of its reads
failing once in turn returns the same result;
- a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a
2-D dataset with one unlimited dimension (Extensible Array, one of them
gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step,
while a Rust reader refreshes and reads them in a loop, and an h5py SWMR
reader does the same as the reference. Every value read must be the value
the writer wrote (a deterministic function of its position), extents
never shrink, and after the writer closes both readers must read the
same data as h5py.
## Status
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
build): no read returned a value the writer had not written at that
position, h5py's reader agreed, and retries were needed but rare (the
failures seen were checksum mismatches, each cured by one retry). A
variant of the test with the chunk cache left on in live mode fails it
(stale chunk index / edge chunk), which is why live files do not use it.
Also found: `File::open` of such a file had been failing since the
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).
+183 -10
View File
@@ -7,6 +7,56 @@ deleting it.
---
## Files a SWMR writer had open could not be read past a stale end of file
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any
release: reads have been bounded by the recorded end of file since
`7d7a7e7` (2026-09-26), which no release contains.
A libhdf5 writer in SWMR mode (h5py `f.swmr_mode = True`) sets the
superblock's SWMR-write flag and does not keep its end-of-file address up to
date: a copy h5py made of its own file mid-write records 715 in a
6 030-byte file (h5py 3.16 / HDF5 2.0, tank). `Superblock::data_end` only
ignored the recorded end when it lay *past* the end of the file, so every
reader bounded such a file at 715 bytes: it listed, but every chunked read
failed ("unexpected EOF: need 787 bytes, have 715") and `h5rs check`
reported the chunk indexes past the end of the file. Never wrong data.
**Fix:** for a v3 superblock with the SWMR-write flag, the data ends at the
end of the file, as libhdf5's SWMR reader reads it. **Test:**
`crates/clawhdf5/tests/swmr_interop.rs` (the mid-write copy,
`tests/fixtures/swmr_mid_write.h5`, through every open path and against
h5py's SWMR reader). A file still being written is read with
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
at its length at open and is not meant for files that change while open.
## Shrinking a chunked dataset with no recorded maximum scrambled it
**Status:** fixed 2026-09-27, before any release (`FileEditor::resize`
shipped on main in PR #18, a4c2ace). Files the writer produced before the
fix still lack the maximum; see *Existing files*.
clawhdf5's writer stored no maximum dimensions for a chunked dataset
created without a `maxshape` (libhdf5 always stores one, equal to the
dimensions when none is given). With none recorded, the maximum is the
current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index
places chunks by the maximum. `FileEditor::resize` changed only the
current dimensions, so a shrink moved every existing chunk and every
reader returned wrong values; after a shrink the dataset could not grow
back. Found by the review of the Python editing work.
**Fix:** before the first resize of such a dataset the editor records the
maximum libhdf5 would have written (the dimensions the index was built
with); the writer now records it for every chunked dataset. **Test:**
`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:**
datasets written before the fix have no recorded maximum. The fixed
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
does not** — it scrambles them the same way, and lets them grow past
their Fixed Array. Resize them once with a fixed `FileEditor` (a resize
to the same shape changes nothing; shrink and grow back) before letting
libhdf5 resize them. A dataset already shrunk by the unfixed
editor holds misplaced chunks; rewrite it from a good copy.
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to
@@ -116,6 +166,57 @@ space it leaves is too small for its next, larger version.
**No journal.** A crash while an edit patches existing structures can leave
the file inconsistent; see the `FileEditor` documentation.
**Renamed files outside Linux.** Edits always go to the file the editor
opened. `FileEditor::reader()` (which the Python `'r+'` handle reads
through) reopens it through `/proc/self/fd` on Linux; elsewhere it reopens
the path, and after the path was renamed or replaced it fails (on Unix;
Windows cannot tell and would read whatever the path names).
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
**Status:** open (added 2026-09-27). The Python bindings edit through
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
before anything is written. On top of them:
- **No new or deleted objects:** `create_dataset`/`create_group` in an
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
`NotImplementedError` (the editor changes values, shapes and
attributes only). Mode `'a'` works on an existing file only.
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
whole elements of the same structured dtype are), HDF5 array-type
elements, variable-length data, strings padded with spaces or
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
ones, h5py's, are writable), compounds containing such strings, null
dataspaces, index-list writes of more than 2²² elements (write them
in slices), and boolean-mask keys (`ds[mask] = v`, and mask reads;
h5py supports both).
- **`str` attributes are fixed-length UTF-8**, where h5py writes
variable-length strings: h5py reads them back as `bytes`
(`numpy.bytes_`), not `str`.
- **Numeric conversion follows libhdf5's native-order results, not its
bugs.** Arrays are converted as libhdf5 converts them (integers
saturate, floats are truncated toward zero and clipped), checked value by
value against h5py 3.16 / HDF5 2.0 in
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
differs from h5py on purpose:
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
minimum or 2⁶³ depending on the type);
- when the dataset or the array is not in native byte order, libhdf5's
"soft" conversions store a float in (-1, 0) as the integer minimum and
wrap an unsigned value too large for the signed type of the same size
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
native order;
- libhdf5's native casts that are undefined in C: half floats into
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
integers (→ minimum), a float equal to the integer maximum rounded up
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
- **Each edit reopens the file** (a new memory map) so that reads see it;
reads from other threads wait while an edit is written.
## Selection reads that decode more than the selection
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
@@ -844,9 +945,14 @@ cache, but:
`read_*_zerocopy`) need the file in memory and answer
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
panics for such a file (`File::contiguous_bytes()` is the fallible form).
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
`remote` feature).
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
(`h5rs` and the Python bindings read through `File::storage`, and take
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
feature; the wasm reader's `openUrl` reads by range requests since
2026-09-27, its `open(bytes)` takes a whole file).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -855,16 +961,25 @@ cache, but:
and so `Dataset::read_*`) lists a damaged dataset's chunks in hash-map
order, so which failing chunk it reports can differ from one `File` to
the next (`cve-2025-2310.h5`); the values of a dataset that reads are
not affected.
not affected. **Fixed 2026-09-27:** the chunk cache keeps the chunks in
the order the index lists them, as the uncached readers do
(`several_damaged_chunks_report_the_same_chunk_every_time`).
## Remote files (`clawhdf5-remote`) limits
**Status:** open (added 2026-09-26, milestone M3 of
`docs/design/range-reads.md`).
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
parses through `File::as_bytes`, which a remote file does not have; the
wasm reader's `openUrl` is milestone M4.
- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
the default wheel reads plain `http://` only: `https://` needs a wheel
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
`File::as_bytes`, which a remote file does not have. (The browser can
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.
@@ -903,10 +1018,68 @@ cache, but:
## `clawhdf5-wasm` (browser) limits
**Status:** open (by design for now; added 2026-09-26).
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
- The whole file is held in memory: `open()` takes its bytes. There are no
HTTP range reads, so a multi-GB file does not fit a browser tab.
- `open()` holds the whole file in memory (it takes its bytes), so a
multi-GB local file does not fit a browser tab. A file on a web server
can be opened with `openUrl()` instead, which fetches only the byte
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
with these limits:
- **Round trips:** a call runs as passes over the blocks fetched so far
and is re-run after each wave of misses, so a call costs one round
trip per wave, not one for everything: a chunk index is walked a
level (or a node) per round trip, while the chunks of a read are
fetched together. Listing a group asks for every child's object
header, and every node of a level of the group's index, in one pass
(since 2026-09-27; it was one round trip per header block): 3000
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
`libver="latest"` file (dense links). Each pass re-parses what the
call reads (CPU, not network). With headers spread through the file
(h5py writes each next to its data) a listing still fetches most of
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
- **Memory:** a call keeps every block it reads until it finishes (the
cache budget applies between calls). It may fetch at most `maxFetch`
bytes (512 MiB by default, at most 1 GiB), and a single read longer
than that fails before anything is fetched: a hostile server cannot
make the page fetch or allocate what a file's lengths claim. `read()`
of a dataset that would take more than 1 GiB while it is decoded
(stored bytes, the values widened to 64 bits, the result) fails,
naming `readHyperslab`; read windows of large datasets. On wasm32 a
buffer past 2 GiB cannot exist and a failed allocation aborts the
whole module (every open file on the page), which these limits keep
from happening; before 2026-09-27 both did abort it.
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
open. The format code turns file offsets into `usize` to use them
(with a clean error past it), which is 32 bits on wasm32, so nothing
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
tested (with a mock server); files above 200 MB have not been served
for real.
- **Cross-origin servers** must allow CORS for the page's origin and
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
answer `HEAD` with `Content-Length`. The file is pinned at open by its
`ETag` or `Last-Modified` (and its length): when the page cannot see
either header, only a change of length is detected.
- **A server without range support** (it answers `200`) costs a whole
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
with `fallback: "error"`. Every body, this one and each `206`, is read
as it arrives and cut off past its limit (the range asked for, or
`maxDownload`): a server cannot make the page buffer more.
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
size is not used. No retries: a failed request fails the call (calling
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included (a page on 127.0.0.1 reading a
file from localhost, with and without exposed headers); not in Firefox
or Safari.
- The native corpus comparison (`tests/lazy.rs` with
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
and which chunk's error is reported depends on the iteration order of
the chunk index (a `HashMap`, seeded per process), so the lazy and
the range-storage reads can name different errors. Both are errors;
not specific to `openUrl` (it predates it). **Fixed 2026-09-27:** the
chunk cache keeps the index's chunk order, so every read path names the
same (first) damaged chunk.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`.
+100 -20
View File
@@ -2,10 +2,12 @@
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
file, browse its groups, and look at a dataset's type, shape, attributes
and values (a 50 x 12 window at a time, read as a hyperslab, with the
leading dimensions of a 3-D+ dataset held at chosen indices). The file never
leaves the page.
file or give a URL, browse its groups, and look at a dataset's type, shape,
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
local file never leaves the page. A file given by URL is not downloaded:
only the byte ranges each view needs are fetched (HTTP range requests),
and the header shows the requests and bytes that has cost so far.
## Build and open
@@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
```
Then open <http://localhost:8000/>. `?file=<url>&path=<object>` opens a
file from a URL (same origin, or one serving CORS headers) and selects an
object in it, e.g. `?file=data/run1.h5&path=/results/energy`.
Then open <http://localhost:8000/>. The URL box, or
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
sending CORS headers, see Limits) and selects an object in it, e.g.
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
answer range requests, so a file served by it is downloaded whole (the
header says so); `test/serve.py` does:
```bash
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
```
## JavaScript API
```js
import init, { open } from "./pkg/clawhdf5_wasm.js";
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
await init();
const f = open(new Uint8Array(await blob.arrayBuffer()));
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
@@ -32,8 +42,41 @@ f.attrs("/grid"); // [{ name, value, dtype }]
f.read("/grid"); // { shape, dtype, data }
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
f.free();
// A file on a web server, read by HTTP range requests as needed. The same
// methods, each returning a promise.
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
await r.list("/");
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
r.free();
```
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
between calls, default 64 MiB), `maxFetch` (bytes one call may fetch, and
so the longest single read, up to 1 GiB, default 512 MiB), `fallback`
(`"download"`, the default, reads the whole file when the server ignores
`Range`, up to `maxDownload` bytes, default 512 MiB, at most 1 GiB;
`"error"` refuses such a server), `headers` (a `Headers`, `[name, value]`
pairs or an object) and `credentials` (passed to every request),
`parallel` (requests in flight, 1 to 1024, default 6; when one fails the
others are aborted), `fetch` (a `fetch`-compatible function to use).
How it works: the reader is synchronous and a page cannot block on the
network, so each call runs as a *pass* over the blocks fetched so far. A
pass that needs a block not yet fetched is abandoned, the missing blocks
are fetched (in parallel, adjacent blocks in one request), and the pass is
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
costs one request (the first block, which also gives the file's size);
listing a group whose metadata is in blocks already fetched costs none,
and otherwise a round trip per level of the group's index plus one for
its children's headers, all fetched together;
reading a chunked dataset costs a round trip for its chunk index (a few
for a deep one) and one batch of requests for its chunks. Every answer is
checked — a `206` with exactly the bytes asked for, from the same file
(ETag or Last-Modified, and length) — or the call fails.
`data` is the typed array of the stored width (`Float64Array`,
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
`BigUint64Array`), or an array of strings for fixed- and variable-length
@@ -43,7 +86,17 @@ else throws an `Error` naming the type.
## Limits
- Read-only, and the whole file is held in memory (no range requests).
- Read-only. `open(bytes)` holds the whole file in memory.
- `openUrl`: a cross-origin server must allow CORS and expose
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
`Last-Modified` either, a file replaced on the server is detected only by
a change of length. A call keeps what it reads until it finishes, and
may fetch at most `maxFetch`; `read()` of a dataset that would take more
than 1 GiB to decode is refused (read windows of large datasets with
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in
`docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with
`value: null` and their `dtype`.
@@ -56,28 +109,55 @@ else throws an `Error` naming the type.
## Tests
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
`fixture.nc` (netCDF4) with `test/make_fixture.py`, then:
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
counter, `/norange/...` for a server without range support,
`/noexpose/...` and `/unexposed/...` for one that does not expose its
headers to CORS), then:
- runs `test/test.mjs` under Node: every dataset (whole and a strided
hyperslab), listing and attribute is compared with what libhdf5 reads
back, error paths are checked, and so are the page's DOM-free helpers
(`viewer-lib.js`);
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
and 512 B blocks), calls in flight at once, the request budget of
`big.h5` (listing it and reading small things of it must take at most 8
requests and under 5% of the file), the download fallback and its
limit, and the errors: HTTP status, a file that changes, a server that
sends the wrong bytes or stops honouring `Range`. Also the cross-origin
path with no exposed headers (`/noexpose/`: length from `HEAD`), the
size limits on `make_fixture.py`'s `limits.h5`, `hostile_vl.h5` (a heap
collection claiming 2 GiB) and `far.h5` (data at 3 GiB, served by a
mock), bodies longer than asked for, `headers` forms, `parallel`, and
sibling requests aborted after a failure.
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
to 16 MiB) read by URL with the same file read from bytes;
- runs `test/browser.sh`: loads the page in headless Chromium with
`?file=fixture.h5&path=...` for eight objects and checks the rendered tree,
types, shapes, attribute and value cells, and the error shown for an
unsupported type. Skipped when no Chromium is found (`CHROME` names one;
a Playwright download under `~/.cache/ms-playwright` is picked up).
Drag-and-drop and the file picker are not driven by it; they share
`load()` with the `?file=` path.
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
tree, types, shapes, attribute and value cells, the request counter, and
the error shown for an unsupported type; then the file from another
origin (localhost), with CORS exposing `Content-Range` and exposing
nothing (`/unexposed/`, where the server must see a `HEAD`), a server
without range support, and `big.h5` (a small dataset and a window of the large one,
with a single-digit percentage of the file fetched). Skipped when no
Chromium is found (`CHROME` names one; a Playwright download under
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
and the URL box are not driven by it; they share `setFile()` with the
`?file=` path.
The same expectations are checked natively, without Node, by
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI
container has no Node or browser).
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
corpus too), which is what CI runs (the CI container has no Node or
browser).
## Size
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`:
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
larger now and the table has not been re-measured: the reader has grown
since, and `openUrl` (2026-09-27) made the facade's range-read path
reachable from JavaScript and added promise glue and `remote.js`.
| | raw | gzip -9 |
|---|---:|---:|
+94 -44
View File
@@ -56,27 +56,39 @@
border: 1px solid var(--line); border-radius: 4px; }
.error { color: var(--bad); font-family: var(--mono); white-space: pre-wrap; }
.muted { color: var(--muted); }
form.url { display: flex; gap: 6px; margin-left: auto; flex: 1 1 320px; max-width: 560px; }
form.url input { flex: 1; min-width: 0; font: 13px var(--mono); padding: 4px 8px; background: var(--panel);
color: var(--ink); border: 1px solid var(--line); border-radius: 6px; }
.netstats { flex-basis: 100%; color: var(--muted); font-size: 12.5px; font-family: var(--mono); }
.netstats:empty { display: none; }
</style>
</head>
<body>
<header>
<h1>HDF5 Viewer</h1>
<span class="file" id="filename">no file</span>
<form class="url" id="urlform">
<input type="url" id="url" placeholder="https://…/file.h5 (read by HTTP range requests)" aria-label="File URL">
<button class="button" type="submit">Open URL</button>
</form>
<label class="button">Open file…<input type="file" id="picker" accept=".h5,.hdf5,.he5,.nc,.nc4,.cdf" hidden></label>
<span class="netstats" id="netstats"></span>
</header>
<main>
<nav id="tree"></nav>
<section id="detail">
<div class="drop">
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, or use “Open file…”.</p>
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; it is not uploaded anywhere.</p>
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, use “Open file…”, or give a URL.</p>
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; a local file is not uploaded anywhere.
A file given by URL is not downloaded: only the byte ranges each view needs are fetched (HTTP range requests),
and the requests and bytes it has cost are shown at the top.</p>
<p class="muted" id="version"></p>
</div>
</section>
</main>
<script type="module">
import init, { open, version } from "./pkg/clawhdf5_wasm.js";
import { joinPath, formatValue, viewWindow, toRows, perElement } from "./viewer-lib.js";
import init, { open, openUrl, version } from "./pkg/clawhdf5_wasm.js";
import { joinPath, formatValue, viewWindow, toRows, perElement, formatStats } from "./viewer-lib.js";
const $ = (id) => document.getElementById(id);
const el = (tag, props = {}, ...kids) => {
@@ -85,30 +97,58 @@ const el = (tag, props = {}, ...kids) => {
return e;
};
// The open file: an H5File (a local file, read in memory; synchronous
// methods) or a RemoteFile (openUrl: bytes fetched by range requests as
// needed; methods return promises). Every call below is awaited, so both
// work the same.
let file = null;
let selectedNode = null;
// Bumped by every selection, so a slow read for an object no longer shown
// does not overwrite the current one.
let showSeq = 0;
await init();
$("version").textContent = `clawhdf5 ${version()}`;
async function load(blob) {
const bytes = new Uint8Array(await blob.arrayBuffer());
// Requests and bytes a remote file has cost so far.
function updateStats() {
$("netstats").textContent = file && file.stats ? formatStats(file.stats()) : "";
}
async function setFile(name, opener) {
if (file) file.free();
file = null;
$("filename").textContent = blob.name;
selectedNode = null;
showSeq++;
$("filename").textContent = name;
$("tree").replaceChildren();
$("netstats").textContent = "";
$("detail").replaceChildren(el("p", { className: "muted", textContent: `Opening ${name}…` }));
try {
file = open(bytes);
file = await opener();
} catch (e) {
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${blob.name}: ${e.message}` }));
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${name}: ${e.message}` }));
return;
}
updateStats();
const root = el("ul", { className: "tree" });
root.append(treeNode("/", "/", "group"));
$("tree").append(root);
root.querySelector(".node").click();
await activate(root.querySelector(".node"));
}
async function load(blob) {
const bytes = new Uint8Array(await blob.arrayBuffer());
await setFile(blob.name, async () => open(bytes));
}
async function loadUrl(url) {
$("url").value = url;
await setFile(url, () => openUrl(new URL(url, location.href).href));
}
// A tree row: clicking it selects it (and expands a group); `activate`
// does the same and resolves once the listing and the detail are shown.
function treeNode(path, name, kind) {
const li = el("li");
const icon = el("span", { className: "icon", textContent: kind === "group" ? "▸" : "·" });
@@ -116,34 +156,39 @@ function treeNode(path, name, kind) {
row.dataset.path = path;
li.append(row);
let children = null;
row.addEventListener("click", () => {
row.activate = async () => {
if (selectedNode) selectedNode.classList.remove("selected");
row.classList.add("selected");
selectedNode = row;
const shown = show(path, kind);
if (kind === "group") {
if (children) {
children.hidden = !children.hidden;
} else {
children = el("ul", { className: "tree" });
li.append(children);
try {
for (const c of file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
for (const c of await file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
} catch (e) {
children.append(el("li", { className: "error", textContent: e.message }));
}
li.append(children);
updateStats();
}
icon.textContent = children.hidden ? "▸" : "▾";
}
show(path, kind);
});
await shown;
};
row.addEventListener("click", () => row.activate());
return li;
}
function attrsTable(path) {
const activate = (row) => row.activate();
async function attrsTable(path) {
let attrs, errors;
try {
attrs = file.attrs(path);
errors = file.attrErrors(path);
attrs = await file.attrs(path);
errors = await file.attrErrors(path);
} catch (e) {
return el("p", { className: "error", textContent: e.message });
}
@@ -157,14 +202,16 @@ function attrsTable(path) {
return el("div", { className: "scroll" }, t);
}
function show(path, kind) {
async function show(path, kind) {
const seq = ++showSeq;
const current = () => seq === showSeq;
const out = [el("h2", { textContent: path })];
if (kind === "dataset") {
let info;
try {
info = file.info(path);
info = await file.info(path);
} catch (e) {
$("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
if (current()) $("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
return;
}
const max = info.maxshape === null ? "—" : `(${info.maxshape.map((d) => d ?? "∞").join(", ")})`;
@@ -172,15 +219,16 @@ function show(path, kind) {
el("dt", { textContent: "type" }), el("dd", { textContent: info.dtype }),
el("dt", { textContent: "shape" }), el("dd", { textContent: `(${info.shape.join(", ")})` }),
el("dt", { textContent: "max shape" }), el("dd", { textContent: max })));
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
out.push(el("h3", { textContent: "Values" }), valuesView(path, info));
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
out.push(el("h3", { textContent: "Values" }), await valuesView(path, info, current));
} else {
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
}
$("detail").replaceChildren(...out);
updateStats();
if (current()) $("detail").replaceChildren(...out);
}
function valuesView(path, info) {
async function valuesView(path, info, current) {
const shape = info.shape;
const per = perElement(info.elementShape);
const state = { row: 0, col: 0, rows: 50, cols: 12, fixed: shape.slice(0, Math.max(0, shape.length - 2)).map(() => 0) };
@@ -202,15 +250,19 @@ function valuesView(path, info) {
if (shape.length >= 2) controls.append(num("column", "col"));
state.fixed.forEach((_, i) => controls.append(num(`dim ${i}`, "fixed", i)));
function render() {
let renderSeq = 0;
async function render() {
const seq = ++renderSeq;
const w = viewWindow(shape, state);
let res;
try {
res = w === null ? file.read(path) : file.readHyperslab(path, w.start, w.count);
res = w === null ? await file.read(path) : await file.readHyperslab(path, w.start, w.count);
} catch (e) {
body.replaceChildren(el("p", { className: "error", textContent: e.message }));
if (seq === renderSeq) body.replaceChildren(el("p", { className: "error", textContent: e.message }));
return;
}
updateStats();
if (seq !== renderSeq || !current()) return;
if (w === null) {
body.replaceChildren(el("pre", { textContent: toRows(res.data, 1, 1, per)[0][0] }));
return;
@@ -229,41 +281,39 @@ function valuesView(path, info) {
(shape.length >= 2 ? `, columns ${w.col}–${w.col + w.cols - 1}` : "") + ` of ${total} values` });
body.replaceChildren(t, note);
}
render();
await render();
return box;
}
// Expand the tree down to `path` and select it.
function reveal(path) {
async function reveal(path) {
const rowFor = (p) => [...document.querySelectorAll(".node")].find((n) => n.dataset.path === p);
let cur = "/";
for (const part of path.split("/").filter(Boolean)) {
const row = rowFor(cur);
if (!row) return;
const kids = row.parentElement.querySelector(":scope > ul");
if (!kids || kids.hidden) row.click();
if (!kids || kids.hidden) await activate(row);
cur = joinPath(cur, part);
}
const target = rowFor(cur);
if (target && target !== selectedNode) target.click();
if (target && target !== selectedNode) await activate(target);
}
// ?file=<url>&path=<object> opens a file from a URL (same origin, or one
// that allows CORS) and selects an object in it.
// that allows CORS) by range requests and selects an object in it.
const params = new URLSearchParams(location.search);
if (params.get("file")) {
const url = params.get("file");
try {
const resp = await fetch(url);
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
const blob = await resp.blob();
await load(new File([blob], url.split("/").pop()));
if (file && params.get("path")) reveal(params.get("path"));
} catch (e) {
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot fetch ${url}: ${e.message}` }));
}
await loadUrl(params.get("file"));
if (file && params.get("path")) await reveal(params.get("path"));
document.body.dataset.ready = "1";
}
$("urlform").addEventListener("submit", (e) => {
e.preventDefault();
const url = $("url").value.trim();
if (url) loadUrl(url);
});
$("picker").addEventListener("change", (e) => e.target.files[0] && load(e.target.files[0]));
document.addEventListener("dragover", (e) => { e.preventDefault(); document.body.classList.add("dragging"); });
document.addEventListener("dragleave", () => document.body.classList.remove("dragging"));
+60 -30
View File
@@ -3,10 +3,13 @@
#
# browser.sh FIXTURE_DIR
#
# FIXTURE_DIR holds fixture.h5 from make_fixture.py; ../pkg must be built.
# The page is opened with ?file=fixture.h5&path=<object>, which fetches the
# file, builds the tree down to <object> and shows it; the rendered DOM is
# dumped and checked for the values libhdf5 reads.
# FIXTURE_DIR holds fixture.h5 from make_fixture.py (and big.h5 when it was
# run with WASM_BIG_MB); ../pkg must be built. test/serve.py serves the
# page and, under /fix, the fixtures, with HTTP range requests (no copies,
# no links). The page is opened with ?file=fix/<name>&path=<object>, which
# opens the file with openUrl (range requests), builds the tree down to
# <object> and shows it; the rendered DOM is dumped and checked for the
# values libhdf5 reads and for the request counter.
#
# Browser: $CHROME, else chromium/google-chrome on PATH, else a Playwright
# download under ~/.cache/ms-playwright. Exit 3 when none is found.
@@ -31,63 +34,90 @@ if [ -z "$chrome" ] || [ ! -x "$chrome" ]; then
exit 3
fi
root="$(mktemp -d)"
profiles="$(mktemp -d)"
server=""
cleanup() {
[ -n "$server" ] && kill "$server" 2>/dev/null || true
rm -rf "$root"
rm -rf "$profiles"
}
trap cleanup EXIT
ln -s "$HERE/../index.html" "$HERE/../viewer-lib.js" "$HERE/../pkg" "$FIX/fixture.h5" "$root/"
port=$("$PY" -c 'import socket; s = socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1])')
"$PY" -m http.server --bind 127.0.0.1 --directory "$root" "$port" >/dev/null 2>&1 &
"$PY" "$HERE/serve.py" --root "fix=$FIX" --root "=$HERE/.." > "$profiles/port" &
server=$!
for _ in $(seq 50); do
"$PY" -c "import urllib.request; urllib.request.urlopen('http://127.0.0.1:$port/index.html')" 2>/dev/null && break
[ -s "$profiles/port" ] && break
sleep 0.1
done
port="$(head -1 "$profiles/port")"
fails=0
# A fresh profile per page: a second instance on the same profile fails.
render() {
local profile
profile="$(mktemp -d "$root/profile.XXXXXX")"
profile="$(mktemp -d "$profiles/profile.XXXXXX")"
"$chrome" --headless --no-sandbox --disable-gpu --user-data-dir="$profile" \
--virtual-time-budget=20000 \
--dump-dom "http://127.0.0.1:$port/index.html?file=fixture.h5&path=$1" 2>/dev/null
--dump-dom "http://127.0.0.1:$port/index.html?file=$1&path=$2" 2>/dev/null
}
# expect PATH TEXT...: every TEXT appears in the page rendered for PATH.
# expect FILE PATH TEXT...: every TEXT appears in the page rendered for PATH
# of FILE (a URL relative to the page); "re:TEXT" is an extended regex.
expect() {
local path="$1" dom
shift
dom="$(render "$path")"
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$path" | tr / _).html"
local file="$1" path="$2" dom
shift 2
dom="$(render "$file" "$path")"
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$file$path" | tr / _).html"
for text in "$@"; do
if ! grep -qF -- "$text" <<<"$dom"; then
echo "FAIL: page for $path lacks: $text" >&2
local flag=-qF
if [ "${text#re:}" != "$text" ]; then flag=-qE; text="${text#re:}"; fi
if ! grep $flag -- "$text" <<<"$dom"; then
echo "FAIL: page for $file $path lacks: $text" >&2
fails=$((fails + 1))
fi
done
echo "rendered $path"
echo "rendered $file $path"
}
H5=fix/fixture.h5
# Tree (root expanded; the group row carries its path) and root attributes.
expect "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
expect $H5 "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
'<td>big</td><td>9223372036854775813</td>' '(compound{x: f64, n: i32})'
# A chunked, deflated 2-D dataset: type, shape, the first window of values.
expect "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
'showing rows 0–5, columns 0–9 of 60 values'
# A chunked, deflated 2-D dataset: type, shape, the first window of values;
# the request counter.
expect $H5 "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
'showing rows 0–5, columns 0–9 of 60 values' 'request' 'fetched (100.0%)'
# Nested path revealed through the tree; big-endian float32.
expect "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
expect $H5 "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
# 64-bit integers stay exact; strings; array datatype cells.
expect "/u64" '<td>18446744073709551615</td>'
expect "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
expect $H5 "/u64" '<td>18446744073709551615</td>'
expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
# 3-D: leading dimension held at 0, window over the last two.
expect "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
# Unsupported type: an error, not values.
expect "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
# exposes Content-Range, and CORS that exposes nothing (/unexposed/: the
# browser hides Content-Range and ETag, so the page learns the length with
# a HEAD request, which the server's log must show).
expect "http://localhost:$port/$H5" "/grid" '<td>14.75</td>' 'request'
get() { "$PY" -c 'import sys, urllib.request; print(urllib.request.urlopen(sys.argv[1]).read().decode())' "$1"; }
get "http://127.0.0.1:$port/__reset" >/dev/null
expect "http://localhost:$port/unexposed/$H5" "/grid" '<td>0.25</td>' '<td>14.75</td>' 'request'
heads="$(get "http://127.0.0.1:$port/__stats" | "$PY" -c \
'import json, sys; print(sum(1 for l in json.load(sys.stdin)["log"] if l[4] == "HEAD" and l[0].endswith("fixture.h5")))')"
if [ "$heads" -lt 1 ]; then
echo "FAIL: no HEAD request for the length with unexposed headers" >&2
fails=$((fails + 1))
fi
# A server without range support: the file is downloaded whole, and says so.
expect norange/$H5 "/grid" '<td>14.75</td>' 'downloaded whole'
# A large file: a small dataset, and a window of the large one, fetch a few
# blocks (the counter shows a small percentage: "(1.5%)", not "(100.0%)").
if [ -f "$FIX/big.h5" ]; then
expect fix/big.h5 "/small" '<td>1.5</td>' '<td>3.25</td>' 're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
expect fix/big.h5 "/big" '<dd>(25000000)</dd>' '<td>0</td>' '<td>24.5</td>' \
're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
fi
if [ "$fails" -gt 0 ]; then
echo "browser: $fails checks failed" >&2
+106 -1
View File
@@ -3,7 +3,9 @@ reads back from them, for the clawhdf5-wasm tests.
python make_fixture.py OUT_DIR
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json.
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json,
and the limit-test files OUT_DIR/limits.h5 and OUT_DIR/hostile_vl.h5 (see
write_limits).
Both the Rust test (crates/clawhdf5-wasm/tests/h5py_interop.rs, native) and
the Node test (test.mjs, the built wasm package) compare against the same
expected.json, so the two check the same values.
@@ -14,6 +16,7 @@ encoded as strings so JSON.parse keeps 64-bit values exact.
"""
import json
import os
import sys
import warnings
from pathlib import Path
@@ -201,3 +204,105 @@ def slab_for(obj):
json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)},
open(out / "expected.json", "w"), indent=1, ensure_ascii=False)
HUGE_U8 = 2**28 + 1024
# The collection size hostile_vl.h5 claims: past 2 GiB, which a wasm32
# buffer cannot hold.
HOSTILE_GCOL_SIZE = 2**31 + 4096
# The file length a server claims for hostile_vl.h5 (the tests' mock fetch
# answers every range with zeros past the real bytes): 3 GiB, within what
# wasm32 opens, and room for the collection.
HOSTILE_LENGTH = 3 << 30
def write_limits(out):
"""Files for the size limits (the tests must get errors, not aborts):
- limits.h5: /huge_u8, 2^28 + 1024 bytes of u8 in compressed chunks
(a small file): read whole it would take over 2 GiB while decoding;
its last value is 7.
- hostile_vl.h5: a variable-length string dataset /a whose global heap
collection claims HOSTILE_GCOL_SIZE bytes, with the superblock's end
of file set to HOSTILE_LENGTH (libhdf5 cannot read it; it is only
served by a mock that claims that length).
- far.h5 and far.json: /x, 16 float64 values, whose contiguous data
address is moved FAR_SHIFT bytes on (past 2 GiB, the sign bit of a
wasm32 isize) in a file whose end of file is moved as far; the tests'
mock serves the data there, to show offsets up to 4 GiB work on
wasm32.
"""
with h5py.File(out / "limits.h5", "w") as f:
d = f.create_dataset("huge_u8", shape=(HUGE_U8,), dtype="u1",
chunks=(1 << 20,), compression="gzip")
d[-1] = 7
path = out / "hostile_vl.h5"
with h5py.File(path, "w", libver="earliest") as f:
f.create_dataset("a", data=["x", "yy"], dtype=h5py.string_dtype())
b = bytearray(path.read_bytes())
assert b[8] == 0, "a version 0 superblock"
b[40:48] = HOSTILE_LENGTH.to_bytes(8, "little") # end of file address
at = b.index(b"GCOL")
b[at + 8:at + 16] = HOSTILE_GCOL_SIZE.to_bytes(8, "little")
path.write_bytes(bytes(b))
path = out / "far.h5"
values = np.arange(16, dtype="<f8") * 1.5
with h5py.File(path, "w", libver="earliest") as f:
f.create_dataset("x", data=values)
data_at = f["x"].id.get_offset()
b = bytearray(path.read_bytes())
# The layout message: the data's address, then its size.
old = data_at.to_bytes(8, "little") + (values.nbytes).to_bytes(8, "little")
assert b.count(old) == 1
at = b.index(old)
b[at:at + 8] = (data_at + FAR_SHIFT).to_bytes(8, "little")
length = len(b) + FAR_SHIFT
b[40:48] = length.to_bytes(8, "little")
path.write_bytes(bytes(b))
json.dump({"data_at": data_at, "far_at": data_at + FAR_SHIFT,
"nbytes": values.nbytes, "length": length,
"values": [float(x) for x in values]},
open(out / "far.json", "w"))
# far.h5's data moves this far: past 2 GiB, below 4 GiB.
FAR_SHIFT = 3 << 30
write_limits(out)
def write_big(path, megabytes):
"""A large file for the range-request tests (`openUrl`): `/big`, about
`megabytes` MB of float64 in 1 MiB chunks, written after a small
dataset and a group, so listing and reading `/small` touch a few blocks
of the file and a window of `/big` one chunk. Returns what h5py reads
back."""
n = megabytes * 1_000_000 // 8
chunk = 1 << 17
with h5py.File(path, "w") as f:
f.attrs["note"] = "large file for range reads"
f.create_dataset("small", data=np.array([1.5, -2.0, 3.25]))
g = f.create_group("meta")
g.attrs["units"] = "m"
g.create_dataset("ids", data=np.arange(10, dtype="<i4"))
big = f.create_dataset("big", shape=(n,), dtype="<f8", chunks=(chunk,))
for s in range(0, n, 1 << 22):
e = min(n, s + (1 << 22))
big[s:e] = np.arange(s, e, dtype="<f8") * 0.5
start = n // 2 + 12_345
with h5py.File(path, "r") as f:
return {
"size": path.stat().st_size,
"list": {"groups": ["meta"], "datasets": ["big", "small"]},
"small": [float(x) for x in f["small"][()]],
"ids": [str(int(x)) for x in f["meta/ids"][()]],
"big_shape": list(f["big"].shape),
"window": {"start": start, "count": 10,
"values": [float(x) for x in f["big"][start:start + 10]]},
}
big_mb = int(os.environ.get("WASM_BIG_MB", "0"))
if big_mb > 0:
json.dump(write_big(out / "big.h5", big_mb), open(out / "big.json", "w"), indent=1)
+26 -5
View File
@@ -1,11 +1,17 @@
#!/usr/bin/env bash
# Build the wasm package (../build.sh), test it under Node against files
# written by h5py and netCDF4 (make_fixture.py), then load the viewer page
# in headless Chromium if one is found (browser.sh).
# written by h5py and netCDF4 (make_fixture.py) — opened from bytes, and
# opened by URL from a local range-capable HTTP server (serve.py) — then
# load the viewer page in headless Chromium if one is found (browser.sh).
#
# Needs node, the wasm-bindgen CLI (see ../build.sh) and a Python with h5py,
# netCDF4 and numpy: CLAWHDF5_PYTHON names it (default python3). Without that
# Python the test is skipped, unless CLAWHDF5_REQUIRE_INTEROP=1.
#
# WASM_BIG_MB (default 200) sizes the large file of the range-request
# budget test (0 leaves it out); it is written under TMPDIR.
# CLAWHDF5_WASM_CORPUS=DIR also compares every HDF5 file under DIR (up to
# 16 MiB) read over HTTP with the same file read from bytes.
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
@@ -23,9 +29,24 @@ fi
bash "$HERE/../build.sh"
fix="$(mktemp -d)"
trap 'rm -rf "$fix"' EXIT
"$PY" "$HERE/make_fixture.py" "$fix"
node "$HERE/test.mjs" "$HERE/../pkg" "$fix"
server=""
cleanup() {
[ -n "$server" ] && kill "$server" 2>/dev/null || true
rm -rf "$fix"
}
trap cleanup EXIT
WASM_BIG_MB="${WASM_BIG_MB:-200}" "$PY" "$HERE/make_fixture.py" "$fix"
# The fixtures (and the corpus) over HTTP with range support.
roots=(--root "fix=$fix")
[ -n "${CLAWHDF5_WASM_CORPUS:-}" ] && roots+=(--root "corpus=$CLAWHDF5_WASM_CORPUS")
"$PY" "$HERE/serve.py" "${roots[@]}" > "$fix/port" &
server=$!
for _ in $(seq 50); do
[ -s "$fix/port" ] && break
sleep 0.1
done
node "$HERE/test.mjs" "$HERE/../pkg" "$fix" "http://127.0.0.1:$(head -1 "$fix/port")"
# The page itself, in headless Chromium when one is available.
status=0
+223
View File
@@ -0,0 +1,223 @@
"""A static HTTP server for the wasm tests, with HTTP Range support and
request counting.
python serve.py [--root PREFIX=DIR ...]
Serves each DIR under URL PREFIX (the first match wins; PREFIX "" is the
site root), prints the port on its first line of stdout, and runs until
killed. No symlinks or copies are made: files are read where they are.
- `Range: bytes=a-b`, `bytes=a-` and `bytes=-n` get 206 with Content-Range,
an unsatisfiable range 416; every file answer carries an ETag, and CORS
headers exposing Content-Range, so a page on another origin can use it.
- Under `/norange/...` the same files are served but Range is ignored
(200 with the whole file), as by a server without range support.
- Under `/noexpose/...` ranges are served without Content-Range, ETag,
Last-Modified or Accept-Ranges, and CORS exposes none of them: what a
page sees of a cross-origin server that does not list them in
Access-Control-Expose-Headers. The length comes from Content-Length of
a HEAD request (always readable).
- Under `/unexposed/...` ranges are served with all those headers, but
CORS exposes none of them: a browser page on another origin cannot read
them (test/browser.sh loads the page from 127.0.0.1 and the file from
localhost), so it has to take the same path.
- `GET /__stats` returns `{"requests": n, "bytes": n, "log": [...]}` for
file requests since the last `GET /__reset`, which zeroes them.
"""
import argparse
import hashlib
import json
import os
import posixpath
import sys
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import unquote, urlsplit
TYPES = {
".html": "text/html; charset=utf-8",
".js": "text/javascript; charset=utf-8",
".mjs": "text/javascript; charset=utf-8",
".wasm": "application/wasm",
".json": "application/json",
".ts": "text/plain; charset=utf-8",
}
lock = threading.Lock()
stats = {"requests": 0, "bytes": 0, "log": []}
def resolve(roots, path):
"""The file for URL `path`, or None. `..` never leaves a root."""
parts = [p for p in posixpath.normpath(unquote(path)).split("/") if p]
if any(p in (".", "..") for p in parts):
return None
for prefix, root in roots:
pre = [p for p in prefix.split("/") if p]
if parts[: len(pre)] == pre:
rest = parts[len(pre):] or ["index.html"]
f = os.path.join(root, *rest)
if os.path.isfile(f):
return f
return None
def parse_range(header, size):
"""(start, end exclusive) for a single `bytes=` range, "bad" when
unsatisfiable, None when absent or unparsable (served whole)."""
if not header or not header.startswith("bytes=") or "," in header:
return None
a, _, b = header[len("bytes="):].strip().partition("-")
try:
if a == "":
n = int(b)
return (max(0, size - n), size) if n > 0 and size > 0 else "bad"
start = int(a)
end = int(b) + 1 if b else size
except ValueError:
return None
if start >= size or end <= start:
return "bad"
return start, min(end, size)
def make_handler(roots):
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def log_message(self, *args):
pass
def cors(self, expose=True):
self.send_header("Access-Control-Allow-Origin", "*")
if expose:
self.send_header("Access-Control-Expose-Headers",
"Content-Range, Content-Length, ETag, Accept-Ranges")
def do_OPTIONS(self):
self.send_response(204)
self.cors()
self.send_header("Access-Control-Allow-Headers", "Range")
self.send_header("Content-Length", "0")
self.end_headers()
def do_HEAD(self):
self.serve(head=True)
def do_GET(self):
self.serve(head=False)
def json(self, obj):
body = json.dumps(obj).encode()
self.send_response(200)
self.cors()
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.send_header("Cache-Control", "no-store")
self.end_headers()
self.wfile.write(body)
def serve(self, head):
path = urlsplit(self.path).path
if path == "/__stats":
with lock:
return self.json(stats)
if path == "/__reset":
with lock:
stats.update(requests=0, bytes=0, log=[])
return self.json({})
# ranges: honour Range; send_all: send Content-Range, ETag and
# Accept-Ranges; expose: list them for CORS.
ranges = send_all = expose = True
if path.startswith("/norange/"):
ranges = False
path = path[len("/norange"):]
elif path.startswith("/noexpose/"):
expose = send_all = False
path = path[len("/noexpose"):]
elif path.startswith("/unexposed/"):
expose = False
path = path[len("/unexposed"):]
f = resolve(roots, path)
if f is None:
self.send_response(404)
self.cors()
self.send_header("Content-Length", "0")
self.end_headers()
return
size = os.path.getsize(f)
st = os.stat(f)
etag = '"%s"' % hashlib.sha1(
f"{f}:{size}:{st.st_mtime_ns}".encode()).hexdigest()[:16]
r = parse_range(self.headers.get("Range"), size) if ranges else None
if r == "bad":
self.send_response(416)
self.cors()
self.send_header("Content-Range", f"bytes */{size}")
self.send_header("Content-Length", "0")
self.end_headers()
return
start, end = r if r else (0, size)
self.send_response(206 if r else 200)
self.cors(expose)
ext = os.path.splitext(f)[1]
self.send_header("Content-Type", TYPES.get(ext, "application/octet-stream"))
self.send_header("Content-Length", str(end - start))
self.send_header("Cache-Control", "no-store")
if send_all:
self.send_header("ETag", etag)
if ranges:
self.send_header("Accept-Ranges", "bytes")
if r:
self.send_header("Content-Range", f"bytes {start}-{end - 1}/{size}")
self.end_headers()
with lock:
# A HEAD is a request too (openUrl makes one when it cannot
# see Content-Range); it sends no bytes.
stats["requests"] += 1
sent = 0 if head else end - start
stats["bytes"] += sent
stats["log"].append([path, start, end, 206 if r else 200, "HEAD" if head else "GET"])
if not head:
with open(f, "rb") as fh:
fh.seek(start)
left = end - start
try:
while left:
buf = fh.read(min(left, 1 << 20))
if not buf:
break
self.wfile.write(buf)
left -= len(buf)
except (BrokenPipeError, ConnectionResetError):
pass
return Handler
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--root", action="append", default=[],
help="PREFIX=DIR: serve DIR under URL PREFIX")
args = ap.parse_args()
roots = []
for spec in args.root:
prefix, _, d = spec.partition("=")
roots.append((prefix, os.path.abspath(d)))
class Server(ThreadingHTTPServer):
def handle_error(self, request, client_address):
# A client that drops a connection (a cancelled download) is
# not an error of the server.
if not isinstance(sys.exc_info()[1], (ConnectionError, TimeoutError)):
super().handle_error(request, client_address)
httpd = Server(("127.0.0.1", 0), make_handler(roots))
httpd.daemon_threads = True
print(httpd.server_address[1], flush=True)
sys.stdout.close()
httpd.serve_forever()
if __name__ == "__main__":
main()
+531 -32
View File
@@ -1,20 +1,31 @@
// Node test of the built wasm package (the exact pkg/ the viewer page loads)
// and the viewer's DOM-free helpers. Run by test/run.sh:
// node test.mjs PKG_DIR FIXTURE_DIR
// node test.mjs PKG_DIR FIXTURE_DIR [SERVER_URL]
// FIXTURE_DIR holds fixture.h5, fixture.nc and expected.json from
// make_fixture.py (values as libhdf5 reads them back).
// make_fixture.py (values as libhdf5 reads them back), and big.h5/big.json
// when it was run with WASM_BIG_MB. SERVER_URL is test/serve.py serving
// FIXTURE_DIR under /fix (and $CLAWHDF5_WASM_CORPUS under /corpus): with
// it, every check is repeated on files opened with openUrl (HTTP range
// requests), and the request budget, the full-download fallback and the
// error paths of openUrl are tested.
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import { join } from "node:path";
import { existsSync, readFileSync, readdirSync, statSync } from "node:fs";
import { join, relative } from "node:path";
import { pathToFileURL } from "node:url";
const [pkgDir, fixDir] = process.argv.slice(2);
const [pkgDir, fixDir, base] = process.argv.slice(2);
const pkg = await import(pathToFileURL(join(pkgDir, "clawhdf5_wasm.js")));
pkg.initSync({ module: readFileSync(join(pkgDir, "clawhdf5_wasm_bg.wasm")) });
const lib = await import(pathToFileURL(join(import.meta.dirname, "..", "viewer-lib.js")));
let checks = 0;
const eq = (a, b, msg) => { assert.deepEqual(a, b, msg); checks++; };
// A call that must fail: a thrown Error (open) or a rejected promise
// (openUrl), whose message matches `re`.
const fails = async (fn, re, msg) => {
await assert.rejects(async () => fn(), (e) => e instanceof Error && re.test(e.message), msg);
checks++;
};
const ARRAY_TYPES = {
f32: Float32Array, f64: Float64Array, i8: Int8Array, i16: Int16Array, i32: Int32Array,
@@ -54,13 +65,13 @@ function checkAttr(ctx, a, want) {
assert.fail(`${ctx}: unknown expectation ${JSON.stringify(want)}`);
}
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
for (const [name, exp] of Object.entries(expected)) {
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
// Every listing, dataset, error and attribute of `file` against what
// libhdf5 reads (expected.json). `file` is an H5File (synchronous methods)
// or a RemoteFile (promises): every call is awaited.
async function checkFile(name, exp, file) {
for (const [path, want] of Object.entries(exp.lists)) {
eq(file.kind(path), "group", `${name}:${path} kind`);
const list = file.list(path);
eq(await file.kind(path), "group", `${name}:${path} kind`);
const list = await file.list(path);
for (const [kind, key] of [["group", "groups"], ["dataset", "datasets"]]) {
eq(list.filter((c) => c.kind === kind).map((c) => c.name).sort(), want[key], `${name}:${path} ${key}`);
}
@@ -68,58 +79,67 @@ for (const [name, exp] of Object.entries(expected)) {
for (const [path, want] of Object.entries(exp.datasets)) {
const ctx = `${name}:${path}`;
eq(file.kind(path), "dataset", `${ctx} kind`);
eq(await file.kind(path), "dataset", `${ctx} kind`);
if (want.unavailable) {
// The wasm build has no zstd (it links C): a clear error, no data.
assert.throws(() => file.read(path), (e) => e.message.includes(want.unavailable), ctx);
checks++;
await fails(() => file.read(path), new RegExp(want.unavailable), ctx);
continue;
}
const info = file.info(path);
const info = await file.info(path);
eq([...info.shape, ...info.elementShape], want.shape, `${ctx} info shape`);
const r = file.read(path);
const r = await file.read(path);
eq(r.shape, want.shape, `${ctx} shape`);
eq(r.dtype, info.dtype, `${ctx} dtype`);
assert.ok(r.data instanceof ARRAY_TYPES[want.kind], `${ctx}: ${r.data.constructor.name} for ${want.kind}`);
eq(values(want.kind, r.data), want.values, ctx);
if (want.slab) {
const s = want.slab;
const part = file.readHyperslab(path, s.start, s.count, s.stride);
const part = await file.readHyperslab(path, s.start, s.count, s.stride);
eq(part.shape, s.shape, `${ctx} slab shape`);
eq(values(want.kind, part.data), s.values, `${ctx} slab`);
}
}
for (const [path, what] of Object.entries(exp.errors)) {
assert.throws(() => file.read(path), (e) => e instanceof Error && e.message.includes(what), `${name}:${path}`);
checks++;
await fails(() => file.read(path), new RegExp(what), `${name}:${path}`);
}
for (const [path, want] of Object.entries(exp.attrs)) {
const attrs = file.attrs(path);
eq(file.attrErrors(path), [], `${name}:${path} attr errors`);
const attrs = await file.attrs(path);
eq(await file.attrErrors(path), [], `${name}:${path} attr errors`);
const seen = attrs.filter((a) => !a.name.startsWith("_") && !exp.skip_attrs.includes(a.name));
eq(seen.map((a) => a.name).sort(), Object.keys(want).sort(), `${name}:${path} attr names`);
for (const a of seen) checkAttr(`${name}:${path}@${a.name}`, a, want[a.name]);
}
}
// The error paths of the reader, for an H5File or a RemoteFile of
// fixture.h5.
async function checkErrors(h5) {
await fails(() => h5.read("/nope"), /./);
await fails(() => h5.list("/grid"), /not a group/);
await fails(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
await fails(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
await fails(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
await fails(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
// Big integers stay exact.
eq((await h5.read("/u64")).data[0], 18446744073709551615n, "u64 max");
}
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
for (const [name, exp] of Object.entries(expected)) {
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
await checkFile(name, exp, file);
file.free();
}
// Errors reach JavaScript as thrown Errors, never as data.
const h5 = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.h5"))));
const throwsMsg = (fn, re) => { assert.throws(fn, (e) => e instanceof Error && re.test(e.message)); checks++; };
throwsMsg(() => pkg.open(new Uint8Array(64)), /./);
throwsMsg(() => h5.read("/nope"), /./);
throwsMsg(() => h5.list("/grid"), /not a group/);
throwsMsg(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
throwsMsg(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
throwsMsg(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
throwsMsg(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
await fails(() => pkg.open(new Uint8Array(64)), /./);
await checkErrors(h5);
// Info for a dataset with an unlimited dimension (netCDF "time").
const nc = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.nc"))));
eq(nc.info("/time").maxshape, [null], "unlimited dimension is null");
// Big integers stay exact.
eq(h5.read("/u64").data[0], 18446744073709551615n, "u64 max");
eq(typeof pkg.version(), "string", "version");
// Viewer helpers.
@@ -137,7 +157,486 @@ eq(lib.toRows(pairs.data, 2, 1, lib.perElement([2])), [["[2, 3]"], ["[4, 5]"]],
eq(lib.formatValue(0.1 + 0.2), "0.3", "float formatting");
eq(lib.formatValue(2n ** 64n - 1n), "18446744073709551615", "bigint formatting");
eq(lib.formatValue("x"), '"x"', "string formatting");
eq(lib.formatBytes(0), "0 B", "bytes");
eq(lib.formatBytes(1536), "1.5 KiB", "KiB");
eq(lib.formatBytes(200 * 1024 * 1024), "200 MiB", "MiB");
eq(lib.formatStats({ lazy: true, requests: 3, bytesFetched: 2 << 20, size: 200 << 20 }),
"3 requests, 2 MiB of 200 MiB fetched (1.0%)", "stats line");
eq(lib.formatStats({ lazy: false, requests: 1, bytesFetched: 1024, size: 1024 }),
"downloaded whole (1 KiB): the server does not support range requests", "stats line, no ranges");
h5.free();
nc.free();
console.log(`wasm package: ${checks} checks passed`);
if (base) await remoteTests();
async function serverStats() {
return (await fetch(`${base}/__stats`)).json();
}
async function remoteTests() {
const before = checks;
await fetch(`${base}/__reset`);
// The same checks over HTTP range requests, at the default block size
// and at 512-byte blocks with a 4 KiB cache (almost every structure read
// a miss, evictions between calls).
for (const opts of [undefined, { blockSize: 512, cacheSize: 4096 }]) {
for (const [name, exp] of Object.entries(expected)) {
const f = await pkg.openUrl(`${base}/fix/${name}`, opts);
await checkFile(`${name} (openUrl ${JSON.stringify(opts ?? {})})`, exp, f);
const st = f.stats();
eq(st.lazy, true, "read by ranges");
eq(st.size, statSync(join(fixDir, name)).size, "size");
f.free();
}
}
const remote = await pkg.openUrl(`${base}/fix/fixture.h5`);
await checkErrors(remote);
eq((await (await pkg.openUrl(`${base}/fix/fixture.nc`)).info("/time")).maxshape, [null], "remote unlimited");
// What the page counts is what the server served.
await fetch(`${base}/__reset`);
const counted = await pkg.openUrl(`${base}/fix/fixture.nc`, { blockSize: 1024 });
await counted.read("/temp");
const server = await serverStats();
eq(counted.stats().requests, server.requests, "requests counted");
eq(counted.stats().bytesFetched, server.bytes, "bytes counted");
// A custom fetch is used for every request, with the caller's headers,
// given in any form fetch takes (a caller's Range is not sent).
for (const headers of [{ "X-Test": "1" }, new Headers({ "X-Test": "1", Range: "bytes=0-0" }), [["X-Test", "1"]]]) {
let calls = 0;
const seen = new Set();
const viaCustom = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 4096,
headers,
fetch: (url, init) => {
calls++;
const h = new Headers(init.headers);
seen.add(`${h.get("X-Test")} ${h.get("Range").startsWith("bytes=0-0") ? "caller's range" : "ours"}`);
return fetch(url, init);
},
});
const kind = headers.constructor.name;
eq((await viaCustom.read("/sensors/temp")).data[0], 21.5, `custom fetch values (${kind})`);
eq(calls, viaCustom.stats().requests, `custom fetch calls (${kind})`);
eq([...seen], ["1 ours"], `headers passed (${kind})`);
}
// parallel must be a positive integer.
for (const parallel of [0, -1, 1.5, "4", NaN, 5000]) {
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { parallel }), /parallel/, `parallel ${String(parallel)}`);
}
// When one range request fails, the others in flight are aborted and no
// more start (fetchRanges, called directly: remote.js is in the package).
{
const snippets = join(pkgDir, "snippets");
const dir = readdirSync(snippets).find((d) => existsSync(join(snippets, d, "js", "remote.js")));
const remote = await import(pathToFileURL(join(snippets, dir, "js", "remote.js")));
let started = 0;
const aborted = [];
const slowFetch = (url, init) => {
const k = started++;
if (k === 2) return Promise.resolve(new Response(null, { status: 500 }));
return new Promise((resolve, reject) => {
const t = setTimeout(() => resolve(new Response(new Uint8Array(10), { status: 206 })), 200);
init.signal?.addEventListener("abort", () => {
clearTimeout(t);
aborted.push(k);
reject(new DOMException("aborted", "AbortError"));
});
});
};
const ranges = Array.from({ length: 20 }, (_, i) => [i * 10, i * 10 + 10]).flat();
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: 3 }, null, 200),
/HTTP 500/, "a failed range request is the error");
await new Promise((r) => setTimeout(r, 300));
eq(started, 3, "no request starts after a failure");
eq(aborted.sort(), [0, 1], "requests in flight are aborted");
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: "x" }, null, 200),
/parallel must be a positive integer/, "fetchRanges checks parallel");
}
// Calls in flight at once share the cache (a 1 KiB budget: nothing is
// evicted while any of them runs) and each gets its own answer.
const both = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, cacheSize: 1024 });
const [g, t, l, a] = await Promise.all([
both.read("/grid"), both.read("/sensors/temp"), both.list("/sensors"), both.attrs("/"),
]);
const exp = expected["fixture.h5"];
eq(Array.from(g.data), exp.datasets["/grid"].values, "concurrent /grid");
eq(Array.from(t.data), exp.datasets["/sensors/temp"].values, "concurrent /sensors/temp");
eq(l.map((c) => c.name).sort(), [...exp.lists["/sensors"].groups, ...exp.lists["/sensors"].datasets].sort(), "concurrent list");
eq(a.length > 0, true, "concurrent attrs");
assert.ok(both.stats().cachedBytes <= 1024, "trimmed to the budget when idle");
// Listing and reading small things of a large file fetches a few blocks,
// not the file.
if (existsSync(join(fixDir, "big.json"))) {
const big = JSON.parse(readFileSync(join(fixDir, "big.json"), "utf8"));
await fetch(`${base}/__reset`);
const f = await pkg.openUrl(`${base}/fix/big.h5`);
const list = await f.list("/");
eq(list.filter((c) => c.kind === "group").map((c) => c.name), big.list.groups, "big: groups");
eq(list.filter((c) => c.kind === "dataset").map((c) => c.name).sort(), big.list.datasets, "big: datasets");
eq(Array.from((await f.read("/small")).data), big.small, "big: /small");
eq(Array.from((await f.read("/meta/ids")).data, String), big.ids, "big: /meta/ids");
eq((await f.info("/big")).shape, big.big_shape, "big: shape");
eq((await f.attrs("/meta"))[0].value, "m", "big: group attribute");
const win = await f.readHyperslab("/big", [big.window.start], [big.window.count]);
eq(Array.from(win.data), big.window.values, "big: window");
const st = f.stats();
const server = await serverStats();
eq(st.requests, server.requests, "big: requests counted");
eq(st.bytesFetched, server.bytes, "big: bytes counted");
console.log(`big.h5 (${big.size} bytes): listed, 2 small reads, attributes, info and a window in ` +
`${server.requests} requests, ${server.bytes} bytes (${(100 * server.bytes / big.size).toFixed(2)}%)`);
assert.ok(server.requests <= 8, `big: ${server.requests} requests`);
assert.ok(server.bytes * 20 < big.size, `big: ${server.bytes} bytes fetched`);
checks += 2;
}
// A cross-origin server that does not expose Content-Range, ETag or
// Last-Modified (serve.py's /noexpose/): the length comes from a HEAD
// request and answers are checked by their length alone. Every fixture
// check, calls in flight at once, and what the page counts.
for (const opts of [undefined, { blockSize: 512, cacheSize: 1024 }]) {
for (const [name, exp] of Object.entries(expected)) {
await fetch(`${base}/__reset`);
const f = await pkg.openUrl(`${base}/noexpose/fix/${name}`, opts);
await checkFile(`${name} (no exposed headers, ${JSON.stringify(opts ?? {})})`, exp, f);
const st = f.stats();
eq(st.lazy, true, "no exposed headers: read by ranges");
eq(st.size, statSync(join(fixDir, name)).size, "no exposed headers: size from HEAD");
const server = await serverStats();
eq(server.log.filter((l) => l[4] === "HEAD").length, 1, "no exposed headers: one HEAD");
eq(st.requests, server.requests, "no exposed headers: requests counted");
eq(st.bytesFetched, server.bytes, "no exposed headers: bytes counted");
f.free();
}
}
{
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, { blockSize: 512, cacheSize: 0, parallel: 3 });
const exp = expected["fixture.h5"];
const paths = ["/grid", "/sensors/temp", "/vlen_str", "/cube"];
const got = await Promise.all([...paths, ...paths].map((p) => f.read(p)));
got.forEach((r, i) => eq(values(exp.datasets[paths[i % 4]].kind, r.data), exp.datasets[paths[i % 4]].values,
`no exposed headers: concurrent ${paths[i % 4]}`));
await checkErrors(f);
}
// Without a validator a changed file cannot be told apart; a short
// answer still can.
await fails(async () => {
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
blockSize: 512,
fetch: async (url, init) => {
const r = await fetch(url, init);
if (init.method === "HEAD" || init.headers.Range === "bytes=0-511") return r;
return new Response((await r.arrayBuffer()).slice(1), { status: 206 });
},
});
await f.read("/grid");
}, /got \d+/, "no exposed headers: short answer");
// No HEAD length either: a clear error.
await fails(() => pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
fetch: async (url, init) => (init.method === "HEAD" ? new Response(null, { status: 405 }) : fetch(url, init)),
}), /cannot learn the file's size/, "no exposed headers, no HEAD");
// A server without range support: downloaded whole (the default), or
// refused.
const whole = await pkg.openUrl(`${base}/norange/fix/fixture.h5`);
await checkFile("fixture.h5 (no range support)", expected["fixture.h5"], whole);
eq(whole.stats().lazy, false, "downloaded whole");
eq(whole.stats().requests, 1, "one request");
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fallback: "error" }),
/does not support HTTP range requests/, "fallback: error");
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000 }),
/more than maxDownload/, "maxDownload");
// Without a Content-Length the body is streamed, and stopped at the limit.
const undeclared = async (url, init) => {
const r = await fetch(url, init);
return new Response(r.body, { status: r.status });
};
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000, fetch: undeclared }),
/more than maxDownload/, "maxDownload, streamed");
const streamed = await pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fetch: undeclared });
eq(Array.from((await streamed.read("/sensors/temp")).data), [21.5, 22, 22.25], "streamed download");
// Errors: HTTP status, not HDF5, bad options, a file that changes, a
// server that answers with the wrong bytes.
await fails(() => pkg.openUrl(`${base}/fix/missing.h5`), /HTTP 404/, "404");
await fails(() => pkg.openUrl(`${base}/fix/expected.json`), /./, "not HDF5");
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 100 }), /blockSize/, "blockSize");
const tamper = (edit) => async (url, init) => {
const r = await fetch(url, init);
return init.headers.Range === "bytes=0-511" ? r : edit(r);
};
const withHeaders = async (r, headers) => {
const h = new Headers(r.headers);
for (const [k, v] of Object.entries(headers)) h.set(k, v);
return new Response(await r.arrayBuffer(), { status: r.status, headers: h });
};
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, fetch: tamper((r) => withHeaders(r, { ETag: '"other"' })) });
await f.read("/grid");
}, /changed on the server/, "changed file");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => new Response((await r.arrayBuffer()).slice(1), { status: 206 })),
});
await f.read("/grid");
}, /got \d+/, "short answer");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => new Response(await r.arrayBuffer(), { status: 200 })),
});
await f.read("/grid");
}, /stopped honouring range requests/, "200 mid-file");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
});
await f.read("/grid");
}, /the server sent 0-511/, "wrong range");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes */25752" })),
});
await f.read("/grid");
}, /unusable Content-Range/, "unusable Content-Range");
await limitTests();
await floodTests();
// Corpus files: what the viewer can show of each is the same read by
// ranges as in memory (an error wherever it gives one).
const corpus = process.env.CLAWHDF5_WASM_CORPUS;
if (corpus) await corpusTests(corpus);
console.log(`openUrl: ${checks - before} checks passed`);
}
// A fetch that serves `buf` as a file of `total` bytes (zeros past the end
// of `buf`, and `extra` = [[offset, bytes], ...] laid over them), counting
// its calls. `hide` leaves out Content-Range and the validators, as a
// cross-origin server that exposes neither does; HEAD then gives the length.
function mockFetch(buf, { total = buf.length, extra = [], hide = false } = {}) {
const f = async (url, init) => {
f.calls++;
if (init.method === "HEAD") {
return new Response(null, { status: 200, headers: { "Content-Length": String(total) } });
}
const m = /^bytes=(\d+)-(\d+)$/.exec(new Headers(init.headers).get("Range"));
const a = Number(m[1]);
const b = Math.min(Number(m[2]) + 1, total);
const out = new Uint8Array(b - a);
for (const [at, bytes] of [[0, buf], ...extra]) {
const from = Math.max(a, at);
const to = Math.min(b, at + bytes.length);
if (from < to) out.set(bytes.subarray(from - at, to - at), from - a);
}
const headers = { "Content-Length": String(b - a) };
if (!hide) headers["Content-Range"] = `bytes ${a}-${b - 1}/${total}`;
return new Response(out, { status: 206, headers });
};
f.calls = 0;
return f;
}
// Sizes a hostile server or a large dataset can name are errors, never an
// allocation that aborts the module (which would take every open file on
// the page with it); see write_limits in make_fixture.py.
async function limitTests() {
// Read whole, /huge_u8 (2^28 + 1024 bytes) would take over 2 GiB while
// decoding: refused before its chunks are fetched; a window reads.
const n = 2 ** 28 + 1024;
const limits = readFileSync(join(fixDir, "limits.h5"));
const local = pkg.open(new Uint8Array(limits));
await fails(() => local.read("/huge_u8"), /readHyperslab/, "huge read, in memory");
const remote = await pkg.openUrl(`${base}/fix/limits.h5`);
const before = remote.stats().requests;
await fails(() => remote.read("/huge_u8"), /readHyperslab/, "huge read, openUrl");
eq(remote.stats().requests, before, "huge read fetched nothing");
eq(Array.from((await remote.readHyperslab("/huge_u8", [n - 4], [4])).data), [0, 0, 0, 7], "huge window");
eq(Array.from(local.readHyperslab("/huge_u8", [n - 4], [4]).data), [0, 0, 0, 7], "huge window, in memory");
local.free();
// A server claiming 3 GiB, a heap collection claiming 2 GiB + 4 KiB: an
// error after a few requests (it used to fetch 2 GiB, then abort).
const hostile = mockFetch(new Uint8Array(readFileSync(join(fixDir, "hostile_vl.h5"))), { total: 3 * 2 ** 30 });
const h = await pkg.openUrl("http://hostile.invalid/h.h5", { fetch: hostile });
await fails(() => h.read("/a"), /maxFetch/, "hostile collection size");
assert.ok(hostile.calls <= 4, `hostile: ${hostile.calls} requests`);
// The module survived: files open and read.
eq((await (await pkg.openUrl(`${base}/fix/fixture.h5`)).read("/sensors/temp")).data[0], 21.5, "alive after hostile");
// A call that would fetch more than maxFetch fails before fetching it.
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, maxFetch: 1024 });
await f.read("/grid");
}, /maxFetch/, "maxFetch");
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxFetch: 2 ** 31 }), /maxFetch/, "maxFetch range");
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxDownload: 2 ** 31 }), /maxDownload/, "maxDownload range");
// Offsets past 2 GiB work on wasm32: far.h5's data sits at 3 GiB.
const far = JSON.parse(readFileSync(join(fixDir, "far.json"), "utf8"));
const farBytes = new Uint8Array(readFileSync(join(fixDir, "far.h5")));
const farData = farBytes.subarray(far.data_at, far.data_at + far.nbytes);
const ff = await pkg.openUrl("http://far.invalid/far.h5", {
fetch: mockFetch(farBytes, { total: far.length, extra: [[far.far_at, farData]] }),
});
eq(Array.from((await ff.read("/x")).data), far.values, "data past 2 GiB");
eq(ff.stats().size, far.length, "size past 2 GiB");
// wasm32's reader holds file offsets in 32 bits: a file of 4 GiB or more
// is refused at open, with the limit in the message (past 2^53 - 1 bytes
// a JavaScript number is not even exact).
for (const total of [2 ** 32, 2 ** 40, 2 ** 53 + 2]) {
await fails(() => pkg.openUrl("http://huge.invalid/x.h5", { fetch: mockFetch(farBytes, { total }) }),
total > 2 ** 53 ? /2\^53 - 1/ : /4 GiB/, `length ${total}`);
}
}
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
// tells how many were.
function flood(total) {
let pulled = 0;
const body = new ReadableStream({
pull(c) {
if (pulled >= total) return c.close();
const n = Math.min(65536, total - pulled);
pulled += n;
c.enqueue(new Uint8Array(n));
},
}, { highWaterMark: 0 });
return { body, pulled: () => pulled };
}
// A 206 whose body runs past the range asked for is cut off as it arrives:
// the page never reads (or holds) more than it asked for, whatever the
// server sends. 64 MiB stands for "gigabytes".
async function floodTests() {
const FLOOD = 64 << 20;
// The probe: asked for the first 1 MiB.
let f = flood(FLOOD);
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
fetch: async () => new Response(f.body, { status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000" } }),
}), /the server sent more/, "flooded probe");
assert.ok(f.pulled() <= (1 << 20) + 65536, `probe: read ${f.pulled()} bytes`);
// A range request after a correct probe.
f = flood(FLOOD);
const file = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: async (url, init) => {
const range = new Headers(init.headers).get("Range");
if (range === "bytes=0-511") return fetch(url, init);
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
},
}).catch((e) => e);
// The open itself may need a second range: flooded either way.
const read = file instanceof Error ? Promise.reject(file) : file.read("/grid");
await fails(() => read, /the server sent more/, "flooded range");
assert.ok(f.pulled() <= 8 * 512 + 65536, `range: read ${f.pulled()} bytes`);
// A declared Content-Length past the range is refused before reading.
f = flood(FLOOD);
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
fetch: async () => new Response(f.body, {
status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000", "Content-Length": String(FLOOD) },
}),
}), /the server sent more/, "declared too long");
assert.ok(f.pulled() <= 65536, `declared: read ${f.pulled()} bytes`);
}
function hdf5Files(dir, out) {
for (const e of readdirSync(dir, { withFileTypes: true })) {
const p = join(dir, e.name);
let st;
try {
st = statSync(p);
} catch {
continue;
}
if (st.isDirectory()) hdf5Files(p, out);
else if (st.size <= 16 << 20) {
const b = readFileSync(p);
const sig = [0x89, 0x48, 0x44, 0x46, 0x0d, 0x0a, 0x1a, 0x0a];
for (let at = 0; at + 8 <= b.length; at = at === 0 ? 512 : at * 2) {
if (sig.every((x, i) => b[at + i] === x)) { out.push(p); break; }
}
}
}
return out;
}
function show(v) {
return JSON.stringify(v, (_, x) => {
if (typeof x === "bigint") return `${x}n`;
if (ArrayBuffer.isView(x)) return Array.from(x, (y) => (typeof y === "bigint" ? `${y}n` : Number.isNaN(y) ? "NaN" : y));
return x;
});
}
async function transcript(f) {
const out = [];
const call = async (what, fn) => {
try {
out.push(`${what}: ${show(await fn())}`);
} catch {
out.push(`${what}: Err`);
}
};
const todo = [["/", 0]];
while (todo.length && out.length < 3000) {
const [path, depth] = todo.pop();
let kind = null;
await call(`${path} kind`, async () => (kind = await f.kind(path)));
await call(`${path} attrs`, () => f.attrs(path));
if (kind === "group") {
let list = [];
await call(`${path} list`, async () => (list = await f.list(path)));
if (depth < 12) for (const c of list.reverse()) todo.push([lib.joinPath(path, c.name), depth + 1]);
} else if (kind === "dataset") {
let info = null;
await call(`${path} info`, async () => (info = await f.info(path)));
if (info && [...info.shape, ...info.elementShape].reduce((a, b) => a * b, 1) <= 1 << 20) {
await call(`${path} read`, () => f.read(path));
}
}
}
return out;
}
async function corpusTests(corpus) {
const files = hdf5Files(corpus, []).sort();
let opened = 0;
let bytes = 0;
let size = 0;
for (const p of files) {
const rel = relative(corpus, p).split("/").map(encodeURIComponent).join("/");
let local;
try {
local = pkg.open(new Uint8Array(readFileSync(p)));
} catch {
await fails(() => pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 }), /./, `${rel}: opens in neither`);
continue;
}
const remote = await pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 });
const want = await transcript(local);
const got = await transcript(remote);
// Errors are compared as errors: a malformed file can fail at another
// check, with another message, when read by ranges.
eq(got, want, `${rel}: transcript`);
opened++;
bytes += remote.stats().bytesFetched;
size += remote.stats().size;
local.free();
remote.free();
}
console.log(`corpus: ${opened} of ${files.length} files agree over HTTP (${bytes} of ${size} bytes fetched)`);
}
+21
View File
@@ -75,3 +75,24 @@ export function toRows(data, rows, cols, per = 1) {
export function perElement(elementShape) {
return elementShape.reduce((a, b) => a * b, 1);
}
/** A byte count for people: `1.5 KiB`, `200 MiB`. */
export function formatBytes(n) {
const units = ["B", "KiB", "MiB", "GiB", "TiB"];
let i = 0;
let v = n;
while (v >= 1024 && i < units.length - 1) {
v /= 1024;
i++;
}
const s = i === 0 ? String(v) : v < 10 ? String(Number(v.toFixed(1))) : String(Math.round(v));
return `${s} ${units[i]}`;
}
/** What reading a remote file has cost, from `RemoteFile.stats()`. */
export function formatStats({ lazy, requests, bytesFetched, size }) {
if (!lazy) return `downloaded whole (${formatBytes(size)}): the server does not support range requests`;
const pct = size > 0 ? (100 * bytesFetched) / size : 0;
return `${requests} request${requests === 1 ? "" : "s"}, ${formatBytes(bytesFetched)} of ` +
`${formatBytes(size)} fetched (${pct.toFixed(1)}%)`;
}
+8 -2
View File
@@ -131,14 +131,17 @@ run_step "cargo clippy (h5rs remote)" cargo clippy \
# js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C.
# clawhdf5-remote is checked by default (plain HTTP) and with its
# object-store feature, and h5rs with URL support (remote); the https
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in.
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. The
# Python bindings (clawhdf5-py, remote reads over plain HTTP) are checked too:
# their https/s3/gcs/azure features are opt-in for the same reason.
no_c_in_default_build() {
local entry crate features found=0
for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \
clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \
clawhdf5-tools \
clawhdf5-wasm \
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote; do
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote \
clawhdf5-py; do
crate=${entry%%:*}
features=()
[ "$entry" != "$crate" ] && features=(--features "${entry#*:}")
@@ -269,6 +272,9 @@ python_package() {
-i "$PYTHON" \
--out "$out/wheel" || return 1
"$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1
# The editing tests run `h5rs check` on every file they edit.
cargo build -q -p clawhdf5-tools || return 1
CLAWHDF5_H5RS="${CARGO_TARGET_DIR:-$root/target}/debug/h5rs" \
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
"$root/crates/clawhdf5-py/tests"
}