Merge pull request 'Lazy remote files in the browser (M4), SWMR reader (M5), Python remote reads and editing' (#19) from feat/p3-wasm-swmr-python into main
Reviewed-on: #19
This commit was merged in pull request #19.
This commit is contained in:
+302
@@ -2,6 +2,308 @@
|
||||
|
||||
## Unreleased
|
||||
|
||||
### Deterministic errors on damaged chunked datasets (2026-09-27)
|
||||
- A read through the file's chunk cache listed a damaged dataset's chunks in
|
||||
hash-map order, seeded per `File`, so two opens of the same file could
|
||||
report different failing chunks (`cve-2025-2310.h5`, and an intermittent
|
||||
failure of the storage and wasm corpus comparisons). The cache now keeps
|
||||
the chunk index's own order, as the uncached readers use: every path names
|
||||
the same first damaged chunk. Values of readable datasets were never
|
||||
affected.
|
||||
|
||||
### Range reads, milestone M5: reading files a SWMR writer is appending to (2026-09-27)
|
||||
Design: `docs/design/swmr.md`.
|
||||
- **Fix: files with the SWMR-write flag were bounded by a stale end of
|
||||
file** (since the end-of-file check of 2026-09-26, unreleased). A
|
||||
libhdf5 SWMR writer (h5py `f.swmr_mode = True`) does not keep
|
||||
the superblock's end-of-file address up to date; a copy h5py made of its
|
||||
own file mid-write records 715 in a 6 030-byte file.
|
||||
`Superblock::data_end` only ignored the recorded end when it lay past the
|
||||
end of the file, so such a file listed, but every chunked read failed
|
||||
("unexpected EOF: need 787 bytes, have 715") and `h5rs check` reported
|
||||
its chunk indexes past the end of the file. For a v3 superblock with the
|
||||
SWMR-write flag the data now ends at the end of the file, as libhdf5's
|
||||
SWMR reader reads it (it skips the end-of-allocation check). Every open
|
||||
path (`File::open`, `open_buffered`, `from_bytes`, `open_storage`,
|
||||
`MmapFile`, `LazyFile`, `h5rs`) reads such a copy now.
|
||||
- **`File::open_swmr(path)` / `File::open_storage_swmr(storage)`**: open a
|
||||
file a SWMR writer may still be appending to, as libhdf5's SWMR reader
|
||||
(h5py `File(path, "r", swmr=True)`) does. The file is read with
|
||||
positioned reads (the new `clawhdf5::FileStorage`: `pread`/`seek_read`,
|
||||
never mapped, `len()` the file's current length), reads are bounded by
|
||||
the file's length at the time of each read, and the chunk cache is not
|
||||
used (a cached chunk index would hide new chunks; a cached edge chunk
|
||||
would read as fill where the writer has since written). A file open for
|
||||
writing without SWMR is refused with `Error::Locked`, as libhdf5 refuses
|
||||
it.
|
||||
Only a file whose superblock has the SWMR-write flag when it is opened is
|
||||
read this way; any other file (its writer has closed it) reads exactly as
|
||||
`File::open` reads it, bounded by its recorded end of file, and
|
||||
`is_swmr_read()` is `false`.
|
||||
- **`Dataset::refresh()`** reads the dataset's object header again
|
||||
(`H5Drefresh`, h5py `Dataset.refresh()`), so `shape()` and later reads
|
||||
see the writer's appends; a handle keeps its extent until refreshed.
|
||||
- **Bounded retries.** On a SWMR-read file, an operation (open, lookups
|
||||
and listings, refresh, every dataset read including strings and
|
||||
variable-length data, attributes, `File::decode_*`) that fails with an error a concurrent write can cause —
|
||||
the failures libhdf5's SWMR reader retries: a checksum mismatch, a read
|
||||
past the file's current end, an object header prefix that does not
|
||||
decode — is run again from the start, up to
|
||||
`File::swmr_read_attempts()` times (`SWMR_READ_ATTEMPTS` = 100,
|
||||
libhdf5's default metadata read attempts for SWMR readers,
|
||||
`H5Pset_metadata_read_attempts`; `set_swmr_read_attempts` changes it),
|
||||
pausing 1 µs doubling to 10 ms between attempts. Results come only from
|
||||
an attempt in which every structure verified, so a torn read is at worst
|
||||
an error, never data. Any other error is returned at once.
|
||||
`File::swmr_retries()` counts the retries.
|
||||
`File::swmr_writer_active()` reads the superblock flags again to tell
|
||||
when the writer has closed the file.
|
||||
- Tests (`crates/clawhdf5/tests/swmr_interop.rs`): the mid-write copy
|
||||
(fixture `tests/fixtures/swmr_mid_write.h5`) through every open path and
|
||||
against h5py's SWMR reader; garbled reads (a storage that corrupts the
|
||||
next reads) retried, never returned, and given up after the attempts;
|
||||
every read path (listings, attributes compact and dense, strings,
|
||||
variable-length data; fixture `tests/fixtures/swmr_strings_attrs.h5`)
|
||||
with each of its reads failing once in turn returns the same result;
|
||||
and a live test: an h5py writer appends to a 1-D and a 2-D dataset with
|
||||
one unlimited dimension (Extensible Array index, one of them gzip) and a
|
||||
2-D dataset with two (v2 B-tree index) for 2 500 steps
|
||||
(`CLAWHDF5_SWMR_STEPS`), flushing after each, while two Rust reader
|
||||
threads refresh and read all three in a loop — every value must be the
|
||||
one the writer wrote at its position, extents never shrink — beside
|
||||
h5py's own SWMR reader doing the same checks; after the writer closes,
|
||||
the live handle and a new `File::open` read exactly what h5py reads.
|
||||
Run on tank, 2026-09-27, with h5py 3.16 / HDF5 2.0:
|
||||
`CLAWHDF5_REQUIRE_INTEROP=1 CLAWHDF5_PYTHON=.venv/bin/python cargo test
|
||||
-p clawhdf5 --test swmr_interop`; also at 20 000 steps in a release
|
||||
build. A test with the chunk cache left on in live mode fails it.
|
||||
- Not covered: SWMR writing, remote SWMR (`BlockCache` caches blocks and
|
||||
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
|
||||
groups or attributes (a SWMR writer cannot add objects or attributes).
|
||||
|
||||
### Correctness: edits planned from another file after a rename or `chdir` (2026-09-27)
|
||||
- **`FileEditor` planned each edit by re-opening its path but wrote
|
||||
through the file it held open** (fixed 2026-09-27; on main since PR #18,
|
||||
no release). When the path came to name another file between edits — a
|
||||
rename or replacement, or, for a relative path, a change of working
|
||||
directory — an edit was laid out from the other file's metadata and
|
||||
written into the held one, corrupting it (h5py: "invalid dataset size,
|
||||
likely file corruption"). The Python `'r+'` handle had the same flaw in
|
||||
its reads: it reopened the path after every edit, so reads came from the
|
||||
other file. The editor now plans every edit from the file it holds, and
|
||||
its path is canonicalised at open. New `FileEditor::reader()` opens the
|
||||
held file anew for reading (on Linux through `/proc/self/fd`, so it
|
||||
follows a renamed file; elsewhere by the path, refused when the path no
|
||||
longer names the held file), without sharing the editor's lock; the
|
||||
Python handle reads through it, and a `'w'` file is written at the
|
||||
absolute path it was opened with. Tests: `edit_tests.rs`'s
|
||||
`edits_go_to_the_file_held_not_the_path`; `test_edit.py`'s
|
||||
`test_relative_path_and_chdir` and `test_path_replaced_between_edits`
|
||||
(the review's repro).
|
||||
|
||||
### Correctness: zero extents in Fixed/Extensible Array chunk indexes (2026-09-27)
|
||||
- **A chunked dataset whose maximum (or, with none recorded, current)
|
||||
extent is 0 along a dimension made the reader divide by zero** (fixed
|
||||
2026-09-27): `h5rs check` panicked ("attempt to divide by zero",
|
||||
`chunk_grid.rs`) and the next `FileEditor::resize` failed with an
|
||||
internal error. The unfixed editor produced such files by resizing a
|
||||
clawhdf5-written dataset to a zero extent (12 of 30 extra random-edit
|
||||
seeds on clawhdf5-written files). Such an index has no slot for any
|
||||
chunk of the dataset, and `ChunkGrid::offsets` now says so instead of
|
||||
dividing by the zero stride. Tests: `chunk_grid`'s
|
||||
`zero_extent_has_no_chunks`, `edit_interop.rs`'s
|
||||
`zero_extent_resizes_without_a_recorded_maximum` (a file the unfixed
|
||||
editor left checks clean and resizes on; a 2.7.0-written file through
|
||||
zero extents checks clean at each step), and `test_edit.py`'s random
|
||||
edits on seeds 10 to 39 of a clawhdf5-written file.
|
||||
|
||||
### Correctness: resizing chunked datasets with no recorded maximum (2026-09-27)
|
||||
- **`FileEditor::resize` scrambled the values of a chunked dataset whose
|
||||
dataspace records no maximum dimensions when it shrank it** (fixed
|
||||
2026-09-27). The editor shipped on main in PR #18 (a4c2ace) and reached
|
||||
Python as `Dataset.resize` in `'r+'` files; no release has it. clawhdf5's
|
||||
writer stores such a dataspace for every chunked dataset created without
|
||||
a `maxshape`, with a Fixed Array (or Single Chunk) chunk index. The
|
||||
editor patched only the current dimensions, and with no maximum
|
||||
recorded the maximum is the current dimensions — which is also what the
|
||||
Fixed Array linearises chunks by — so a shrink moved every chunk after
|
||||
the first row: h5py, h5dump and our reader all read wrong values
|
||||
without complaint (20x20, chunks 6x6, resized to 15x15: row 6 read
|
||||
`0 0 0 0 0 0 120 ...`). After a shrink the dataset could not grow back
|
||||
either (`3 exceeds the maximum 0`). libhdf5 itself never writes such a
|
||||
dataspace (`H5S_set_extent_simple` records the maximum, equal to the
|
||||
dimensions when none is given); reading one, `H5S_extent_get_dims`
|
||||
reports the current dimensions as the maximum and `H5S_set_extent`
|
||||
checks against no maximum at all, so libhdf5's own `H5Dset_extent`
|
||||
scrambles such a file the same way (and lets it grow past its Fixed
|
||||
Array). The editor now records the maximum libhdf5 would have written
|
||||
— the dimensions before the first resize, the ones the index was built
|
||||
with — then changes the current ones (the dataspace message grows by one
|
||||
length per dimension and moves in the header when it has to). The
|
||||
dataset then shrinks, grows back to that extent and refuses more, as
|
||||
one libhdf5 wrote would. The writer (`FileBuilder`) now records the
|
||||
maximum of every chunked dataset too, as libhdf5 does, so h5py can
|
||||
resize what it writes (8 more bytes per dimension). Tests:
|
||||
`crates/clawhdf5/tests/edit_resize_interop.rs` (a 2.7.0-written fixture,
|
||||
new `FileBuilder` files and h5py files through shrinks, zero extents
|
||||
and growth, checked against a model with our reader and h5py; h5py
|
||||
resizing a `FileBuilder` file) and `test_edit.py`'s numpy-model checks.
|
||||
|
||||
### Python bindings: in-place editing (2026-09-27)
|
||||
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
|
||||
file for editing through `clawhdf5::FileEditor`, holding its exclusive
|
||||
lock until `close()`:
|
||||
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
|
||||
one increasing index list) and broadcasting (numpy's rules for slices
|
||||
and integers, allowing extra leading length-1 axes; the exact shape for
|
||||
an index list, a scalar only where h5py expands it). A numpy array is
|
||||
converted to the dataset's dtype as libhdf5 converts it in native byte
|
||||
order (integers saturate; floats are truncated toward zero and clipped;
|
||||
integers go into h5py's bool enum by value, as libhdf5 stores them);
|
||||
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
|
||||
does. NaN into an integer dataset is a `ValueError`.
|
||||
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
|
||||
and errors (`TypeError` for a dataset that is not chunked).
|
||||
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
|
||||
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
|
||||
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
|
||||
enum, complex as the `r`/`i` compound), except that `str` becomes
|
||||
fixed-length UTF-8.
|
||||
- Every edit is written and synced before it returns, then the file is
|
||||
reopened: datasets and attrs objects taken earlier see new shapes and
|
||||
attributes, and reads on other threads wait while an edit is written.
|
||||
- What the editor cannot do raises `NotImplementedError` and writes
|
||||
nothing (deleting attributes or objects, creating datasets or groups,
|
||||
compound fields by name, variable-length data, ...;
|
||||
`docs/known-issues.md`).
|
||||
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
|
||||
- Boolean-mask keys (`ds[mask]`, `ds[mask] = v`), which h5py supports,
|
||||
raise `NotImplementedError` (they raised `TypeError`).
|
||||
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
|
||||
by h5py and by clawhdf5, and both read back through h5py after every
|
||||
edit, on files h5py writes with `libver` earliest, v114 and latest and on
|
||||
a clawhdf5-written one: a fixed sequence over every chunk index kind,
|
||||
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
|
||||
and compound types, and 16 random sequences of 40 edits (writes,
|
||||
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
|
||||
and leave its file unchanged. A matrix of every numeric source dtype into
|
||||
every numeric dataset dtype at the edge values, dense attribute storage,
|
||||
locking, objects seeing each other's edits, readers racing a writer
|
||||
(never a partly written dataset). Every edited file must pass `h5dump`
|
||||
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
|
||||
a file opened `'r+'`.
|
||||
|
||||
### Python bindings: remote files (2026-09-27)
|
||||
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
|
||||
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
|
||||
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
|
||||
through the block cache, the whole read API (groups, attributes, every
|
||||
dataset type and index the local reader handles). A URL is any
|
||||
`scheme://…`; a remote file is read-only (another mode is a
|
||||
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
|
||||
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
|
||||
`allow_full_download`, `max_full_download`, `require_validator`,
|
||||
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
|
||||
cache's counters. The default wheel builds plain HTTP only (no C: rustls
|
||||
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
|
||||
check now covers `clawhdf5-py`.
|
||||
- **Every read parses through `File::storage()`** instead of
|
||||
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
|
||||
group listings, variable-length data through the global heap), inside one
|
||||
shared file handle that releases the GIL for all file access, not only
|
||||
dataset reads: a read waiting on the network lets other Python threads
|
||||
run. A failed read of the storage (a network error, a file changed on
|
||||
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
|
||||
data; `key in group` raises it instead of answering `False`.
|
||||
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
|
||||
blocks and 1 KiB blocks) against a range-capable `http.server` in the
|
||||
test process, plus `tests/test_remote.py`: request counts of a small
|
||||
read, cache hits, a server without `Range` support (refused, or a
|
||||
whole download when allowed), a file changed on the server, a server
|
||||
that hangs up, 16 threads on one remote file, and a thread that keeps
|
||||
running while a read waits on 0.2 s requests.
|
||||
|
||||
### Range reads, milestone M4: remote files in the browser (2026-09-27)
|
||||
- **`openUrl(url, opts)`** in `clawhdf5-wasm` opens an HDF5/NetCDF-4 file
|
||||
on a web server without downloading it: the returned `RemoteFile` has
|
||||
`H5File`'s methods (`kind`, `list`, `info`, `attrs`, `attrErrors`,
|
||||
`read`, `readHyperslab`), each returning a promise, and `stats()`
|
||||
(requests, bytes fetched, file size). Only the byte ranges a call needs
|
||||
are fetched, with `fetch` and `Range` headers, in 1 MiB blocks
|
||||
(`blockSize`) kept in a 64 MiB cache (`cacheSize`). `open(bytes)` is
|
||||
unchanged.
|
||||
- **How, on the browser's main thread:** the design's restartable
|
||||
"NeedBytes" mode (`clawhdf5_wasm::lazy::LazyStorage`). A call runs as a
|
||||
pass over the blocks fetched so far; a read that misses records the
|
||||
missing blocks and fails, the pass's result is dropped (even if a parser
|
||||
caught the error and carried on), the blocks are fetched and the pass is
|
||||
re-run. No Web Worker and no synchronous XHR (what h5wasm's lazy files
|
||||
need). No block is evicted while a call is in flight, so every call ends;
|
||||
the cache is trimmed between calls, raw-data blocks before metadata.
|
||||
`clawhdf5-remote`'s `BlockCache` is not reused: it fetches by blocking,
|
||||
and it evicts, and does not keep large reads, during a read, where a
|
||||
restartable pass needs every block it read to stay until it finishes.
|
||||
- **Every answer is checked** (`js/remote.js`): a `206` with exactly the
|
||||
bytes asked for (`Content-Range`, when the page can see it, and the body
|
||||
length), and the same `ETag`/`Last-Modified` and length as at open, or
|
||||
the call fails: never data from another file or offset. A server that
|
||||
answers `200` to a `Range` request is downloaded whole (up to
|
||||
`maxDownload`, 512 MiB) unless `fallback: "error"`. Other options:
|
||||
`headers`, `credentials`, `parallel` (6 requests at a time), `fetch`.
|
||||
- **The viewer** (`examples/wasm-viewer`) has a URL box, opens
|
||||
`?file=<url>` lazily, and shows the requests and bytes fetched.
|
||||
- Counted on tank, 2026-09-27 (`WASM_BIG_MB=200 bash
|
||||
examples/wasm-viewer/test/run.sh`, requests as `test/serve.py` counted
|
||||
them): in a 200 MB h5py file, listing the root, reading two small
|
||||
datasets, a group's attributes, the large dataset's shape and a
|
||||
10-value window of it (25 million values) took 5 requests and 6 MiB.
|
||||
With
|
||||
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, the 622 corpus files
|
||||
up to 16 MiB that open read over HTTP as from bytes: the same listings,
|
||||
attributes and values (datasets up to 2^20 values), and an error
|
||||
wherever bytes give one. Natively, `cargo test -p clawhdf5-wasm --test
|
||||
lazy` with the same variable compares 656 files (up to 64 MiB, error
|
||||
messages included, against the facade's range-storage path).
|
||||
- `Reader::open_storage` (the wasm crate's core over any `Storage`), and
|
||||
variable-length strings resolve through the file's storage rather than
|
||||
`File::as_bytes`.
|
||||
- The package is larger: the facade's `Storage` read path is now
|
||||
reachable from JavaScript (it was compiled out before), and the promise
|
||||
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
|
||||
build machine was shared); the viewer README's size table predates M4.
|
||||
- **Hardened after review (2026-09-27):**
|
||||
- Sizes a server or a dataset names are errors, never an abort of the
|
||||
wasm module (which took every open file on the page with it): a read
|
||||
past 2 GiB aborted in the lazy cache, reachable by a hostile server
|
||||
claiming a large file and a 2 GiB heap collection, and `read()` of a
|
||||
256 MiB `u8` dataset aborted widening it to 64 bits. New option
|
||||
`maxFetch` (512 MiB, at most 1 GiB): what one call may fetch, and the
|
||||
longest single read, refused before fetching. `read()` refuses a
|
||||
dataset that would take more than 1 GiB to decode, naming
|
||||
`readHyperslab`. A file of 4 GiB or more is refused at open (wasm32
|
||||
reads offsets as 32-bit); `maxDownload` is at most 1 GiB.
|
||||
- Response bodies are read as they arrive and cut off at the length
|
||||
asked for (`maxDownload` for a `200`): a `206` with a gigabyte body
|
||||
was buffered whole before its length was checked.
|
||||
- Listing a group reads every child's header, and every node of each
|
||||
level of the group's index, in one pass: 3000 datasets (h5py, 198 MB)
|
||||
listed in 6 passes and 73 requests at 1 MiB blocks instead of 185
|
||||
passes and 184 serial requests (`libver="latest"`: 9 passes instead of
|
||||
189). In `clawhdf5-format`, the B-tree v1/v2 collectors, the symbol
|
||||
table node loop and the dense-link loop read (without using) the
|
||||
siblings after the first that fails, then return that error: same
|
||||
results and errors, more reads only on failure (free in memory). The
|
||||
lazy cache no longer re-fetches a cached block to merge two requests.
|
||||
- `headers` may be a `Headers` instance or `[name, value]` pairs (a
|
||||
`Headers` was silently dropped); when one range request fails the
|
||||
others in flight are aborted; `parallel` must be an integer from 1 to
|
||||
1024.
|
||||
- Tests: `test/serve.py` serves ranges without exposed `Content-Range`/
|
||||
`ETag` (`/noexpose/`, and `/unexposed/` for a real cross-origin page
|
||||
in Chromium), so the HEAD-length path runs end to end; hostile and
|
||||
oversized files (`make_fixture.py`'s `write_limits`), flooding
|
||||
bodies, aborted siblings.
|
||||
|
||||
### Range reads, milestone M3: remote files (2026-09-26)
|
||||
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
||||
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
|
||||
|
||||
@@ -181,14 +181,19 @@ Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal F
|
||||
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
|
||||
file over HTTP with `File::open`.
|
||||
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
|
||||
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only, file held in memory;
|
||||
no Zstd/SZIP since they link C) and the `examples/wasm-viewer/` page.
|
||||
`examples/wasm-viewer/test/run.sh` builds the package (needs the
|
||||
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node
|
||||
and headless Chromium (a Playwright download in `~/.cache/ms-playwright`
|
||||
on tank); the CI container has neither, so CI runs the native
|
||||
`clawhdf5-wasm` `h5py_interop` test on the same fixture. Size numbers are
|
||||
in the example's README.
|
||||
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
|
||||
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
|
||||
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
|
||||
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
|
||||
call is re-run after each wave of misses; no block evicted while a call
|
||||
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
|
||||
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
|
||||
version) and tests it under Node and headless Chromium (a Playwright
|
||||
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
|
||||
(range server with request counts, 200 MB budget file); the CI container
|
||||
has neither, so CI runs the native `h5py_interop` and `lazy` tests
|
||||
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
|
||||
numbers in the example's README predate `openUrl`.
|
||||
- Python and Node.js bindings for cross-language use
|
||||
- NetCDF-4 compatibility for scientific data interop
|
||||
|
||||
|
||||
@@ -101,6 +101,10 @@ breaking change, are in [CHANGELOG.md](CHANGELOG.md).
|
||||
HTTP server (or in S3/GCS/Azure, opt-in) by range requests through a
|
||||
block cache, without downloading it; `h5rs` takes URLs with its `remote`
|
||||
feature. See [Reading remote files](#reading-remote-files).
|
||||
- `File::open_swmr` follows a file an h5py/libhdf5 SWMR writer is still
|
||||
appending to (`Dataset::refresh`, bounded retries); copies of such files
|
||||
taken mid-write read with every open path. See
|
||||
[Following a file a SWMR writer is appending to](#following-a-file-a-swmr-writer-is-appending-to).
|
||||
|
||||
**Tooling**
|
||||
- CI now runs the h5py/netCDF4 interop suites for real (they had been skipping
|
||||
@@ -504,6 +508,43 @@ built with `--features remote` takes the same URLs:
|
||||
`https://` is the `https` feature (rustls with ring, which compiles C).
|
||||
Limits are in [known issues](docs/known-issues.md).
|
||||
|
||||
### Following a file a SWMR writer is appending to
|
||||
|
||||
`File::open_swmr` reads a file that a libhdf5 writer in SWMR mode (h5py
|
||||
`f.swmr_mode = True`) is still appending to, as h5py's
|
||||
`File(path, "r", swmr=True)` does: `Dataset::refresh()` picks up the new
|
||||
extent, every read reads the chunk index as it is now, and a read that
|
||||
races the writer (a checksum that fails mid-flush) is retried, up to 100
|
||||
attempts as in libhdf5, and never returned torn.
|
||||
|
||||
```rust
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
let file = clawhdf5::File::open_swmr("live.h5")?;
|
||||
let mut ds = file.dataset("samples")?;
|
||||
let (mut seen, mut last_growth) = (0, Instant::now());
|
||||
// Stop when the writer closes the file, or when the dataset has not grown
|
||||
// for a minute: a writer that crashed or was killed never clears the
|
||||
// SWMR-write flag, so `swmr_writer_active()` alone can stay true forever.
|
||||
while file.swmr_writer_active()? && last_growth.elapsed() < Duration::from_secs(60) {
|
||||
ds.refresh()?;
|
||||
let n = ds.shape()?[0];
|
||||
if n > seen {
|
||||
// read rows seen..n ...
|
||||
(seen, last_growth) = (n, Instant::now());
|
||||
}
|
||||
std::thread::sleep(Duration::from_millis(100));
|
||||
}
|
||||
ds.refresh()?; // the final extent
|
||||
```
|
||||
|
||||
`swmr_writer_active()` reads the superblock's SWMR-write flag, which libhdf5
|
||||
clears only when the writer closes the file; a file whose writer died keeps
|
||||
it set (as the mid-write copy in `tests/fixtures/swmr_mid_write.h5` does),
|
||||
so a follower needs its own stop condition, like the idle timeout above.
|
||||
|
||||
Design and limits: [docs/design/swmr.md](docs/design/swmr.md).
|
||||
|
||||
### Python
|
||||
|
||||
`crates/clawhdf5-py` is a Python package (PyO3 + numpy) that reads HDF5 with
|
||||
@@ -533,8 +574,36 @@ with clawhdf5.File("data.h5", "r") as f:
|
||||
|
||||
records = f["table"] # compound -> numpy structured array
|
||||
ids = records["id"] # one field
|
||||
|
||||
# A file on a web server: range requests through a block cache, nothing
|
||||
# downloaded up front; the same read API. The GIL is released while waiting.
|
||||
with clawhdf5.File("http://data.example.org/run42.h5") as f:
|
||||
first = f["group/temperatures"][0]
|
||||
f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024,
|
||||
headers={"Authorization": "Bearer ..."})
|
||||
```
|
||||
|
||||
An existing file opened with `"r+"` is edited in place (through
|
||||
`clawhdf5::FileEditor`), with h5py's indexing, broadcasting and numeric
|
||||
conversion; each edit is on disk when the statement returns:
|
||||
|
||||
```python
|
||||
with clawhdf5.File("data.h5", "r+") as f:
|
||||
f["group/temperatures"][100:200, ::4] = 0.0
|
||||
f["series"].resize(5000, axis=0) # chunked datasets, within maxshape
|
||||
f["series"][4000:] = new_values
|
||||
f["group"].attrs["calibrated"] = True
|
||||
```
|
||||
|
||||
Creating or deleting datasets, groups and attributes in an existing file is
|
||||
not supported (`NotImplementedError`); limits are in
|
||||
[known issues](docs/known-issues.md).
|
||||
|
||||
The default build reads `http://` URLs only; build with
|
||||
`maturin develop --release --features https` (rustls with ring, which
|
||||
compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for
|
||||
object-store URLs.
|
||||
|
||||
Reads cover integers and IEEE floats of every width in either byte order,
|
||||
`bool`, enums, complex, fixed and variable-length strings, variable-length
|
||||
sequences, opaque, HDF5 array types and compounds; other types (references,
|
||||
@@ -549,7 +618,8 @@ non-default fill value (`docs/known-issues.md`). An index list is read one
|
||||
group of neighbouring chunks at a time.
|
||||
Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`)
|
||||
covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests
|
||||
in `crates/clawhdf5-py/tests` compare every read with h5py; run them with
|
||||
in `crates/clawhdf5-py/tests` compare every read and every in-place edit
|
||||
with h5py; run them with
|
||||
`pip install pytest h5py && pytest crates/clawhdf5-py/tests`.
|
||||
|
||||
### Agent Memory
|
||||
@@ -757,7 +827,7 @@ clawhdf5 workspace (19 crates, ~86K lines of Rust in src/, ~104K with tests
|
||||
├── Bindings
|
||||
│ ├── clawhdf5-py — Python (PyO3)
|
||||
│ ├── clawhdf5-napi — Node.js (napi-rs)
|
||||
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only)
|
||||
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only; remote files by HTTP range requests)
|
||||
│
|
||||
└── Tooling
|
||||
├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check
|
||||
|
||||
@@ -576,7 +576,14 @@ pub fn find_attribute_in_file(
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
) -> Result<Option<AttributeMessage>, FormatError> {
|
||||
find_attribute_core(file_data, header, name, offset_size, length_size)
|
||||
find_attribute_core(
|
||||
file_data,
|
||||
header,
|
||||
name,
|
||||
offset_size,
|
||||
length_size,
|
||||
&mut Vec::new(),
|
||||
)
|
||||
}
|
||||
|
||||
/// [`find_attribute_in_file`] over any [`Storage`] (see
|
||||
@@ -594,16 +601,48 @@ pub fn find_attribute_in<S: Storage + ?Sized>(
|
||||
) -> Result<Option<AttributeMessage>, FormatError> {
|
||||
match file_data.as_contiguous() {
|
||||
Some(all) => find_attribute_in_file(all, header, name, offset_size, length_size),
|
||||
None => find_attribute_core(file_data, header, name, offset_size, length_size),
|
||||
None => find_attribute_core(
|
||||
file_data,
|
||||
header,
|
||||
name,
|
||||
offset_size,
|
||||
length_size,
|
||||
&mut Vec::new(),
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
/// [`find_attribute_in`], also returning the errors of the attributes it
|
||||
/// could not read on the way (which it leaves out rather than failing
|
||||
/// the call): the attribute asked for may be one of them. A reader of a
|
||||
/// file that is being written uses them to tell a read that raced the
|
||||
/// writer from an absent attribute.
|
||||
pub fn find_attribute_reporting_in<S: Storage + ?Sized>(
|
||||
file_data: &S,
|
||||
header: &ObjectHeader,
|
||||
name: &str,
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
) -> Result<(Option<AttributeMessage>, Vec<FormatError>), FormatError> {
|
||||
let mut errors = Vec::new();
|
||||
let found = find_attribute_core(
|
||||
file_data,
|
||||
header,
|
||||
name,
|
||||
offset_size,
|
||||
length_size,
|
||||
&mut errors,
|
||||
)?;
|
||||
Ok((found, errors))
|
||||
}
|
||||
|
||||
fn find_attribute_core<S: Storage + ?Sized>(
|
||||
file_data: &S,
|
||||
header: &ObjectHeader,
|
||||
name: &str,
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
errors: &mut Vec<FormatError>,
|
||||
) -> Result<Option<AttributeMessage>, FormatError> {
|
||||
let attr_info = find_attribute_info(header, offset_size)?;
|
||||
let dense = attr_info
|
||||
@@ -612,12 +651,10 @@ fn find_attribute_core<S: Storage + ?Sized>(
|
||||
let Some((fh_addr, btree_addr)) = dense else {
|
||||
// Compact only (or dense storage without a name index, which a
|
||||
// listing reports): as a listing finds it.
|
||||
return Ok(
|
||||
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?
|
||||
.0
|
||||
.into_iter()
|
||||
.find(|a| a.name == name),
|
||||
);
|
||||
let (attrs, errs) =
|
||||
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?;
|
||||
errors.extend(errs);
|
||||
return Ok(attrs.into_iter().find(|a| a.name == name));
|
||||
};
|
||||
let btree_hdr = BTreeV2Header::parse_in(
|
||||
file_data,
|
||||
@@ -627,12 +664,10 @@ fn find_attribute_core<S: Storage + ?Sized>(
|
||||
)?;
|
||||
let fh = FractalHeapHeader::parse_in(file_data, fh_addr, offset_size, length_size)?;
|
||||
if btree_hdr.tree_type != ATTRIBUTE_NAME_INDEX || btree_hdr.record_size < 4 {
|
||||
return Ok(
|
||||
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?
|
||||
.0
|
||||
.into_iter()
|
||||
.find(|a| a.name == name),
|
||||
);
|
||||
let (attrs, errs) =
|
||||
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?;
|
||||
errors.extend(errs);
|
||||
return Ok(attrs.into_iter().find(|a| a.name == name));
|
||||
}
|
||||
|
||||
// A listing has the compact attributes first.
|
||||
@@ -671,10 +706,10 @@ fn find_attribute_core<S: Storage + ?Sized>(
|
||||
AttributeMessage::parse_in_storage(&d, file_data, offset_size, length_size)
|
||||
});
|
||||
// One that cannot be read is left out, as from a listing.
|
||||
if let Ok(attr) = attr
|
||||
&& attr.name == name
|
||||
{
|
||||
return Ok(Some(attr));
|
||||
match attr {
|
||||
Ok(attr) if attr.name == name => return Ok(Some(attr)),
|
||||
Ok(_) => {}
|
||||
Err(e) => errors.push(e),
|
||||
}
|
||||
}
|
||||
Ok(None)
|
||||
|
||||
@@ -199,19 +199,32 @@ fn collect_symbol_table_nodes_inner<S: Storage + ?Sized>(
|
||||
// Leaf: children are SNOD addresses
|
||||
Ok(node.children)
|
||||
} else {
|
||||
// Internal: recurse into children
|
||||
// Internal: recurse into children. After the first child that
|
||||
// fails, the others are only read (as `storage::touch` does), not
|
||||
// descended into; that error is returned.
|
||||
let mut result = Vec::new();
|
||||
let mut failed = None;
|
||||
for &child_addr in &node.children {
|
||||
let child_snods = collect_symbol_table_nodes_inner(
|
||||
if failed.is_some() {
|
||||
// Parsing reads the node's header, then its body.
|
||||
let _ = BTreeV1Node::parse_in(file, child_addr, offset_size, length_size);
|
||||
continue;
|
||||
}
|
||||
match collect_symbol_table_nodes_inner(
|
||||
file,
|
||||
child_addr,
|
||||
offset_size,
|
||||
length_size,
|
||||
depth + 1,
|
||||
)?;
|
||||
result.extend(child_snods);
|
||||
) {
|
||||
Ok(child_snods) => result.extend(child_snods),
|
||||
Err(e) => failed = Some(e),
|
||||
}
|
||||
}
|
||||
match failed {
|
||||
Some(e) => Err(e),
|
||||
None => Ok(result),
|
||||
}
|
||||
Ok(result)
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -504,45 +504,63 @@ fn collect_internal_records<S: Storage + ?Sized>(
|
||||
|
||||
// Interleave: child[0], record[0], child[1], record[1], ..., child[nr]
|
||||
// We collect child[0] records, then record[0], then child[1], etc.
|
||||
// After the first child that fails, the others are only touched (see
|
||||
// `storage::touch`); that error is returned.
|
||||
let mut failed = None;
|
||||
for (i, &(child_addr, child_nrec)) in node.children.iter().enumerate() {
|
||||
if child_depth == 0 {
|
||||
// Before parsing, so a refused tree is not also a large allocation.
|
||||
spend(budget, usize::from(child_nrec))?;
|
||||
let leaf_recs = parse_leaf_records(
|
||||
file,
|
||||
to_usize(child_addr)?,
|
||||
child_nrec,
|
||||
record_size,
|
||||
node_size,
|
||||
)?;
|
||||
out.extend(leaf_recs);
|
||||
} else {
|
||||
collect_internal_records(
|
||||
file,
|
||||
to_usize(child_addr)?,
|
||||
child_nrec,
|
||||
child_depth,
|
||||
record_size,
|
||||
node_size,
|
||||
offset_size,
|
||||
length_size,
|
||||
max_leaf_nrec,
|
||||
budget,
|
||||
out,
|
||||
)?;
|
||||
if failed.is_some() {
|
||||
let len = usize::try_from(node_size)
|
||||
.unwrap_or(usize::MAX)
|
||||
.min(1 << 16);
|
||||
crate::storage::touch(file, child_addr, len);
|
||||
continue;
|
||||
}
|
||||
if let Err(e) = (|| -> Result<(), FormatError> {
|
||||
if child_depth == 0 {
|
||||
// Before parsing, so a refused tree is not also a large allocation.
|
||||
spend(budget, usize::from(child_nrec))?;
|
||||
let leaf_recs = parse_leaf_records(
|
||||
file,
|
||||
to_usize(child_addr)?,
|
||||
child_nrec,
|
||||
record_size,
|
||||
node_size,
|
||||
)?;
|
||||
out.extend(leaf_recs);
|
||||
} else {
|
||||
collect_internal_records(
|
||||
file,
|
||||
to_usize(child_addr)?,
|
||||
child_nrec,
|
||||
child_depth,
|
||||
record_size,
|
||||
node_size,
|
||||
offset_size,
|
||||
length_size,
|
||||
max_leaf_nrec,
|
||||
budget,
|
||||
out,
|
||||
)?;
|
||||
}
|
||||
|
||||
// Add record[i] (except after the last child)
|
||||
if i < nr {
|
||||
let data = node.record(i, rs)?;
|
||||
spend(budget, 1)?;
|
||||
out.push(BTreeV2Record {
|
||||
data: data.to_vec(),
|
||||
});
|
||||
// Add record[i] (except after the last child)
|
||||
if i < nr {
|
||||
let data = node.record(i, rs)?;
|
||||
spend(budget, 1)?;
|
||||
out.push(BTreeV2Record {
|
||||
data: data.to_vec(),
|
||||
});
|
||||
}
|
||||
Ok(())
|
||||
})() {
|
||||
failed = Some(e);
|
||||
}
|
||||
}
|
||||
|
||||
Ok(())
|
||||
match failed {
|
||||
Some(e) => Err(e),
|
||||
None => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
/// The records of a B-tree v2 that fall in one key range, found by
|
||||
|
||||
@@ -262,6 +262,11 @@ struct CachedChunk {
|
||||
struct DatasetEntry {
|
||||
/// Chunk coordinate -> ChunkInfo (offset + size in file).
|
||||
index: Option<Arc<HashMap<ChunkCoord, ChunkInfo>>>,
|
||||
/// The same chunks in the order the chunk index lists them: what
|
||||
/// [`ChunkCache::chunks_for`] returns, so a cached read walks (and, on a
|
||||
/// damaged file, fails at) the chunks in the same order as an uncached
|
||||
/// one, rather than in hash-map order.
|
||||
ordered: Option<Arc<Vec<ChunkInfo>>>,
|
||||
/// Pre-built chunk index for O(1) coordinate lookups.
|
||||
chunk_index: Option<Arc<ChunkIndex>>,
|
||||
/// Pre-computed chunk layout for fast assembly.
|
||||
@@ -274,6 +279,7 @@ struct DatasetEntry {
|
||||
impl DatasetEntry {
|
||||
fn weight(&self) -> usize {
|
||||
self.index.as_ref().map_or(0, |m| m.len())
|
||||
+ self.ordered.as_ref().map_or(0, |o| o.len())
|
||||
+ self.chunk_index.as_ref().map_or(0, |c| c.num_chunks())
|
||||
}
|
||||
}
|
||||
@@ -562,11 +568,22 @@ impl ChunkCache {
|
||||
rank: usize,
|
||||
build: impl FnOnce() -> Result<Vec<ChunkInfo>, E>,
|
||||
) -> Result<Vec<ChunkInfo>, E> {
|
||||
Ok(self
|
||||
.index_for(addr, rank, build)?
|
||||
.values()
|
||||
.cloned()
|
||||
.collect())
|
||||
if let Some(ordered) = self.lock().touch(addr).ordered.clone() {
|
||||
return Ok(ordered.as_ref().clone());
|
||||
}
|
||||
let chunks = build()?;
|
||||
let map: HashMap<ChunkCoord, ChunkInfo> = chunks
|
||||
.iter()
|
||||
.map(|ci| (ci.offsets.iter().take(rank).copied().collect(), ci.clone()))
|
||||
.collect();
|
||||
let mut inner = self.lock();
|
||||
let entry = inner.touch(addr);
|
||||
// Another thread may have built this dataset's index meanwhile: keep
|
||||
// the first one, so every reader sees the same order.
|
||||
let ordered = Arc::clone(entry.ordered.get_or_insert_with(|| Arc::new(chunks)));
|
||||
entry.index.get_or_insert_with(|| Arc::new(map));
|
||||
inner.trim_datasets(addr);
|
||||
Ok(ordered.as_ref().clone())
|
||||
}
|
||||
|
||||
fn index_for<E>(
|
||||
@@ -580,11 +597,14 @@ impl ChunkCache {
|
||||
}
|
||||
let chunks = build()?;
|
||||
let map: HashMap<ChunkCoord, ChunkInfo> = chunks
|
||||
.into_iter()
|
||||
.map(|ci| (ci.offsets.iter().take(rank).copied().collect(), ci))
|
||||
.iter()
|
||||
.map(|ci| (ci.offsets.iter().take(rank).copied().collect(), ci.clone()))
|
||||
.collect();
|
||||
let mut inner = self.lock();
|
||||
let entry = inner.touch(addr);
|
||||
if entry.index.is_none() {
|
||||
entry.ordered = Some(Arc::new(chunks));
|
||||
}
|
||||
let index = Arc::clone(entry.index.get_or_insert_with(|| Arc::new(map)));
|
||||
inner.trim_datasets(addr);
|
||||
Ok(index)
|
||||
|
||||
@@ -135,6 +135,12 @@ impl ChunkGrid {
|
||||
let mut rem = index;
|
||||
for p in 0..rank {
|
||||
let d = self.order[p];
|
||||
// A zero stride: a later dimension has no chunks (its maximum,
|
||||
// or with none recorded its current extent, is 0), so no slot of
|
||||
// the index is a chunk of the dataset.
|
||||
if self.down[p] == 0 {
|
||||
return None;
|
||||
}
|
||||
let scaled = rem / self.down[p];
|
||||
rem %= self.down[p];
|
||||
if scaled >= self.cur_chunks[d] {
|
||||
@@ -193,6 +199,28 @@ mod tests {
|
||||
assert_eq!(g.offsets(11), Some(vec![2, 3]));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn zero_extent_has_no_chunks() {
|
||||
// No maximum recorded and a zero current dimension: every stride
|
||||
// before it is 0 (this divided by zero).
|
||||
let g = ChunkGrid::fixed_array(&[1, 0], None, &[6, 6]).unwrap();
|
||||
for i in 0..16 {
|
||||
assert_eq!(g.offsets(i), None);
|
||||
}
|
||||
let g = ChunkGrid::fixed_array(&[0, 0, 3], Some(&[4, 0, 3]), &[2, 2, 3]).unwrap();
|
||||
for i in 0..16 {
|
||||
assert_eq!(g.offsets(i), None);
|
||||
}
|
||||
let g = ChunkGrid::extensible_array(&[0, 5], Some(&[u64::MAX, 0]), &[2, 2]).unwrap();
|
||||
for i in 0..16 {
|
||||
assert_eq!(g.offsets(i), None);
|
||||
}
|
||||
// A zero last dimension leaves the other strides alone.
|
||||
let g = ChunkGrid::fixed_array(&[4, 0], Some(&[4, 6]), &[2, 3]).unwrap();
|
||||
assert_eq!(g.offsets(0), None);
|
||||
assert_eq!(g.linear_index(&[1, 1]), 3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rejects_two_unlimited_dims_after_the_first() {
|
||||
assert!(ChunkGrid::fixed_array(&[4, 6], Some(&[u64::MAX, u64::MAX]), &[2, 3]).is_err());
|
||||
|
||||
@@ -130,6 +130,22 @@ pub(crate) fn build_chunked_dataset_oh(
|
||||
fill_message: &[u8],
|
||||
refcount: u32,
|
||||
) -> Result<Vec<u8>, FormatError> {
|
||||
// libhdf5 records every simple dataspace's maximum dimensions (the
|
||||
// dimensions themselves when none are given, `H5S_set_extent_simple`).
|
||||
// Without them libhdf5 takes the maximum to be the current dimensions,
|
||||
// so a resize by libhdf5 (h5py's `Dataset.resize`) would also change
|
||||
// the maximum a Fixed Array chunk index is laid out by, and move every
|
||||
// chunk already written.
|
||||
let recorded;
|
||||
let ds = if ds.space_type == DataspaceType::Simple && ds.max_dimensions.is_none() {
|
||||
recorded = Dataspace {
|
||||
max_dimensions: Some(ds.dimensions.clone()),
|
||||
..ds.clone()
|
||||
};
|
||||
&recorded
|
||||
} else {
|
||||
ds
|
||||
};
|
||||
let mut w = ObjectHeaderWriter::new();
|
||||
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
|
||||
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
|
||||
|
||||
@@ -79,27 +79,51 @@ pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
|
||||
length_size,
|
||||
)?;
|
||||
|
||||
// The names are read one by one from the heap's data segment; read
|
||||
// (up to 1 MiB of) it first, so a storage that records what it lacks
|
||||
// asks for it at once (see `storage::touch`).
|
||||
if !snod_addrs.is_empty() {
|
||||
let len = usize::try_from(heap.data_segment_size).map_or(1 << 20, |n| n.min(1 << 20));
|
||||
crate::storage::touch(file_data, heap.data_segment_address, len);
|
||||
}
|
||||
|
||||
let mut entries = Vec::new();
|
||||
let mut heap_checked = false;
|
||||
// After the first node that fails, the others are only read (as
|
||||
// `storage::touch` does); that error is returned.
|
||||
let mut failed = None;
|
||||
for snod_addr in snod_addrs {
|
||||
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
|
||||
for entry in &snod.entries {
|
||||
// Like libhdf5, look at the heap's free list only once a name is
|
||||
// needed: an empty group with a damaged heap still lists.
|
||||
if !heap_checked {
|
||||
heap.validate_free_list_in(file_data, length_size)?;
|
||||
heap_checked = true;
|
||||
if failed.is_some() {
|
||||
let _ = SymbolTableNode::parse_in(file_data, snod_addr, offset_size);
|
||||
continue;
|
||||
}
|
||||
let mut node = || -> Result<(), FormatError> {
|
||||
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
|
||||
for entry in &snod.entries {
|
||||
// Like libhdf5, look at the heap's free list only once a name
|
||||
// is needed: an empty group with a damaged heap still lists.
|
||||
if !heap_checked {
|
||||
heap.validate_free_list_in(file_data, length_size)?;
|
||||
heap_checked = true;
|
||||
}
|
||||
let name = heap.read_string_in(file_data, entry.link_name_offset)?;
|
||||
entries.push(GroupEntry {
|
||||
name,
|
||||
object_header_address: entry.object_header_address,
|
||||
cache_type: entry.cache_type,
|
||||
});
|
||||
}
|
||||
let name = heap.read_string_in(file_data, entry.link_name_offset)?;
|
||||
entries.push(GroupEntry {
|
||||
name,
|
||||
object_header_address: entry.object_header_address,
|
||||
cache_type: entry.cache_type,
|
||||
});
|
||||
Ok(())
|
||||
};
|
||||
if let Err(e) = node() {
|
||||
failed = Some(e);
|
||||
}
|
||||
}
|
||||
|
||||
Ok(entries)
|
||||
match failed {
|
||||
Some(e) => Err(e),
|
||||
None => Ok(entries),
|
||||
}
|
||||
}
|
||||
|
||||
/// Symbol table cache type for a soft link: the scratch pad's first four bytes
|
||||
|
||||
@@ -126,6 +126,9 @@ fn for_each_dense_link<S: Storage + ?Sized>(
|
||||
)?;
|
||||
let records = collect_btree_v2_records_in(file_data, &btree_hdr, offset_size, length_size)?;
|
||||
|
||||
// After the first link that fails, the others are only read, not
|
||||
// visited (a touch, see `storage::touch`); that error is returned.
|
||||
let mut failed = None;
|
||||
for record in &records {
|
||||
// For type 5 (name index): hash(4) + heap_id(heap_id_length)
|
||||
// For type 6 (creation order): creation_order(8) + heap_id(heap_id_length)
|
||||
@@ -141,12 +144,20 @@ fn for_each_dense_link<S: Storage + ?Sized>(
|
||||
let id_bytes = &record.data[id_offset..id_offset + fh.heap_id_length as usize];
|
||||
|
||||
// Read managed object from fractal heap
|
||||
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size)?;
|
||||
if let Some(link) = parse_link(&link_data, offset_size)? {
|
||||
visit(link);
|
||||
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size);
|
||||
if failed.is_some() {
|
||||
continue;
|
||||
}
|
||||
match link_data.and_then(|d| parse_link(&d, offset_size)) {
|
||||
Ok(Some(link)) => visit(link),
|
||||
Ok(None) => {}
|
||||
Err(e) => failed = Some(e),
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
match failed {
|
||||
Some(e) => Err(e),
|
||||
None => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
/// Resolve entries from dense storage (fractal heap + B-tree v2).
|
||||
|
||||
@@ -202,6 +202,19 @@ pub(crate) fn len_usize<S: Storage + ?Sized>(file: &S) -> usize {
|
||||
usize::try_from(file.len()).unwrap_or(usize::MAX)
|
||||
}
|
||||
|
||||
/// Read `len` bytes at `offset` and drop them, ignoring any error.
|
||||
///
|
||||
/// For a traversal that has failed on one sibling (a B-tree child, a
|
||||
/// symbol table node, a heap object) and would stop there: it first
|
||||
/// touches the siblings it did not get to, so a storage that records what
|
||||
/// it lacks — the browser's restartable reader, which fetches over the
|
||||
/// network between attempts — learns about all of them in one attempt
|
||||
/// instead of one per attempt. Results and errors are unchanged (the first
|
||||
/// error is still the one returned); an in-memory read is free.
|
||||
pub fn touch<S: Storage + ?Sized>(file: &S, offset: u64, len: usize) {
|
||||
let _ = file.read_at(offset, len);
|
||||
}
|
||||
|
||||
/// Bytes `[offset, offset + len)`, all of them.
|
||||
///
|
||||
/// A range that runs past the end of the storage is
|
||||
|
||||
@@ -116,22 +116,27 @@ impl Superblock {
|
||||
/// [`FormatError::TruncatedFile`]. Bytes past that address are not part
|
||||
/// of the file: libhdf5 fails any read of them ("addr overflow" /
|
||||
/// "address plus size exceeds file eoa"), so a reader should parse only
|
||||
/// the data up to the returned end. As libhdf5 does for a SWMR reader,
|
||||
/// the check is skipped for a version-3 superblock whose writer is still
|
||||
/// writing it in SWMR mode (it extends the file as it goes); the data
|
||||
/// then ends at the end of the file.
|
||||
/// the data up to the returned end.
|
||||
///
|
||||
/// A version-3 superblock with the SWMR-write flag set belongs to a file
|
||||
/// a SWMR writer has open (or had, and did not close). That writer does
|
||||
/// not keep the recorded end of file up to date — a copy taken mid-write
|
||||
/// can record an end of a few hundred bytes in a file of tens of
|
||||
/// kilobytes — and libhdf5's SWMR reader skips its end-of-allocation
|
||||
/// check for every read (`H5FD_read`). For such a superblock the data
|
||||
/// ends at the end of the file, whatever end it records.
|
||||
///
|
||||
/// When the superblock's recorded base address differs from where the
|
||||
/// superblock actually is (a user block added or removed after the file
|
||||
/// was written), libhdf5 moves the recorded end of file by the same
|
||||
/// amount, and so does this.
|
||||
pub fn data_end(&self, user_block: u64, file_len: u64) -> Result<u64, FormatError> {
|
||||
if self.version >= 3 && self.is_swmr_write() {
|
||||
return Ok(file_len.saturating_sub(user_block));
|
||||
}
|
||||
let eof =
|
||||
i128::from(self.eof_address) - i128::from(self.base_address) + i128::from(user_block);
|
||||
if eof < 0 || eof > i128::from(file_len) {
|
||||
if self.version >= 3 && self.is_swmr_write() {
|
||||
return Ok(file_len.saturating_sub(user_block));
|
||||
}
|
||||
return Err(FormatError::TruncatedFile {
|
||||
stored_eof: u64::try_from(eof).unwrap_or(self.eof_address),
|
||||
actual_len: file_len,
|
||||
@@ -617,6 +622,13 @@ mod tests {
|
||||
let mut swmr = Superblock::parse(&build_v2_bytes(8, 3), 0).unwrap();
|
||||
swmr.consistency_flags = swmr_flags::WRITE_ACCESS | swmr_flags::SWMR_WRITE;
|
||||
assert_eq!(swmr.data_end(0, 1000), Ok(1000));
|
||||
// ... nor bounded by its recorded end, which the writer does not
|
||||
// keep up to date (2048 here).
|
||||
assert_eq!(swmr.data_end(0, 17_857), Ok(17_857));
|
||||
assert_eq!(swmr.data_end(512, 17_857), Ok(17_345));
|
||||
// Without the SWMR-write flag the recorded end bounds the data.
|
||||
swmr.consistency_flags = swmr_flags::WRITE_ACCESS;
|
||||
assert_eq!(swmr.data_end(0, 17_857), Ok(2048));
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
@@ -39,6 +39,23 @@ impl MmapReader {
|
||||
Ok(Self { _file: file, mmap })
|
||||
}
|
||||
|
||||
/// Memory-map a file that is already open (for reading).
|
||||
///
|
||||
/// The mapping references `file`'s open file description for as long as
|
||||
/// it lives, so a `flock` taken through that description (or a
|
||||
/// `try_clone` of it) is held until the reader is dropped.
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// The same contract as [`open`](Self::open): the file must not be
|
||||
/// modified while the mapping is active.
|
||||
pub fn from_file(file: fs::File) -> io::Result<Self> {
|
||||
// SAFETY: a read-only mapping; the caller keeps the file unmodified
|
||||
// while it is alive.
|
||||
let mmap = unsafe { Mmap::map(&file)? };
|
||||
Ok(Self { _file: file, mmap })
|
||||
}
|
||||
|
||||
/// Zero-copy access to the entire file contents.
|
||||
pub fn as_bytes(&self) -> &[u8] {
|
||||
&self.mmap
|
||||
|
||||
@@ -17,11 +17,20 @@ crate-type = ["cdylib", "rlib"]
|
||||
[dependencies]
|
||||
clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" }
|
||||
clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
|
||||
# Remote files (`clawhdf5.File(url)`): plain HTTP by default, which builds
|
||||
# no C. HTTPS and the object stores are opt-in features below.
|
||||
clawhdf5-remote = { path = "../clawhdf5-remote", version = "2.7.0" }
|
||||
pyo3 = "0.29"
|
||||
numpy = "0.29"
|
||||
|
||||
[features]
|
||||
extension-module = ["pyo3/extension-module"]
|
||||
# https:// URLs (rustls with ring, which compiles C and assembly).
|
||||
https = ["clawhdf5-remote/https"]
|
||||
# s3://, gs://, az:// URLs (object_store; its cloud clients build aws-lc-rs, C).
|
||||
s3 = ["clawhdf5-remote/s3"]
|
||||
gcs = ["clawhdf5-remote/gcs"]
|
||||
azure = ["clawhdf5-remote/azure"]
|
||||
|
||||
[package.metadata.docs.rs]
|
||||
features = []
|
||||
|
||||
@@ -43,8 +43,9 @@ with clawhdf5.File("data.h5", "r") as f:
|
||||
Other types raise `TypeError`.
|
||||
- Keys are h5py's: integers, slices with a positive step, `...`, one
|
||||
increasing list of integers, compound field names. Each maps onto a
|
||||
hyperslab selection. `None`, negative steps and boolean masks are refused
|
||||
with h5py's errors.
|
||||
hyperslab selection. `None` and negative steps are refused
|
||||
with h5py's errors; boolean masks (which h5py supports) raise
|
||||
`NotImplementedError`, for reads and writes.
|
||||
- What is read from the file: a selection whose bounding box covers at
|
||||
most half the dataset decodes only the chunks (or contiguous rows) the box
|
||||
overlaps. The library decodes the whole dataset for a larger box
|
||||
@@ -61,6 +62,36 @@ with clawhdf5.File("data.h5", "r") as f:
|
||||
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
|
||||
dataspace (h5py's `Empty`).
|
||||
|
||||
## Remote files
|
||||
|
||||
A URL instead of a path reads the file where it is, through
|
||||
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
|
||||
64 MiB budget by default), fetching only the blocks a read needs. The whole
|
||||
read API works the same, and the GIL is released while waiting on the
|
||||
network.
|
||||
|
||||
```python
|
||||
f = clawhdf5.File("http://host/data.h5") # default options
|
||||
f = clawhdf5.File.open_url(
|
||||
"http://host/data.h5",
|
||||
block_size=256 * 1024, cache_size=128 << 20, # the block cache
|
||||
headers={"Authorization": "Bearer ..."}, # sent to this origin only
|
||||
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
|
||||
allow_full_download=False, # a server without Range support: refuse
|
||||
require_validator=False, # refuse servers without ETag/Last-Modified
|
||||
)
|
||||
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
|
||||
```
|
||||
|
||||
- The file is pinned when opened (ETag or Last-Modified, and length): if it
|
||||
changes on the server, reads raise `OSError` instead of mixing versions.
|
||||
Network failures are `OSError` too.
|
||||
- Remote files are read-only.
|
||||
- Schemes: the default build (no C) reads `http://`. `https://` needs
|
||||
`maturin develop --release --features https` (rustls with ring, which
|
||||
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
|
||||
`azure` features (credentials from the environment; aws-lc-rs, C).
|
||||
|
||||
## Writing
|
||||
|
||||
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
|
||||
@@ -68,6 +99,41 @@ chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
|
||||
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
|
||||
written on `close()`.
|
||||
|
||||
## Editing a file in place
|
||||
|
||||
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
|
||||
where it is, through clawhdf5's `FileEditor`; the file is locked until
|
||||
`close()`, and every edit is written and synced before the statement
|
||||
returns.
|
||||
|
||||
```python
|
||||
with clawhdf5.File("data.h5", "r+") as f:
|
||||
ds = f["grid"]
|
||||
ds[10:20, ::2] = 0 # h5py keys and broadcasting
|
||||
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
|
||||
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
|
||||
f["series"].attrs["units"] = "K"
|
||||
f.attrs.create("version", 2, dtype="u1")
|
||||
```
|
||||
|
||||
- Values: a numpy array is converted to the dataset's dtype as libhdf5
|
||||
converts it (integers saturate at the target's limits; floats are
|
||||
truncated toward zero and clipped); anything else goes through
|
||||
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
|
||||
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
|
||||
value). A few libhdf5 edge cases differ on purpose; see
|
||||
`docs/known-issues.md`.
|
||||
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
|
||||
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
|
||||
see its result.
|
||||
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
|
||||
`str` is stored as a fixed-length UTF-8 string (h5py stores a
|
||||
variable-length one), so h5py reads it back as `bytes`.
|
||||
- Not supported (`NotImplementedError`, nothing written): creating or
|
||||
deleting datasets, groups and attributes, writing compound fields by
|
||||
name, variable-length data, HDF5 array types, and whatever
|
||||
`FileEditor` refuses (listed in `docs/known-issues.md`).
|
||||
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
@@ -76,7 +142,12 @@ pytest crates/clawhdf5-py/tests
|
||||
```
|
||||
|
||||
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
|
||||
writes. `scripts/ci-test.sh` builds the wheel and runs these in CI.
|
||||
writes, opened locally and over HTTP (an in-process range server,
|
||||
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
|
||||
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
|
||||
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
|
||||
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
|
||||
the wheel and runs these in CI.
|
||||
|
||||
## License
|
||||
|
||||
|
||||
+166
-65
@@ -1,22 +1,62 @@
|
||||
//! PyAttrs — dict-like access to HDF5 attributes.
|
||||
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::sync::{Arc, Mutex, PoisonError};
|
||||
|
||||
use clawhdf5_format::attribute::AttributeMessage;
|
||||
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
|
||||
use pyo3::exceptions::{PyKeyError, PyNotImplementedError, PyTypeError, PyValueError};
|
||||
use pyo3::prelude::*;
|
||||
use pyo3::types::{PyList, PyTuple};
|
||||
|
||||
use crate::convert::{Converter, Elements, resolve_vl};
|
||||
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value};
|
||||
use crate::handle::Handle;
|
||||
use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, edit, node, py_to_attr_value};
|
||||
|
||||
/// The attributes of an object in a file opened for reading (or editing).
|
||||
struct ReadAttrs {
|
||||
handle: Arc<Handle>,
|
||||
addr: u64,
|
||||
path: String,
|
||||
/// Sorted by name, with the file generation they were read at: an edit
|
||||
/// (`attrs[name] = value`, here or through another handle on the same
|
||||
/// object) makes them re-read.
|
||||
cache: Mutex<(u64, Arc<Vec<AttributeMessage>>)>,
|
||||
}
|
||||
|
||||
impl ReadAttrs {
|
||||
fn current(&self, py: Python<'_>) -> PyResult<Arc<Vec<AttributeMessage>>> {
|
||||
let generation = self.handle.generation();
|
||||
{
|
||||
let cached = self.cache.lock().unwrap_or_else(PoisonError::into_inner);
|
||||
if cached.0 == generation {
|
||||
return Ok(Arc::clone(&cached.1));
|
||||
}
|
||||
}
|
||||
let (addr, path) = (self.addr, &self.path);
|
||||
let attrs = Arc::new(self.handle.with(py, |f| node::attributes(f, addr, path))?);
|
||||
*self.cache.lock().unwrap_or_else(PoisonError::into_inner) =
|
||||
(generation, Arc::clone(&attrs));
|
||||
Ok(attrs)
|
||||
}
|
||||
|
||||
fn check_writable(&self) -> PyResult<()> {
|
||||
if self.handle.is_writable() {
|
||||
return Ok(());
|
||||
}
|
||||
Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||
"cannot set attributes on a read-only file (open it with mode 'r+')",
|
||||
))
|
||||
}
|
||||
|
||||
fn set(&self, py: Python<'_>, name: &str, value: clawhdf5_rs::AttrValue) -> PyResult<()> {
|
||||
self.check_writable()?;
|
||||
let path = node::name(&self.path);
|
||||
self.handle.edit(py, |ed| ed.set_attr(&path, name, &value))
|
||||
}
|
||||
}
|
||||
|
||||
/// Backing storage for attributes.
|
||||
enum AttrsInner {
|
||||
/// Attributes of an object in a file opened for reading, sorted by name.
|
||||
Read {
|
||||
file: Arc<clawhdf5_rs::File>,
|
||||
attrs: Vec<AttributeMessage>,
|
||||
},
|
||||
Read(ReadAttrs),
|
||||
/// Writable attribute list shared with a parent (PyFile or PyGroup).
|
||||
Write(Arc<Mutex<Vec<(String, OwnedAttrValue)>>>),
|
||||
}
|
||||
@@ -26,8 +66,11 @@ enum AttrsInner {
|
||||
/// In read mode, values are what h5py returns: numpy scalars for scalar
|
||||
/// attributes, numpy arrays otherwise, `str` for variable-length strings,
|
||||
/// `numpy.bytes_` for fixed-length ones, and `Empty` for a null dataspace.
|
||||
/// In write mode, attributes set here are accumulated and written when
|
||||
/// the parent file is closed.
|
||||
/// In a file opened with `'r+'`, `attrs[name] = value` adds or replaces an
|
||||
/// attribute in the file at once (as h5py stores it, except that `str`
|
||||
/// values become fixed-length UTF-8 strings). In write mode (`'w'`),
|
||||
/// attributes set here are accumulated and written when the parent file is
|
||||
/// closed.
|
||||
#[pyclass(name = "Attrs")]
|
||||
pub struct PyAttrs {
|
||||
inner: AttrsInner,
|
||||
@@ -36,10 +79,21 @@ pub struct PyAttrs {
|
||||
impl PyAttrs {
|
||||
/// The attributes of the object at `addr` (whose path is `path`) in a
|
||||
/// file opened for reading.
|
||||
pub(crate) fn read(file: Arc<clawhdf5_rs::File>, addr: u64, path: &str) -> PyResult<Self> {
|
||||
let attrs = node::attributes(&file, addr, path)?;
|
||||
pub(crate) fn read(
|
||||
py: Python<'_>,
|
||||
handle: Arc<Handle>,
|
||||
addr: u64,
|
||||
path: &str,
|
||||
) -> PyResult<Self> {
|
||||
let generation = handle.generation();
|
||||
let attrs = Arc::new(handle.with(py, |f| node::attributes(f, addr, path))?);
|
||||
Ok(Self {
|
||||
inner: AttrsInner::Read { file, attrs },
|
||||
inner: AttrsInner::Read(ReadAttrs {
|
||||
handle,
|
||||
addr,
|
||||
path: path.to_string(),
|
||||
cache: Mutex::new((generation, attrs)),
|
||||
}),
|
||||
})
|
||||
}
|
||||
|
||||
@@ -49,14 +103,48 @@ impl PyAttrs {
|
||||
inner: AttrsInner::Write(store),
|
||||
}
|
||||
}
|
||||
|
||||
fn set_value(
|
||||
&self,
|
||||
py: Python<'_>,
|
||||
key: &str,
|
||||
value: &Bound<'_, PyAny>,
|
||||
dtype: Option<&Bound<'_, PyAny>>,
|
||||
shape: Option<&Bound<'_, PyAny>>,
|
||||
) -> PyResult<()> {
|
||||
match &self.inner {
|
||||
AttrsInner::Read(r) => {
|
||||
r.check_writable()?;
|
||||
let value = edit::attr_value(py, value, dtype, shape)?;
|
||||
r.set(py, key, value)
|
||||
}
|
||||
AttrsInner::Write(store) => {
|
||||
if dtype.is_some() || shape.is_some() {
|
||||
return Err(PyNotImplementedError::new_err(
|
||||
"attrs.create with a dtype or shape is only supported in a file opened \
|
||||
with 'r+'",
|
||||
));
|
||||
}
|
||||
let owned = py_to_attr_value(value)?;
|
||||
let mut guard = store.lock().unwrap();
|
||||
// Replace existing key if present.
|
||||
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
|
||||
entry.1 = owned;
|
||||
} else {
|
||||
guard.push((key.to_string(), owned));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[pymethods]
|
||||
impl PyAttrs {
|
||||
fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
||||
match &self.inner {
|
||||
AttrsInner::Read { file, attrs } => match attrs.iter().find(|a| a.name == key) {
|
||||
Some(attr) => Ok(attr_to_py(py, file, attr)?.unbind()),
|
||||
AttrsInner::Read(r) => match r.current(py)?.iter().find(|a| a.name == key) {
|
||||
Some(attr) => Ok(attr_to_py(py, &r.handle, attr)?.unbind()),
|
||||
None => Err(PyKeyError::new_err(format!(
|
||||
"Can't open attribute (can't locate attribute: '{key}')"
|
||||
))),
|
||||
@@ -74,36 +162,63 @@ impl PyAttrs {
|
||||
}
|
||||
}
|
||||
|
||||
fn __setitem__(&self, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
||||
/// `attrs[name] = value`. In a file opened with `'r+'` this writes the
|
||||
/// attribute (numeric, bool, complex, bytes and str data, any shape)
|
||||
/// into the file before returning; see the class docs.
|
||||
fn __setitem__(&self, py: Python<'_>, key: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
||||
self.set_value(py, key, value, None, None)
|
||||
}
|
||||
|
||||
/// Deleting attributes is not supported in a file (the in-place editor
|
||||
/// cannot remove them); in write mode it removes a pending attribute.
|
||||
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||
match &self.inner {
|
||||
AttrsInner::Read { .. } => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||
"cannot set attributes on a read-only file",
|
||||
)),
|
||||
AttrsInner::Read(_) => Err(PyNotImplementedError::new_err(format!(
|
||||
"cannot delete attribute '{key}': deleting attributes is not supported by \
|
||||
clawhdf5's in-place editor"
|
||||
))),
|
||||
AttrsInner::Write(store) => {
|
||||
let owned = py_to_attr_value(value)?;
|
||||
let mut guard = store.lock().unwrap();
|
||||
// Replace existing key if present.
|
||||
if let Some(entry) = guard.iter_mut().find(|(k, _)| k == key) {
|
||||
entry.1 = owned;
|
||||
} else {
|
||||
guard.push((key.to_string(), owned));
|
||||
let before = guard.len();
|
||||
guard.retain(|(k, _)| k != key);
|
||||
if guard.len() == before {
|
||||
return Err(PyKeyError::new_err(key.to_string()));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn __len__(&self) -> usize {
|
||||
/// h5py's `attrs.create(name, data, shape=None, dtype=None)`: `data`
|
||||
/// converted to `dtype` and reshaped to `shape` first.
|
||||
#[pyo3(signature = (name, data, shape=None, dtype=None))]
|
||||
fn create(
|
||||
&self,
|
||||
py: Python<'_>,
|
||||
name: &str,
|
||||
data: &Bound<'_, PyAny>,
|
||||
shape: Option<&Bound<'_, PyAny>>,
|
||||
dtype: Option<&Bound<'_, PyAny>>,
|
||||
) -> PyResult<()> {
|
||||
self.set_value(py, name, data, dtype, shape)
|
||||
}
|
||||
|
||||
/// h5py's `attrs.modify(name, value)`: same as `attrs[name] = value`.
|
||||
fn modify(&self, py: Python<'_>, name: &str, value: &Bound<'_, PyAny>) -> PyResult<()> {
|
||||
self.set_value(py, name, value, None, None)
|
||||
}
|
||||
|
||||
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||
match &self.inner {
|
||||
AttrsInner::Read { attrs, .. } => attrs.len(),
|
||||
AttrsInner::Write(store) => store.lock().unwrap().len(),
|
||||
AttrsInner::Read(r) => Ok(r.current(py)?.len()),
|
||||
AttrsInner::Write(store) => Ok(store.lock().unwrap().len()),
|
||||
}
|
||||
}
|
||||
|
||||
fn __contains__(&self, key: &str) -> bool {
|
||||
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||
match &self.inner {
|
||||
AttrsInner::Read { attrs, .. } => attrs.iter().any(|a| a.name == key),
|
||||
AttrsInner::Write(store) => store.lock().unwrap().iter().any(|(k, _)| k == key),
|
||||
AttrsInner::Read(r) => Ok(r.current(py)?.iter().any(|a| a.name == key)),
|
||||
AttrsInner::Write(store) => Ok(store.lock().unwrap().iter().any(|(k, _)| k == key)),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -113,15 +228,17 @@ impl PyAttrs {
|
||||
Ok(iter)
|
||||
}
|
||||
|
||||
fn __repr__(&self) -> String {
|
||||
let n = self.__len__();
|
||||
format!("<HDF5 Attrs ({n} members)>")
|
||||
fn __repr__(&self, py: Python<'_>) -> String {
|
||||
match self.__len__(py) {
|
||||
Ok(n) => format!("<HDF5 Attrs ({n} members)>"),
|
||||
Err(_) => "<HDF5 Attrs>".to_string(),
|
||||
}
|
||||
}
|
||||
|
||||
/// The value of `key`, or `default` if there is no such attribute.
|
||||
#[pyo3(signature = (key, default=None))]
|
||||
fn get(&self, py: Python<'_>, key: &str, default: Option<Py<PyAny>>) -> PyResult<Py<PyAny>> {
|
||||
if self.__contains__(key) {
|
||||
if self.__contains__(py, key)? {
|
||||
self.__getitem__(py, key)
|
||||
} else {
|
||||
Ok(default.unwrap_or_else(|| py.None()))
|
||||
@@ -131,7 +248,7 @@ impl PyAttrs {
|
||||
/// Return attribute names as a list.
|
||||
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||
let names: Vec<String> = match &self.inner {
|
||||
AttrsInner::Read { attrs, .. } => attrs.iter().map(|a| a.name.clone()).collect(),
|
||||
AttrsInner::Read(r) => r.current(py)?.iter().map(|a| a.name.clone()).collect(),
|
||||
AttrsInner::Write(store) => store
|
||||
.lock()
|
||||
.unwrap()
|
||||
@@ -146,9 +263,10 @@ impl PyAttrs {
|
||||
/// Return attribute values as a list.
|
||||
fn values(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||
let vals: Vec<Py<PyAny>> = match &self.inner {
|
||||
AttrsInner::Read { file, attrs } => attrs
|
||||
AttrsInner::Read(r) => r
|
||||
.current(py)?
|
||||
.iter()
|
||||
.map(|a| attr_to_py(py, file, a).map(Bound::unbind))
|
||||
.map(|a| attr_to_py(py, &r.handle, a).map(Bound::unbind))
|
||||
.collect::<PyResult<_>>()?,
|
||||
AttrsInner::Write(store) => store
|
||||
.lock()
|
||||
@@ -167,9 +285,10 @@ impl PyAttrs {
|
||||
/// Return attribute (key, value) pairs as a list of tuples.
|
||||
fn items(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||
let pairs: Vec<(String, Py<PyAny>)> = match &self.inner {
|
||||
AttrsInner::Read { file, attrs } => attrs
|
||||
AttrsInner::Read(r) => r
|
||||
.current(py)?
|
||||
.iter()
|
||||
.map(|a| Ok((a.name.clone(), attr_to_py(py, file, a)?.unbind())))
|
||||
.map(|a| Ok((a.name.clone(), attr_to_py(py, &r.handle, a)?.unbind())))
|
||||
.collect::<PyResult<_>>()?,
|
||||
AttrsInner::Write(store) => store
|
||||
.lock()
|
||||
@@ -189,12 +308,11 @@ impl PyAttrs {
|
||||
/// An attribute's value as h5py returns it.
|
||||
fn attr_to_py<'py>(
|
||||
py: Python<'py>,
|
||||
file: &clawhdf5_rs::File,
|
||||
handle: &Handle,
|
||||
attr: &AttributeMessage,
|
||||
) -> PyResult<Bound<'py, PyAny>> {
|
||||
crate::no_panic(|| {
|
||||
let sb = file.superblock();
|
||||
let conv = Converter::new(py, &attr.datatype, sb.offset_size)
|
||||
let conv = Converter::new(py, &attr.datatype, handle.offset_size)
|
||||
.map_err(|e| prefix_err(py, &attr.name, e))?;
|
||||
if node::is_null(&attr.dataspace) {
|
||||
return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any());
|
||||
@@ -216,12 +334,11 @@ fn attr_to_py<'py>(
|
||||
)));
|
||||
}
|
||||
let raw = &attr.raw_data[..want];
|
||||
let file_data = file.as_bytes();
|
||||
let (osz, lsz, unit) = (sb.offset_size, sb.length_size, conv.vl_unit);
|
||||
Elements::Vl(
|
||||
py.detach(|| resolve_vl(file_data, raw, n, osz, lsz, unit))
|
||||
.map_err(|e| PyValueError::new_err(format!("attribute {}: {e}", attr.name)))?,
|
||||
)
|
||||
let (osz, lsz, unit) = (handle.offset_size, handle.length_size, conv.vl_unit);
|
||||
let what = format!("attribute {}", attr.name);
|
||||
Elements::Vl(handle.with(py, |f| {
|
||||
resolve_vl(f.storage(), raw, n, osz, lsz, unit).map_err(|e| e.into_py(&what))
|
||||
})?)
|
||||
} else {
|
||||
Elements::Bytes(attr.raw_data.clone())
|
||||
};
|
||||
@@ -244,19 +361,3 @@ fn prefix_err(py: Python<'_>, name: &str, e: PyErr) -> PyErr {
|
||||
PyValueError::new_err(msg)
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn write_attrs_len() {
|
||||
let store = Arc::new(Mutex::new(Vec::new()));
|
||||
store
|
||||
.lock()
|
||||
.unwrap()
|
||||
.push(("key".into(), OwnedAttrValue::I64(99)));
|
||||
let attrs = PyAttrs::from_write(store);
|
||||
assert_eq!(attrs.__len__(), 1);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -17,6 +17,7 @@ use std::collections::HashMap;
|
||||
|
||||
use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder};
|
||||
use clawhdf5_format::global_heap::GlobalHeapCollection;
|
||||
use clawhdf5_format::storage::Storage;
|
||||
use numpy::PyArray1;
|
||||
use pyo3::exceptions::{PyTypeError, PyValueError};
|
||||
use pyo3::prelude::*;
|
||||
@@ -538,16 +539,16 @@ fn object_array<'py>(
|
||||
/// their bytes: each element's stored length times `unit` (1 for strings,
|
||||
/// the base type's size for sequences). Pure Rust, so it runs without the
|
||||
/// GIL.
|
||||
pub(crate) fn resolve_vl(
|
||||
file_data: &[u8],
|
||||
pub(crate) fn resolve_vl<S: Storage + ?Sized>(
|
||||
file: &S,
|
||||
raw: &[u8],
|
||||
count: usize,
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
unit: usize,
|
||||
) -> Result<Vec<Vec<u8>>, String> {
|
||||
) -> Result<Vec<Vec<u8>>, VlError> {
|
||||
let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size)
|
||||
.map_err(|e| e.to_string())?;
|
||||
.map_err(|e| VlError::Invalid(e.to_string()))?;
|
||||
let undefined = match offset_size {
|
||||
2 => 0xFFFF,
|
||||
4 => 0xFFFF_FFFF,
|
||||
@@ -558,43 +559,70 @@ pub(crate) fn resolve_vl(
|
||||
for vl in &refs {
|
||||
if vl.collection_address == 0 || vl.collection_address == undefined {
|
||||
if vl.length != 0 {
|
||||
return Err(format!(
|
||||
return Err(VlError::Invalid(format!(
|
||||
"variable-length element of length {} has no heap address",
|
||||
vl.length
|
||||
));
|
||||
)));
|
||||
}
|
||||
out.push(Vec::new());
|
||||
continue;
|
||||
}
|
||||
let coll = match collections.entry(vl.collection_address) {
|
||||
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
|
||||
std::collections::hash_map::Entry::Vacant(e) => {
|
||||
let addr = usize::try_from(vl.collection_address)
|
||||
.map_err(|_| "global heap address out of range".to_string())?;
|
||||
e.insert(
|
||||
GlobalHeapCollection::parse(file_data, addr, length_size)
|
||||
.map_err(|e| e.to_string())?,
|
||||
)
|
||||
}
|
||||
std::collections::hash_map::Entry::Vacant(e) => e.insert(
|
||||
GlobalHeapCollection::parse_in(file, vl.collection_address, length_size)
|
||||
.map_err(VlError::from_format)?,
|
||||
),
|
||||
};
|
||||
let index = u16::try_from(vl.object_index)
|
||||
.map_err(|_| format!("global heap object index {} out of range", vl.object_index))?;
|
||||
let index = u16::try_from(vl.object_index).map_err(|_| {
|
||||
VlError::Invalid(format!(
|
||||
"global heap object index {} out of range",
|
||||
vl.object_index
|
||||
))
|
||||
})?;
|
||||
let obj = coll.get_object(index).ok_or_else(|| {
|
||||
format!(
|
||||
VlError::Invalid(format!(
|
||||
"global heap object {index} not found in the collection at {}",
|
||||
vl.collection_address
|
||||
)
|
||||
))
|
||||
})?;
|
||||
let need = (vl.length as usize)
|
||||
.checked_mul(unit)
|
||||
.ok_or("variable-length element too long")?;
|
||||
.ok_or_else(|| VlError::Invalid("variable-length element too long".into()))?;
|
||||
if need > obj.data.len() {
|
||||
return Err(format!(
|
||||
return Err(VlError::Invalid(format!(
|
||||
"variable-length element of {need} bytes in a {}-byte heap object",
|
||||
obj.data.len()
|
||||
));
|
||||
)));
|
||||
}
|
||||
out.push(obj.data[..need].to_vec());
|
||||
}
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
/// Why variable-length elements could not be resolved.
|
||||
#[derive(Debug)]
|
||||
pub(crate) enum VlError {
|
||||
/// Reading the file failed (a network error on a remote file).
|
||||
Storage(String),
|
||||
/// The references or the heap are not valid.
|
||||
Invalid(String),
|
||||
}
|
||||
|
||||
impl VlError {
|
||||
fn from_format(e: clawhdf5_format::error::FormatError) -> Self {
|
||||
match e {
|
||||
clawhdf5_format::error::FormatError::Storage(_) => VlError::Storage(e.to_string()),
|
||||
e => VlError::Invalid(e.to_string()),
|
||||
}
|
||||
}
|
||||
|
||||
/// As a Python exception, the message prefixed with `what`: a storage
|
||||
/// failure is an `OSError`, anything else a `ValueError`.
|
||||
pub(crate) fn into_py(self, what: &str) -> PyErr {
|
||||
match self {
|
||||
VlError::Storage(m) => pyo3::exceptions::PyOSError::new_err(format!("{what}: {m}")),
|
||||
VlError::Invalid(m) => PyValueError::new_err(format!("{what}: {m}")),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -6,20 +6,52 @@
|
||||
//! whole dataset instead); the
|
||||
//! bytes it returns become the numpy array's buffer without a copy (see
|
||||
//! `convert`). All file access and decoding runs with the GIL released, so
|
||||
//! Python threads reading the same or different datasets run in parallel.
|
||||
//! Python threads reading the same or different datasets run in parallel,
|
||||
//! and a remote file's network reads never hold the GIL.
|
||||
|
||||
use std::sync::Arc;
|
||||
use std::sync::{Arc, Mutex, PoisonError};
|
||||
|
||||
use clawhdf5_format::datatype::Datatype;
|
||||
use clawhdf5_format::object_header::ObjectHeader;
|
||||
use pyo3::exceptions::{PyTypeError, PyValueError};
|
||||
use clawhdf5_rs::File;
|
||||
use pyo3::exceptions::{PyNotImplementedError, PyOSError, PyTypeError, PyValueError};
|
||||
use pyo3::prelude::*;
|
||||
use pyo3::types::{PyList, PyTuple};
|
||||
|
||||
use crate::attrs::PyAttrs;
|
||||
use crate::convert::{Converter, Elements, resolve_vl};
|
||||
use crate::convert::{Converter, Elements, VlError, resolve_vl};
|
||||
use crate::handle::Handle;
|
||||
use crate::select::{self, Plan};
|
||||
use crate::{PyEmpty, node, to_py_err};
|
||||
use crate::{PyEmpty, edit, node, to_py_err};
|
||||
|
||||
/// What opening a dataset reads from the file (without the GIL).
|
||||
pub(crate) struct DatasetMeta {
|
||||
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
|
||||
shape: Option<Vec<u64>>,
|
||||
chunks: Option<Vec<u64>>,
|
||||
datatype: Datatype,
|
||||
}
|
||||
|
||||
impl DatasetMeta {
|
||||
pub(crate) fn load(f: &File, addr: u64, hdr: &ObjectHeader, path: &str) -> PyResult<Self> {
|
||||
let null = node::is_null(&node::dataspace(f, hdr, path)?);
|
||||
let ds = f.dataset_at(addr).map_err(to_py_err)?;
|
||||
let shape = if null {
|
||||
None
|
||||
} else {
|
||||
Some(ds.shape().map_err(to_py_err)?)
|
||||
};
|
||||
let datatype = ds.raw_datatype().map_err(to_py_err)?;
|
||||
let chunks = shape
|
||||
.as_ref()
|
||||
.and_then(|s| node::chunk_shape(f, hdr, s.len()));
|
||||
Ok(Self {
|
||||
shape,
|
||||
chunks,
|
||||
datatype,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// A dataset in a file opened for reading.
|
||||
///
|
||||
@@ -30,13 +62,14 @@ use crate::{PyEmpty, node, to_py_err};
|
||||
/// ```
|
||||
#[pyclass(name = "Dataset")]
|
||||
pub struct PyDataset {
|
||||
file: Arc<clawhdf5_rs::File>,
|
||||
handle: Arc<Handle>,
|
||||
path: String,
|
||||
/// Where the dataset's object header is: reads open it from here rather
|
||||
/// than resolve `path` again.
|
||||
addr: u64,
|
||||
/// `None` for a dataset with a null dataspace (h5py's `Empty`).
|
||||
shape: Option<Vec<u64>>,
|
||||
/// The shape (`None` for a null dataspace, h5py's `Empty`), with the
|
||||
/// file generation it was read at: an edit (a resize) may change it.
|
||||
shape: Mutex<(u64, Option<Vec<u64>>)>,
|
||||
/// The chunk shape, for a chunked dataset.
|
||||
chunks: Option<Vec<u64>>,
|
||||
datatype: Datatype,
|
||||
@@ -45,39 +78,25 @@ pub struct PyDataset {
|
||||
}
|
||||
|
||||
impl PyDataset {
|
||||
pub(crate) fn open(
|
||||
pub(crate) fn new(
|
||||
py: Python<'_>,
|
||||
file: Arc<clawhdf5_rs::File>,
|
||||
handle: Arc<Handle>,
|
||||
path: String,
|
||||
addr: u64,
|
||||
hdr: &ObjectHeader,
|
||||
) -> PyResult<Self> {
|
||||
crate::no_panic(|| {
|
||||
let null = node::is_null(&node::dataspace(&file, hdr)?);
|
||||
let (shape, datatype) = {
|
||||
let ds = file.dataset_at(addr).map_err(to_py_err)?;
|
||||
let shape = if null {
|
||||
None
|
||||
} else {
|
||||
Some(ds.shape().map_err(to_py_err)?)
|
||||
};
|
||||
(shape, ds.raw_datatype().map_err(to_py_err)?)
|
||||
};
|
||||
let conv = Converter::new(py, &datatype, file.superblock().offset_size)
|
||||
.map_err(|e| e.value(py).to_string());
|
||||
let chunks = shape
|
||||
.as_ref()
|
||||
.and_then(|s| node::chunk_shape(&file, hdr, s.len()));
|
||||
Ok(Self {
|
||||
file,
|
||||
path,
|
||||
addr,
|
||||
shape,
|
||||
chunks,
|
||||
datatype,
|
||||
conv,
|
||||
})
|
||||
})
|
||||
meta: DatasetMeta,
|
||||
) -> Self {
|
||||
let generation = handle.generation();
|
||||
let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size))
|
||||
.map_err(|e| e.value(py).to_string());
|
||||
Self {
|
||||
handle,
|
||||
path,
|
||||
addr,
|
||||
shape: Mutex::new((generation, meta.shape)),
|
||||
chunks: meta.chunks,
|
||||
datatype: meta.datatype,
|
||||
conv,
|
||||
}
|
||||
}
|
||||
|
||||
fn converter(&self) -> PyResult<&Converter> {
|
||||
@@ -86,10 +105,54 @@ impl PyDataset {
|
||||
.map_err(|msg| PyTypeError::new_err(format!("{}: {msg}", node::name(&self.path))))
|
||||
}
|
||||
|
||||
/// The current shape: the one read at open, or re-read after an edit.
|
||||
fn dims(&self, py: Python<'_>) -> PyResult<Option<Vec<u64>>> {
|
||||
let generation = self.handle.generation();
|
||||
{
|
||||
let cached = self.shape.lock().unwrap_or_else(PoisonError::into_inner);
|
||||
if cached.0 == generation {
|
||||
return Ok(cached.1.clone());
|
||||
}
|
||||
}
|
||||
let addr = self.addr;
|
||||
let null = self
|
||||
.shape
|
||||
.lock()
|
||||
.unwrap_or_else(PoisonError::into_inner)
|
||||
.1
|
||||
.is_none();
|
||||
let shape = if null {
|
||||
None
|
||||
} else {
|
||||
Some(self.handle.with(py, |f| {
|
||||
f.dataset_at(addr)
|
||||
.and_then(|ds| ds.shape())
|
||||
.map_err(to_py_err)
|
||||
})?)
|
||||
};
|
||||
*self.shape.lock().unwrap_or_else(PoisonError::into_inner) = (generation, shape.clone());
|
||||
Ok(shape)
|
||||
}
|
||||
|
||||
fn check_writable(&self) -> PyResult<()> {
|
||||
if self.handle.is_writable() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(PyOSError::new_err(format!(
|
||||
"{}: the file is open read-only; open it with mode 'r+' to change it",
|
||||
node::name(&self.path)
|
||||
)))
|
||||
}
|
||||
}
|
||||
|
||||
/// Read the selection described by `plan` into a numpy array.
|
||||
fn read_plan<'py>(&self, py: Python<'py>, plan: &Plan) -> PyResult<Bound<'py, PyAny>> {
|
||||
fn read_plan<'py>(
|
||||
&self,
|
||||
py: Python<'py>,
|
||||
plan: &Plan,
|
||||
dims: &[u64],
|
||||
) -> PyResult<Bound<'py, PyAny>> {
|
||||
let conv = self.converter()?;
|
||||
let dims = self.shape.as_deref().unwrap_or(&[]);
|
||||
let out_shape = plan.out_shape();
|
||||
|
||||
let arr = if plan.is_empty() {
|
||||
@@ -103,10 +166,10 @@ impl PyDataset {
|
||||
};
|
||||
let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size);
|
||||
let read_shape = plan.read_shape();
|
||||
let file = &*self.file;
|
||||
let handle = &*self.handle;
|
||||
let addr = self.addr;
|
||||
// Everything below touches only Rust data: release the GIL.
|
||||
let read = || -> Result<Elements, ReadError> {
|
||||
let read = |file: &File| -> Result<Elements, ReadError> {
|
||||
let ds = file.dataset_at(addr)?;
|
||||
let mut blocks = Vec::with_capacity(reads.len());
|
||||
for read in reads {
|
||||
@@ -147,7 +210,7 @@ impl PyDataset {
|
||||
let sb = file.superblock();
|
||||
let n = read_shape.iter().product();
|
||||
resolve_vl(
|
||||
file.as_bytes(),
|
||||
file.storage(),
|
||||
&raw,
|
||||
n,
|
||||
sb.offset_size,
|
||||
@@ -155,12 +218,20 @@ impl PyDataset {
|
||||
unit,
|
||||
)
|
||||
.map(Elements::Vl)
|
||||
.map_err(ReadError::Other)
|
||||
.map_err(ReadError::Vl)
|
||||
};
|
||||
let data = py
|
||||
.detach(|| {
|
||||
std::panic::catch_unwind(std::panic::AssertUnwindSafe(read))
|
||||
.unwrap_or_else(|p| Err(ReadError::Panic(crate::panic_text(&*p))))
|
||||
handle
|
||||
.with_detached(|f| {
|
||||
Ok(
|
||||
std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| read(f)))
|
||||
.unwrap_or_else(|p| {
|
||||
Err(ReadError::Panic(crate::panic_text(&*p)))
|
||||
}),
|
||||
)
|
||||
})
|
||||
.unwrap_or_else(|e| Err(ReadError::Py(e)))
|
||||
})
|
||||
.map_err(|e| e.into_py(&self.path))?;
|
||||
let joined = conv.to_array(py, data, &read_shape, false)?;
|
||||
@@ -183,6 +254,8 @@ impl PyDataset {
|
||||
/// An error from the read closure, turned into a Python error with the GIL.
|
||||
enum ReadError {
|
||||
Lib(clawhdf5_rs::Error),
|
||||
Vl(VlError),
|
||||
Py(PyErr),
|
||||
Other(String),
|
||||
Panic(String),
|
||||
}
|
||||
@@ -197,6 +270,8 @@ impl ReadError {
|
||||
fn into_py(self, path: &str) -> PyErr {
|
||||
match self {
|
||||
ReadError::Lib(e) => to_py_err(e),
|
||||
ReadError::Vl(e) => e.into_py(&node::name(path)),
|
||||
ReadError::Py(e) => e,
|
||||
ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))),
|
||||
ReadError::Panic(msg) => crate::InternalError::new_err(format!(
|
||||
"{}: clawhdf5 internal error (please report it): {msg}",
|
||||
@@ -243,7 +318,7 @@ impl PyDataset {
|
||||
/// The shape of the dataset (`None` for an empty/null dataspace).
|
||||
#[getter]
|
||||
fn shape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||
match &self.shape {
|
||||
match self.dims(py)? {
|
||||
Some(s) => Ok(PyTuple::new(py, s)?.into_any()),
|
||||
None => Ok(py.None().into_bound(py)),
|
||||
}
|
||||
@@ -252,22 +327,32 @@ impl PyDataset {
|
||||
/// The maximum shape (`None` per unlimited dimension), like h5py.
|
||||
#[getter]
|
||||
fn maxshape<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||
crate::no_panic(|| {
|
||||
let Some(shape) = &self.shape else {
|
||||
return Ok(py.None().into_bound(py));
|
||||
};
|
||||
let max = self
|
||||
.file
|
||||
.dataset_at(self.addr)
|
||||
.and_then(|ds| ds.max_dimensions())
|
||||
.map_err(to_py_err)?
|
||||
.unwrap_or_else(|| shape.clone());
|
||||
let items: Vec<Option<u64>> = max
|
||||
.into_iter()
|
||||
.map(|d| (d != u64::MAX).then_some(d))
|
||||
.collect();
|
||||
Ok(PyTuple::new(py, items)?.into_any())
|
||||
})
|
||||
let Some(shape) = self.dims(py)? else {
|
||||
return Ok(py.None().into_bound(py));
|
||||
};
|
||||
let addr = self.addr;
|
||||
let max = self
|
||||
.handle
|
||||
.with(py, |f| {
|
||||
f.dataset_at(addr)
|
||||
.and_then(|ds| ds.max_dimensions())
|
||||
.map_err(to_py_err)
|
||||
})?
|
||||
.unwrap_or(shape);
|
||||
let items: Vec<Option<u64>> = max
|
||||
.into_iter()
|
||||
.map(|d| (d != u64::MAX).then_some(d))
|
||||
.collect();
|
||||
Ok(PyTuple::new(py, items)?.into_any())
|
||||
}
|
||||
|
||||
/// The chunk shape, or `None` for a dataset that is not chunked.
|
||||
#[getter]
|
||||
fn chunks<'py>(&self, py: Python<'py>) -> PyResult<Bound<'py, PyAny>> {
|
||||
match &self.chunks {
|
||||
Some(c) => Ok(PyTuple::new(py, c)?.into_any()),
|
||||
None => Ok(py.None().into_bound(py)),
|
||||
}
|
||||
}
|
||||
|
||||
/// The dataset's numpy dtype, as h5py reports it.
|
||||
@@ -277,14 +362,14 @@ impl PyDataset {
|
||||
}
|
||||
|
||||
#[getter]
|
||||
fn ndim(&self) -> usize {
|
||||
self.shape.as_ref().map_or(0, Vec::len)
|
||||
fn ndim(&self, py: Python<'_>) -> PyResult<usize> {
|
||||
Ok(self.dims(py)?.map_or(0, |s| s.len()))
|
||||
}
|
||||
|
||||
/// Number of elements (`None` for an empty/null dataspace, as h5py).
|
||||
#[getter]
|
||||
fn size(&self) -> Option<u64> {
|
||||
self.shape.as_ref().map(|s| s.iter().product())
|
||||
fn size(&self, py: Python<'_>) -> PyResult<Option<u64>> {
|
||||
Ok(self.dims(py)?.map(|s| s.iter().product()))
|
||||
}
|
||||
|
||||
/// The dataset's full name, e.g. `/group/data`.
|
||||
@@ -293,10 +378,11 @@ impl PyDataset {
|
||||
node::name(&self.path)
|
||||
}
|
||||
|
||||
/// The dataset's attributes (read-only, dict-like).
|
||||
/// The dataset's attributes (dict-like; writable in a file opened with
|
||||
/// `'r+'`).
|
||||
#[getter]
|
||||
fn attrs(&self) -> PyResult<PyAttrs> {
|
||||
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
|
||||
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
|
||||
}
|
||||
|
||||
/// Read with h5py indexing: integers, slices with positive steps,
|
||||
@@ -308,7 +394,7 @@ impl PyDataset {
|
||||
py: Python<'py>,
|
||||
key: &Bound<'py, PyAny>,
|
||||
) -> PyResult<Bound<'py, PyAny>> {
|
||||
let Some(dims) = &self.shape else {
|
||||
let Some(dims) = self.dims(py)? else {
|
||||
let is_empty_tuple = key.cast::<PyTuple>().is_ok_and(|t| t.is_empty());
|
||||
let is_ellipsis = key.is_instance_of::<pyo3::types::PyEllipsis>();
|
||||
if is_empty_tuple || is_ellipsis {
|
||||
@@ -317,8 +403,109 @@ impl PyDataset {
|
||||
}
|
||||
return Err(PyValueError::new_err("Empty datasets cannot be sliced"));
|
||||
};
|
||||
let plan = select::parse(key, dims)?;
|
||||
self.read_plan(py, &plan)
|
||||
let plan = select::parse(key, &dims)?;
|
||||
self.read_plan(py, &plan, &dims)
|
||||
}
|
||||
|
||||
/// Write with h5py indexing (file opened with `'r+'`): `ds[key] = value`.
|
||||
///
|
||||
/// The key is what `ds[key]` reads (without compound field names). The
|
||||
/// value is converted to the dataset's dtype as h5py converts it (a
|
||||
/// numpy array as libhdf5 does, clipping out-of-range numbers; anything
|
||||
/// else through `numpy.asarray(value, dtype=ds.dtype)`), and broadcast
|
||||
/// to the selection as h5py broadcasts. The edit is written and synced
|
||||
/// before this returns; what the in-place editor cannot write raises
|
||||
/// `NotImplementedError` and leaves the file as it was.
|
||||
fn __setitem__(
|
||||
&self,
|
||||
py: Python<'_>,
|
||||
key: &Bound<'_, PyAny>,
|
||||
value: &Bound<'_, PyAny>,
|
||||
) -> PyResult<()> {
|
||||
self.check_writable()?;
|
||||
let Some(dims) = self.dims(py)? else {
|
||||
return Err(PyNotImplementedError::new_err(
|
||||
"writing to an empty (null dataspace) dataset is not supported",
|
||||
));
|
||||
};
|
||||
let plan = select::parse(key, &dims)?;
|
||||
if !plan.fields.is_empty() {
|
||||
return Err(PyNotImplementedError::new_err(
|
||||
"writing compound fields by name is not supported by clawhdf5's in-place editor; \
|
||||
write whole elements",
|
||||
));
|
||||
}
|
||||
let category = edit::category(&self.datatype)?;
|
||||
let conv = self.converter()?;
|
||||
let bytes = edit::dataset_bytes(
|
||||
py,
|
||||
value,
|
||||
conv.dtype.bind(py),
|
||||
category,
|
||||
&plan,
|
||||
self.chunks.as_deref(),
|
||||
)?;
|
||||
if plan.is_empty() {
|
||||
return Ok(());
|
||||
}
|
||||
let sel = edit::selection(&plan, &dims)?;
|
||||
let path = node::name(&self.path);
|
||||
self.handle
|
||||
.edit(py, |ed| ed.write_selection(&path, &sel, &bytes))
|
||||
}
|
||||
|
||||
/// Change the dataset's shape (file opened with `'r+'`), as h5py's
|
||||
/// `Dataset.resize`: `ds.resize((100, 20))`, or `ds.resize(100, axis=0)`.
|
||||
/// Only chunked datasets, within their maximum shape; new elements read
|
||||
/// as the fill value.
|
||||
#[pyo3(signature = (size, axis=None))]
|
||||
fn resize(&self, py: Python<'_>, size: &Bound<'_, PyAny>, axis: Option<isize>) -> PyResult<()> {
|
||||
self.check_writable()?;
|
||||
let Some(dims) = self.dims(py)? else {
|
||||
return Err(PyTypeError::new_err("Empty datasets cannot be resized"));
|
||||
};
|
||||
if self.chunks.is_none() {
|
||||
return Err(PyTypeError::new_err("Only chunked datasets can be resized"));
|
||||
}
|
||||
let shape: Vec<u64> = match axis {
|
||||
Some(axis) => {
|
||||
let rank = dims.len();
|
||||
let a = usize::try_from(axis)
|
||||
.ok()
|
||||
.filter(|&a| a < rank)
|
||||
.ok_or_else(|| {
|
||||
PyValueError::new_err(format!(
|
||||
"Invalid axis (0 to {} allowed)",
|
||||
rank.saturating_sub(1)
|
||||
))
|
||||
})?;
|
||||
let n: u64 = size.extract().map_err(|_| {
|
||||
PyTypeError::new_err("Argument must be a single int if axis is specified")
|
||||
})?;
|
||||
let mut s = dims.clone();
|
||||
s[a] = n;
|
||||
s
|
||||
}
|
||||
// As h5py: without `axis` the size is a sequence (`tuple(size)`).
|
||||
None => size.extract().map_err(|_| {
|
||||
PyTypeError::new_err(format!(
|
||||
"'{}' object is not iterable",
|
||||
size.get_type()
|
||||
.name()
|
||||
.map(|n| n.to_string())
|
||||
.unwrap_or_default()
|
||||
))
|
||||
})?,
|
||||
};
|
||||
if shape.len() != dims.len() {
|
||||
return Err(PyValueError::new_err(format!(
|
||||
"new shape {shape:?} has {} dimensions, the dataset {}",
|
||||
shape.len(),
|
||||
dims.len()
|
||||
)));
|
||||
}
|
||||
let path = node::name(&self.path);
|
||||
self.handle.edit(py, |ed| ed.resize(&path, &shape))
|
||||
}
|
||||
|
||||
/// `numpy.asarray(ds)` reads the whole dataset.
|
||||
@@ -330,20 +517,20 @@ impl PyDataset {
|
||||
copy: Option<bool>,
|
||||
) -> PyResult<Bound<'py, PyAny>> {
|
||||
let _ = copy; // every read is a fresh array
|
||||
let Some(dims) = &self.shape else {
|
||||
let Some(dims) = self.dims(py)? else {
|
||||
return Err(PyValueError::new_err("an empty dataset has no array value"));
|
||||
};
|
||||
let ellipsis = pyo3::types::PyEllipsis::get(py).to_owned().into_any();
|
||||
let plan = select::parse(&ellipsis, dims)?;
|
||||
let arr = self.read_plan(py, &plan)?;
|
||||
let plan = select::parse(&ellipsis, &dims)?;
|
||||
let arr = self.read_plan(py, &plan, &dims)?;
|
||||
match dtype {
|
||||
Some(dt) => arr.call_method1("astype", (dt,)),
|
||||
None => Ok(arr),
|
||||
}
|
||||
}
|
||||
|
||||
fn __len__(&self) -> PyResult<usize> {
|
||||
match self.shape.as_deref() {
|
||||
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||
match self.dims(py)?.as_deref() {
|
||||
Some([first, ..]) => Ok(*first as usize),
|
||||
_ => Err(PyTypeError::new_err(
|
||||
"Attempt to take len() of scalar dataset",
|
||||
@@ -361,9 +548,10 @@ impl PyDataset {
|
||||
.unwrap_or_default(),
|
||||
Err(_) => format!("{:?}", self.datatype),
|
||||
};
|
||||
let shape = match &self.shape {
|
||||
Some(s) => format!("{s:?}"),
|
||||
None => "None".to_string(),
|
||||
let shape = match self.dims(py) {
|
||||
Ok(Some(s)) => format!("{s:?}"),
|
||||
Ok(None) => "None".to_string(),
|
||||
Err(_) => "?".to_string(),
|
||||
};
|
||||
format!(
|
||||
"<HDF5 dataset \"{}\": shape {shape}, type \"{dtype}\">",
|
||||
|
||||
@@ -0,0 +1,387 @@
|
||||
//! In-place editing (`clawhdf5.File(path, 'r+')`) through `FileEditor`:
|
||||
//! turning what Python assigns into the bytes, selections and attribute
|
||||
//! values the editor takes.
|
||||
//!
|
||||
//! Value conversion follows h5py (see `edit_helpers.py`, run inside the
|
||||
//! extension module); what `FileEditor` cannot do is `NotImplementedError`
|
||||
//! before anything is written.
|
||||
|
||||
use std::ffi::CString;
|
||||
|
||||
use clawhdf5_format::datatype::{
|
||||
CharacterSet, CompoundMember, Datatype, DatatypeByteOrder, EnumMember, StringPadding,
|
||||
};
|
||||
use clawhdf5_format::selection::Selection;
|
||||
use clawhdf5_rs::AttrValue;
|
||||
use pyo3::exceptions::{PyNotImplementedError, PyTypeError};
|
||||
use pyo3::prelude::*;
|
||||
use pyo3::sync::PyOnceLock;
|
||||
use pyo3::types::{PyBytes, PyModule, PyTuple};
|
||||
|
||||
use crate::select::{Axis, Plan};
|
||||
|
||||
/// The helper module, compiled once.
|
||||
pub(crate) fn helpers(py: Python<'_>) -> PyResult<&Bound<'_, PyModule>> {
|
||||
static HELPERS: PyOnceLock<Py<PyModule>> = PyOnceLock::new();
|
||||
let module = HELPERS.get_or_try_init(py, || -> PyResult<Py<PyModule>> {
|
||||
let code = CString::new(include_str!("edit_helpers.py"))
|
||||
.map_err(|e| PyTypeError::new_err(e.to_string()))?;
|
||||
Ok(PyModule::from_code(
|
||||
py,
|
||||
&code,
|
||||
c"clawhdf5/edit_helpers.py",
|
||||
c"clawhdf5._edit_helpers",
|
||||
)?
|
||||
.unbind())
|
||||
})?;
|
||||
Ok(module.bind(py))
|
||||
}
|
||||
|
||||
fn not_implemented(what: impl std::fmt::Display) -> PyErr {
|
||||
PyNotImplementedError::new_err(format!(
|
||||
"{what} is not supported by clawhdf5's in-place editor"
|
||||
))
|
||||
}
|
||||
|
||||
/// How values for a dataset of type `dt` are converted (a category of
|
||||
/// `edit_helpers._convert_array`), or why they cannot be written.
|
||||
pub(crate) fn category(dt: &Datatype) -> PyResult<&'static str> {
|
||||
match dt {
|
||||
Datatype::FixedPoint { .. } => Ok("int"),
|
||||
Datatype::FloatingPoint { .. } => Ok("float"),
|
||||
Datatype::Enumeration {
|
||||
base_type, members, ..
|
||||
} => {
|
||||
let is_bool = base_type.type_size() == 1
|
||||
&& members.len() == 2
|
||||
&& members
|
||||
.iter()
|
||||
.any(|m| m.name == "FALSE" && m.value.first() == Some(&0))
|
||||
&& members
|
||||
.iter()
|
||||
.any(|m| m.name == "TRUE" && m.value.first() == Some(&1));
|
||||
Ok(if is_bool { "bool" } else { "enum" })
|
||||
}
|
||||
Datatype::String {
|
||||
padding: StringPadding::NullPad,
|
||||
..
|
||||
} => Ok("string"),
|
||||
Datatype::String { padding, .. } => Err(not_implemented(format!(
|
||||
"writing fixed-length strings padded {padding:?} (libhdf5 converts them \
|
||||
differently from numpy)"
|
||||
))),
|
||||
Datatype::Compound { size, members } => {
|
||||
if is_complex(*size, members) {
|
||||
return Ok("complex");
|
||||
}
|
||||
check_exact(dt)?;
|
||||
Ok("exact")
|
||||
}
|
||||
Datatype::Opaque { .. } => Ok("exact"),
|
||||
Datatype::Array { .. } => Err(not_implemented("writing HDF5 array-type elements")),
|
||||
Datatype::VariableLength { .. } => Err(not_implemented("writing variable-length data")),
|
||||
Datatype::Reference { .. } => Err(not_implemented("writing references")),
|
||||
Datatype::BitField { .. } => Err(not_implemented("writing bitfields")),
|
||||
Datatype::Time { .. } => Err(not_implemented("writing time values")),
|
||||
}
|
||||
}
|
||||
|
||||
/// h5py's complex numbers: a compound of two identical floats `r`, `i`.
|
||||
fn is_complex(size: u32, members: &[CompoundMember]) -> bool {
|
||||
matches!(members, [r, i] if r.name == "r" && i.name == "i"
|
||||
&& r.datatype == i.datatype
|
||||
&& matches!(r.datatype, Datatype::FloatingPoint { size: fs, .. }
|
||||
if r.byte_offset == 0 && i.byte_offset == u64::from(fs) && size == 2 * fs))
|
||||
}
|
||||
|
||||
/// Compound members written byte for byte from the same numpy dtype: fine
|
||||
/// unless libhdf5 would convert them on the way (strings padded other than
|
||||
/// with NULs), or the editor cannot write them at all.
|
||||
fn check_exact(dt: &Datatype) -> PyResult<()> {
|
||||
match dt {
|
||||
Datatype::Compound { members, .. } => {
|
||||
members.iter().try_for_each(|m| check_exact(&m.datatype))
|
||||
}
|
||||
Datatype::Array { base_type, .. } => check_exact(base_type),
|
||||
Datatype::String {
|
||||
padding: StringPadding::NullPad,
|
||||
..
|
||||
}
|
||||
| Datatype::FixedPoint { .. }
|
||||
| Datatype::FloatingPoint { .. }
|
||||
| Datatype::Enumeration { .. }
|
||||
| Datatype::Opaque { .. }
|
||||
| Datatype::BitField { .. } => Ok(()),
|
||||
Datatype::String { .. } => Err(not_implemented(
|
||||
"writing compounds with strings not padded with NULs",
|
||||
)),
|
||||
Datatype::VariableLength { .. } => Err(not_implemented(
|
||||
"writing compounds with variable-length members",
|
||||
)),
|
||||
Datatype::Reference { .. } => Err(not_implemented("writing references")),
|
||||
Datatype::Time { .. } => Err(not_implemented("writing time values")),
|
||||
}
|
||||
}
|
||||
|
||||
/// Largest point selection an index-list write builds (one coordinate
|
||||
/// vector per element).
|
||||
const MAX_POINTS: usize = 1 << 22;
|
||||
|
||||
/// The selection `plan` writes, whose elements are numbered as the value's
|
||||
/// (row-major over the selection's shape).
|
||||
pub(crate) fn selection(plan: &Plan, dims: &[u64]) -> PyResult<Selection> {
|
||||
if plan.axes.is_empty() {
|
||||
return Ok(Selection::All);
|
||||
}
|
||||
if plan.list_axis().is_none() {
|
||||
let (reads, _) = plan.reads(dims, None, 1);
|
||||
return match <[_; 1]>::try_from(reads) {
|
||||
Ok([read]) => Ok(read.sel),
|
||||
Err(_) => Err(PyTypeError::new_err("internal error: several hyperslabs")),
|
||||
};
|
||||
}
|
||||
// An index list: the points, in the value's order.
|
||||
let per_axis: Vec<Vec<u64>> = plan
|
||||
.axes
|
||||
.iter()
|
||||
.map(|a| match a {
|
||||
Axis::Index(i) => vec![*i],
|
||||
Axis::Slice { start, step, count } => (0..*count).map(|k| start + k * step).collect(),
|
||||
Axis::List(v) => v.clone(),
|
||||
})
|
||||
.collect();
|
||||
let n = per_axis
|
||||
.iter()
|
||||
.try_fold(1usize, |acc, v| acc.checked_mul(v.len()))
|
||||
.filter(|&n| n <= MAX_POINTS)
|
||||
.ok_or_else(|| {
|
||||
not_implemented(format!(
|
||||
"an index-list write of more than {MAX_POINTS} elements (write it in slices)"
|
||||
))
|
||||
})?;
|
||||
let mut points = Vec::with_capacity(n);
|
||||
let mut at = vec![0usize; per_axis.len()];
|
||||
for _ in 0..n {
|
||||
points.push(at.iter().zip(&per_axis).map(|(&i, v)| v[i]).collect());
|
||||
for d in (0..at.len()).rev() {
|
||||
at[d] += 1;
|
||||
if at[d] < per_axis[d].len() {
|
||||
break;
|
||||
}
|
||||
at[d] = 0;
|
||||
}
|
||||
}
|
||||
Ok(Selection::Points(points))
|
||||
}
|
||||
|
||||
/// The bytes to write for `value` under `plan`, in the dataset's dtype.
|
||||
pub(crate) fn dataset_bytes(
|
||||
py: Python<'_>,
|
||||
value: &Bound<'_, PyAny>,
|
||||
dtype: &Bound<'_, PyAny>,
|
||||
category: &str,
|
||||
plan: &Plan,
|
||||
chunks: Option<&[u64]>,
|
||||
) -> PyResult<Vec<u8>> {
|
||||
let shape = PyTuple::new(py, plan.out_shape())?;
|
||||
let fancy = plan.list_axis().is_some();
|
||||
let chunk_elems = chunks.map_or(0, |c| c.iter().fold(1u64, |a, &d| a.saturating_mul(d)));
|
||||
let bytes = helpers(py)?.call_method1(
|
||||
"dataset_values",
|
||||
(value, dtype, category, shape, fancy, chunk_elems),
|
||||
)?;
|
||||
Ok(bytes.cast::<PyBytes>()?.as_bytes().to_vec())
|
||||
}
|
||||
|
||||
fn ieee_float(size: u32, byte_order: DatatypeByteOrder) -> Option<Datatype> {
|
||||
let (exponent_location, exponent_size, mantissa_size, exponent_bias) = match size {
|
||||
2 => (10, 5, 10, 15),
|
||||
4 => (23, 8, 23, 127),
|
||||
8 => (52, 11, 52, 1023),
|
||||
_ => return None,
|
||||
};
|
||||
Some(Datatype::FloatingPoint {
|
||||
size,
|
||||
byte_order,
|
||||
bit_offset: 0,
|
||||
bit_precision: (size * 8) as u16,
|
||||
exponent_location,
|
||||
exponent_size,
|
||||
mantissa_location: 0,
|
||||
mantissa_size,
|
||||
exponent_bias,
|
||||
})
|
||||
}
|
||||
|
||||
/// The HDF5 datatype h5py writes for a numpy dtype string (`'<i4'`,
|
||||
/// `'|b1'`, `'>f8'`, `'<c16'`, `'|S5'`).
|
||||
fn datatype_of(dtype: &str) -> Option<Datatype> {
|
||||
let order = match dtype.as_bytes().first()? {
|
||||
b'<' | b'|' | b'=' => DatatypeByteOrder::LittleEndian,
|
||||
b'>' => DatatypeByteOrder::BigEndian,
|
||||
_ => return None,
|
||||
};
|
||||
let kind = dtype.as_bytes().get(1)?;
|
||||
let size: u32 = dtype.get(2..)?.parse().ok()?;
|
||||
match kind {
|
||||
b'b' if size == 1 => Some(Datatype::Enumeration {
|
||||
size: 1,
|
||||
base_type: Box::new(Datatype::FixedPoint {
|
||||
size: 1,
|
||||
byte_order: DatatypeByteOrder::LittleEndian,
|
||||
signed: true,
|
||||
bit_offset: 0,
|
||||
bit_precision: 8,
|
||||
}),
|
||||
members: vec![
|
||||
EnumMember {
|
||||
name: "FALSE".into(),
|
||||
value: vec![0],
|
||||
},
|
||||
EnumMember {
|
||||
name: "TRUE".into(),
|
||||
value: vec![1],
|
||||
},
|
||||
],
|
||||
}),
|
||||
b'i' | b'u' if matches!(size, 1 | 2 | 4 | 8) => Some(Datatype::FixedPoint {
|
||||
size,
|
||||
byte_order: order,
|
||||
signed: *kind == b'i',
|
||||
bit_offset: 0,
|
||||
bit_precision: (size * 8) as u16,
|
||||
}),
|
||||
b'f' => ieee_float(size, order),
|
||||
b'c' => {
|
||||
let part = ieee_float(size / 2, order)?;
|
||||
Some(Datatype::Compound {
|
||||
size,
|
||||
members: vec![
|
||||
CompoundMember {
|
||||
name: "r".into(),
|
||||
byte_offset: 0,
|
||||
datatype: part.clone(),
|
||||
},
|
||||
CompoundMember {
|
||||
name: "i".into(),
|
||||
byte_offset: u64::from(size / 2),
|
||||
datatype: part,
|
||||
},
|
||||
],
|
||||
})
|
||||
}
|
||||
b'S' if size > 0 => Some(Datatype::String {
|
||||
size,
|
||||
padding: StringPadding::NullPad,
|
||||
charset: CharacterSet::Ascii,
|
||||
}),
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// An attribute value as h5py would store it (`attrs[name] = value`, or
|
||||
/// `attrs.create(name, data, shape, dtype)`), except that `str` data is
|
||||
/// stored as fixed-length UTF-8 strings (h5py stores variable-length ones,
|
||||
/// which the editor cannot write).
|
||||
pub(crate) fn attr_value(
|
||||
py: Python<'_>,
|
||||
value: &Bound<'_, PyAny>,
|
||||
dtype: Option<&Bound<'_, PyAny>>,
|
||||
shape: Option<&Bound<'_, PyAny>>,
|
||||
) -> PyResult<AttrValue> {
|
||||
if value.is_instance_of::<crate::PyEmpty>() {
|
||||
return Err(not_implemented(
|
||||
"writing an empty (null dataspace) attribute",
|
||||
));
|
||||
}
|
||||
let (kind, dt, dims, data): (String, String, Vec<u64>, Vec<u8>) = helpers(py)?
|
||||
.call_method1("attr_value", (value, dtype, shape))?
|
||||
.extract()?;
|
||||
let datatype = if kind == "str" {
|
||||
let size: u32 = dt
|
||||
.parse()
|
||||
.map_err(|_| PyTypeError::new_err("bad string size"))?;
|
||||
Datatype::String {
|
||||
size,
|
||||
padding: StringPadding::NullPad,
|
||||
charset: CharacterSet::Utf8,
|
||||
}
|
||||
} else {
|
||||
datatype_of(&dt).ok_or_else(|| not_implemented(format!("an attribute of dtype {dt}")))?
|
||||
};
|
||||
Ok(AttrValue::Raw {
|
||||
datatype,
|
||||
shape: dims,
|
||||
data,
|
||||
})
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn numpy_dtypes_map_to_h5py_types() {
|
||||
assert!(matches!(
|
||||
datatype_of("<i4"),
|
||||
Some(Datatype::FixedPoint {
|
||||
size: 4,
|
||||
signed: true,
|
||||
byte_order: DatatypeByteOrder::LittleEndian,
|
||||
..
|
||||
})
|
||||
));
|
||||
assert!(matches!(
|
||||
datatype_of(">u2"),
|
||||
Some(Datatype::FixedPoint {
|
||||
size: 2,
|
||||
signed: false,
|
||||
byte_order: DatatypeByteOrder::BigEndian,
|
||||
..
|
||||
})
|
||||
));
|
||||
assert!(matches!(
|
||||
datatype_of("<f2"),
|
||||
Some(Datatype::FloatingPoint { size: 2, .. })
|
||||
));
|
||||
let c = datatype_of("<c16").unwrap();
|
||||
assert!(matches!(&c, Datatype::Compound { size: 16, members } if is_complex(16, members)));
|
||||
assert_eq!(category(&c).unwrap(), "complex");
|
||||
let b = datatype_of("|b1").unwrap();
|
||||
assert_eq!(category(&b).unwrap(), "bool");
|
||||
assert!(matches!(
|
||||
datatype_of("|S5"),
|
||||
Some(Datatype::String { size: 5, .. })
|
||||
));
|
||||
assert!(datatype_of("<f16").is_none());
|
||||
assert!(datatype_of("<M8").is_none());
|
||||
assert!(datatype_of("|S0").is_none());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn list_writes_become_points_in_value_order() {
|
||||
let plan = Plan {
|
||||
axes: vec![
|
||||
Axis::List(vec![1, 4]),
|
||||
Axis::Slice {
|
||||
start: 0,
|
||||
step: 2,
|
||||
count: 2,
|
||||
},
|
||||
Axis::Index(3),
|
||||
],
|
||||
fields: vec![],
|
||||
scalar: false,
|
||||
};
|
||||
let sel = selection(&plan, &[5, 4, 4]).unwrap();
|
||||
assert_eq!(
|
||||
sel,
|
||||
Selection::Points(vec![
|
||||
vec![1, 0, 3],
|
||||
vec![1, 2, 3],
|
||||
vec![4, 0, 3],
|
||||
vec![4, 2, 3]
|
||||
])
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,175 @@
|
||||
"""Values for in-place writes (clawhdf5.File(path, 'r+')), prepared the way
|
||||
h5py prepares them, so `ds[key] = value` stores what h5py would store.
|
||||
|
||||
Loaded by the extension module (src/edit.rs); not a public API.
|
||||
|
||||
h5py converts in two ways, and so does this module:
|
||||
|
||||
- a value that is not a numpy array (a list, a Python or numpy scalar) is
|
||||
converted by numpy straight to the dataset's dtype
|
||||
(`numpy.asarray(value, dtype=ds.dtype)`), with numpy's rules and errors;
|
||||
- a numpy array is converted by libhdf5, whose numeric conversions clip to
|
||||
the target's range instead of wrapping: integers saturate, floats are
|
||||
truncated toward zero and clipped, a double too large for a float becomes
|
||||
infinity. That is what `_convert_array` reproduces. Where libhdf5 has no
|
||||
meaningful answer — NaN into an integer, for which it writes a different
|
||||
arbitrary value per type — this raises ValueError instead of guessing.
|
||||
"""
|
||||
|
||||
import numpy as np
|
||||
|
||||
|
||||
def _no_path(src, dst):
|
||||
return TypeError(f"No conversion path for dtype: {src!r} -> {dst!r}")
|
||||
|
||||
|
||||
def _to_int(arr, dtype):
|
||||
"""Integer target: libhdf5's saturating conversion."""
|
||||
info = np.iinfo(dtype)
|
||||
kind = arr.dtype.kind
|
||||
if kind == "b":
|
||||
return arr.astype(dtype)
|
||||
if kind in "iu":
|
||||
src = np.iinfo(arr.dtype)
|
||||
lo = max(info.min, src.min)
|
||||
hi = min(info.max, src.max)
|
||||
clipped = np.clip(arr, np.array(lo, arr.dtype), np.array(hi, arr.dtype))
|
||||
return clipped.astype(dtype)
|
||||
if kind == "f":
|
||||
if np.isnan(arr).any():
|
||||
raise ValueError(
|
||||
"cannot write NaN to an integer dataset (libhdf5 would store an arbitrary value)"
|
||||
)
|
||||
t = np.trunc(arr.astype(np.float64))
|
||||
# info.max + 1 and info.min are powers of two: exact as floats.
|
||||
over = t >= float(info.max + 1)
|
||||
under = t < float(info.min)
|
||||
out = np.where(over | under, 0.0, t).astype(dtype)
|
||||
out[over] = info.max
|
||||
out[under] = info.min
|
||||
return out
|
||||
raise _no_path(arr.dtype, dtype)
|
||||
|
||||
|
||||
def _convert_array(arr, dtype, category):
|
||||
kind = arr.dtype.kind
|
||||
if category in ("int", "enum"):
|
||||
if arr.dtype == dtype and kind in "iu":
|
||||
return arr
|
||||
if category == "enum" and kind not in "iu":
|
||||
raise _no_path(arr.dtype, dtype)
|
||||
return _to_int(arr, dtype)
|
||||
if category == "bool":
|
||||
if kind == "b":
|
||||
return arr.astype(dtype)
|
||||
if kind in "iu":
|
||||
# h5py's bool is an enum over int8; libhdf5 converts integers
|
||||
# into it by value (saturating), not to FALSE/TRUE, so 3 is
|
||||
# stored as 3. Keep those bytes: a view, not a cast.
|
||||
return _to_int(arr, np.dtype("i1")).view(dtype)
|
||||
raise _no_path(arr.dtype, dtype)
|
||||
if category == "float":
|
||||
if kind not in "biuf":
|
||||
raise _no_path(arr.dtype, dtype)
|
||||
with np.errstate(over="ignore", invalid="ignore"):
|
||||
return arr.astype(dtype)
|
||||
if category == "complex":
|
||||
if kind != "c":
|
||||
raise _no_path(arr.dtype, dtype)
|
||||
with np.errstate(over="ignore", invalid="ignore"):
|
||||
return arr.astype(dtype)
|
||||
if category == "string":
|
||||
if kind != "S":
|
||||
raise _no_path(arr.dtype, dtype)
|
||||
return arr.astype(dtype)
|
||||
# "exact": compound and opaque types, written only from the same dtype.
|
||||
if arr.dtype == dtype:
|
||||
return arr
|
||||
raise _no_path(arr.dtype, dtype)
|
||||
|
||||
|
||||
def _convert_other(value, dtype, category):
|
||||
if category == "string":
|
||||
items = np.asarray(value, dtype=object)
|
||||
if any(isinstance(x, str) for x in items.flat):
|
||||
meta = dtype.metadata or {}
|
||||
if meta.get("h5py_encoding") == "utf-8":
|
||||
enc = [x.encode("utf-8") if isinstance(x, str) else x for x in items.flat]
|
||||
return np.array(enc, dtype=dtype).reshape(items.shape)
|
||||
return np.asarray(value, dtype=dtype)
|
||||
|
||||
|
||||
def _broadcast(arr, shape, fancy, chunk_elems):
|
||||
"""h5py's broadcasting: numpy's rules against the selection's shape
|
||||
(extra leading length-1 axes allowed) for slices and integers. For an
|
||||
index list, the exact shape; a scalar only where h5py expands it to the
|
||||
whole selection (a chunked dataset whose chunk holds at least as many
|
||||
elements as the selection)."""
|
||||
if arr.shape == shape:
|
||||
return arr
|
||||
if fancy:
|
||||
size = int(np.prod(shape))
|
||||
if arr.ndim == 0 and ((chunk_elems > 0 and size <= chunk_elems) or len(shape) == 1):
|
||||
return np.broadcast_to(arr, shape)
|
||||
raise TypeError("Broadcasting is not supported for complex selections")
|
||||
if arr.ndim == 0:
|
||||
return np.broadcast_to(arr, shape)
|
||||
err = TypeError(f"Can't broadcast {arr.shape} -> {shape}")
|
||||
src = arr.shape
|
||||
while len(src) > len(shape) and src[0] == 1:
|
||||
src = src[1:]
|
||||
if len(src) > len(shape):
|
||||
raise err
|
||||
try:
|
||||
return np.broadcast_to(arr.reshape(src), shape)
|
||||
except ValueError:
|
||||
raise err from None
|
||||
|
||||
|
||||
def dataset_values(value, dtype, category, shape, fancy, chunk_elems):
|
||||
"""The bytes to write for `value` under a selection of `shape`, as a
|
||||
C-ordered array of the dataset's dtype."""
|
||||
if isinstance(value, np.ndarray):
|
||||
arr = _convert_array(value, dtype, category)
|
||||
else:
|
||||
arr = _convert_other(value, dtype, category)
|
||||
arr = _broadcast(arr, tuple(shape), fancy, chunk_elems)
|
||||
return np.ascontiguousarray(arr, dtype=dtype).tobytes()
|
||||
|
||||
|
||||
def attr_value(value, dtype=None, shape=None):
|
||||
"""(kind, dtype string, shape, bytes) for an attribute value, h5py's
|
||||
`attrs[name] = value` / `attrs.create(name, data, shape, dtype)`:
|
||||
|
||||
- "str": `str` data (h5py would store a variable-length string; this
|
||||
stores a fixed-length UTF-8 string, which clawhdf5 can write);
|
||||
the dtype string is the byte length of the longest element;
|
||||
- "raw": a numeric, bool or bytes array, as numpy lays it out.
|
||||
"""
|
||||
if dtype is not None:
|
||||
arr = np.asarray(value, dtype=dtype, order="C")
|
||||
else:
|
||||
arr = np.asarray(value, order="C")
|
||||
if shape is not None:
|
||||
arr = arr.reshape(shape)
|
||||
kind = arr.dtype.kind
|
||||
if kind == "O":
|
||||
if arr.size and all(isinstance(x, str) for x in arr.flat):
|
||||
kind = "U"
|
||||
elif arr.size and all(isinstance(x, bytes) for x in arr.flat):
|
||||
arr = arr.astype(bytes)
|
||||
kind = "S"
|
||||
else:
|
||||
raise TypeError(
|
||||
f"clawhdf5 cannot write an attribute of Python objects ({value!r:.60})"
|
||||
)
|
||||
if kind == "U":
|
||||
enc = [str(x).encode("utf-8") for x in arr.flat]
|
||||
size = max([len(b) for b in enc] + [1])
|
||||
data = np.array(enc, dtype=f"S{size}").reshape(arr.shape)
|
||||
return ("str", str(size), arr.shape, data.tobytes())
|
||||
if kind in "biufcS":
|
||||
return ("raw", arr.dtype.str, arr.shape, np.ascontiguousarray(arr).tobytes())
|
||||
raise NotImplementedError(
|
||||
f"clawhdf5 cannot write an attribute of dtype {arr.dtype} in place"
|
||||
)
|
||||
+228
-32
@@ -1,13 +1,17 @@
|
||||
//! PyFile — the main entry point for opening and creating HDF5 files.
|
||||
|
||||
use std::collections::HashMap;
|
||||
use std::path::PathBuf;
|
||||
use std::sync::{Arc, Mutex};
|
||||
use std::time::Duration;
|
||||
|
||||
use pyo3::exceptions::{PyNotImplementedError, PyValueError};
|
||||
use pyo3::prelude::*;
|
||||
use pyo3::types::PyList;
|
||||
use pyo3::types::{PyDict, PyList};
|
||||
|
||||
use crate::attrs::PyAttrs;
|
||||
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
|
||||
use crate::handle::Handle;
|
||||
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
|
||||
|
||||
/// Internal state for write mode.
|
||||
@@ -23,8 +27,10 @@ struct WriteState {
|
||||
/// Mirrors the h5py.File interface:
|
||||
///
|
||||
/// ```python
|
||||
/// # Reading
|
||||
/// # Reading, a local file or a URL (range requests, nothing downloaded
|
||||
/// # up front)
|
||||
/// f = clawhdf5.File('data.h5', 'r')
|
||||
/// f = clawhdf5.File('https://example.org/data.h5')
|
||||
/// ds = f['dataset']
|
||||
/// f.close()
|
||||
///
|
||||
@@ -44,49 +50,200 @@ enum FileInner {
|
||||
Write(WriteState),
|
||||
}
|
||||
|
||||
/// Whether `s` is a URL (`scheme://…`) rather than a path: the scheme is a
|
||||
/// letter followed by letters, digits, `+`, `-` or `.` (RFC 3986).
|
||||
fn is_url(s: &str) -> bool {
|
||||
let Some((scheme, _)) = s.split_once("://") else {
|
||||
return false;
|
||||
};
|
||||
let mut chars = scheme.chars();
|
||||
chars.next().is_some_and(|c| c.is_ascii_alphabetic())
|
||||
&& chars.all(|c| c.is_ascii_alphanumeric() || matches!(c, '+' | '-' | '.'))
|
||||
}
|
||||
|
||||
impl PyFile {
|
||||
fn from_handle(handle: Arc<Handle>, filename: String) -> Self {
|
||||
let root = handle.root;
|
||||
Self {
|
||||
inner: Some(FileInner::Read(ReadGroup::new(handle, String::new(), root))),
|
||||
filename,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[pymethods]
|
||||
impl PyFile {
|
||||
/// Open or create an HDF5 file.
|
||||
///
|
||||
/// Parameters:
|
||||
/// path: file path
|
||||
/// path: file path, or a URL (`http://`, `https://`, `s3://`, `gs://`,
|
||||
/// `az://`; which schemes work depends on how the wheel was built)
|
||||
/// to read the file remotely with default options (see `open_url`)
|
||||
/// mode: 'r' for read (default), 'w' for write
|
||||
#[new]
|
||||
#[pyo3(signature = (path, mode="r"))]
|
||||
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
|
||||
let filename = path.to_string();
|
||||
match mode {
|
||||
"r" => {
|
||||
let file = py.detach(|| {
|
||||
crate::no_panic(|| clawhdf5_rs::File::open(path).map_err(to_py_err))
|
||||
})?;
|
||||
Ok(Self {
|
||||
inner: Some(FileInner::Read(root_group(Arc::new(file)))),
|
||||
filename,
|
||||
})
|
||||
if is_url(path) {
|
||||
if mode != "r" {
|
||||
return Err(PyValueError::new_err(format!(
|
||||
"remote files are read-only: mode '{mode}' is not supported for a URL"
|
||||
)));
|
||||
}
|
||||
let handle = Handle::open_url(py, path, &clawhdf5_remote::Options::default())?;
|
||||
return Ok(Self::from_handle(handle, filename));
|
||||
}
|
||||
match mode {
|
||||
"r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)),
|
||||
"r+" => Ok(Self::from_handle(
|
||||
Handle::open_editable(py, path)?,
|
||||
filename,
|
||||
)),
|
||||
"a" if std::path::Path::new(path).exists() => Ok(Self::from_handle(
|
||||
Handle::open_editable(py, path)?,
|
||||
filename,
|
||||
)),
|
||||
"a" => Err(PyNotImplementedError::new_err(format!(
|
||||
"mode 'a' on {path}, which does not exist: clawhdf5 can only edit an existing \
|
||||
file in place; create a new one with mode 'w'"
|
||||
))),
|
||||
"w" => Ok(Self {
|
||||
filename,
|
||||
inner: Some(FileInner::Write(WriteState {
|
||||
path: PathBuf::from(path),
|
||||
// Absolute now: the file is written at close, possibly
|
||||
// after the working directory changed.
|
||||
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
|
||||
root_datasets: Vec::new(),
|
||||
root_attrs: Arc::new(Mutex::new(Vec::new())),
|
||||
groups: Vec::new(),
|
||||
})),
|
||||
}),
|
||||
other => Err(PyErr::new::<pyo3::exceptions::PyValueError, _>(format!(
|
||||
"unsupported mode '{other}'; expected 'r' or 'w'"
|
||||
other => Err(PyValueError::new_err(format!(
|
||||
"unsupported mode '{other}'; expected 'r', 'r+', 'a' or 'w'"
|
||||
))),
|
||||
}
|
||||
}
|
||||
|
||||
/// Open a remote file for reading, with options.
|
||||
///
|
||||
/// The file is read through a block cache with range requests: opening
|
||||
/// costs one request (it also fetches the first block), and a read
|
||||
/// fetches only the blocks it needs. The GIL is released while waiting
|
||||
/// on the network.
|
||||
///
|
||||
/// Parameters (all optional):
|
||||
/// block_size: bytes per cached block (default 1 MiB)
|
||||
/// cache_size: byte budget of the block cache (default 64 MiB)
|
||||
/// headers: dict of extra HTTP headers (e.g. Authorization), sent only
|
||||
/// to the URL's own origin
|
||||
/// retries: retries of a request that failed transiently (default 3)
|
||||
/// timeout: seconds to connect and receive response headers (default 30)
|
||||
/// allow_full_download: when the server ignores Range requests,
|
||||
/// download the whole file once instead of failing (default False)
|
||||
/// max_full_download: largest file such a download may fetch
|
||||
/// (default 1 GiB)
|
||||
/// require_validator: refuse a server that sends neither ETag nor
|
||||
/// Last-Modified (default False)
|
||||
/// max_redirects: redirects followed per request (default 5)
|
||||
/// max_parallel: requests of one read in flight at once (default 8)
|
||||
#[staticmethod]
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
#[pyo3(signature = (url, *, block_size=None, cache_size=None, headers=None, retries=None,
|
||||
timeout=None, allow_full_download=None, max_full_download=None,
|
||||
require_validator=None, max_redirects=None, max_parallel=None))]
|
||||
fn open_url(
|
||||
py: Python<'_>,
|
||||
url: &str,
|
||||
block_size: Option<u64>,
|
||||
cache_size: Option<u64>,
|
||||
headers: Option<HashMap<String, String>>,
|
||||
retries: Option<u32>,
|
||||
timeout: Option<f64>,
|
||||
allow_full_download: Option<bool>,
|
||||
max_full_download: Option<u64>,
|
||||
require_validator: Option<bool>,
|
||||
max_redirects: Option<u32>,
|
||||
max_parallel: Option<usize>,
|
||||
) -> PyResult<Self> {
|
||||
let mut options = clawhdf5_remote::Options::default();
|
||||
if let Some(b) = block_size {
|
||||
if b == 0 {
|
||||
return Err(PyValueError::new_err("block_size must be positive"));
|
||||
}
|
||||
options.cache.block_size = b;
|
||||
options.cache.coalesce_gap = b;
|
||||
// The opening request fetches the first block, not 1 MiB.
|
||||
options.http.first_request = b;
|
||||
}
|
||||
if let Some(c) = cache_size {
|
||||
options.cache.capacity = c;
|
||||
}
|
||||
let http = &mut options.http;
|
||||
if let Some(h) = headers {
|
||||
http.headers = h.into_iter().collect();
|
||||
}
|
||||
if let Some(r) = retries {
|
||||
http.retries = r;
|
||||
}
|
||||
if let Some(t) = timeout {
|
||||
if !(t.is_finite() && t > 0.0) {
|
||||
return Err(PyValueError::new_err("timeout must be a positive number"));
|
||||
}
|
||||
http.timeout = Duration::from_secs_f64(t);
|
||||
}
|
||||
if let Some(a) = allow_full_download {
|
||||
http.allow_full_download = a;
|
||||
}
|
||||
if let Some(m) = max_full_download {
|
||||
http.max_full_download = m;
|
||||
}
|
||||
if let Some(v) = require_validator {
|
||||
http.require_validator = v;
|
||||
}
|
||||
if let Some(r) = max_redirects {
|
||||
http.max_redirects = r;
|
||||
}
|
||||
if let Some(p) = max_parallel {
|
||||
if p == 0 {
|
||||
return Err(PyValueError::new_err("max_parallel must be positive"));
|
||||
}
|
||||
http.max_parallel = p;
|
||||
}
|
||||
let handle = Handle::open_url(py, url, &options)?;
|
||||
Ok(Self::from_handle(handle, url.to_string()))
|
||||
}
|
||||
|
||||
/// For a remote file, what its block cache has done so far (reads,
|
||||
/// hits, misses, requests, bytes fetched, ...); `None` for a local file.
|
||||
#[getter]
|
||||
fn remote_stats<'py>(&self, py: Python<'py>) -> PyResult<Option<Bound<'py, PyDict>>> {
|
||||
let Some(storage) = self.read_file()?.handle.remote_storage() else {
|
||||
return Ok(None);
|
||||
};
|
||||
let s = storage.stats();
|
||||
let d = PyDict::new(py);
|
||||
d.set_item("reads", s.reads)?;
|
||||
d.set_item("hits", s.hits)?;
|
||||
d.set_item("misses", s.misses)?;
|
||||
d.set_item("waits", s.waits)?;
|
||||
d.set_item("requests", s.requests)?;
|
||||
d.set_item("fetch_calls", s.fetch_calls)?;
|
||||
d.set_item("bytes_fetched", s.bytes_fetched)?;
|
||||
d.set_item("evictions", s.evictions)?;
|
||||
d.set_item("cached_bytes", s.cached_bytes)?;
|
||||
Ok(Some(d))
|
||||
}
|
||||
|
||||
/// Close the file. In write mode, this finalizes and writes the file.
|
||||
fn close(&mut self) -> PyResult<()> {
|
||||
let inner = self.inner.take().ok_or_else(|| {
|
||||
PyErr::new::<pyo3::exceptions::PyIOError, _>("file is already closed")
|
||||
})?;
|
||||
match inner {
|
||||
FileInner::Read(_) => Ok(()),
|
||||
FileInner::Read(root) => {
|
||||
root.handle.close();
|
||||
Ok(())
|
||||
}
|
||||
FileInner::Write(state) => finalize_write(state),
|
||||
}
|
||||
}
|
||||
@@ -121,7 +278,7 @@ impl PyFile {
|
||||
|
||||
/// List the names of all children in the root group.
|
||||
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||
let names = self.read_file()?.member_names()?;
|
||||
let names = self.read_file()?.member_names(py)?;
|
||||
Ok(PyList::new(py, names)?.into_any().unbind())
|
||||
}
|
||||
|
||||
@@ -139,8 +296,8 @@ impl PyFile {
|
||||
self.keys(py)?.call_method0(py, "__iter__")
|
||||
}
|
||||
|
||||
fn __len__(&self) -> PyResult<usize> {
|
||||
Ok(self.read_file()?.member_names()?.len())
|
||||
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||
Ok(self.read_file()?.member_names(py)?.len())
|
||||
}
|
||||
|
||||
/// The root group's name, `/`.
|
||||
@@ -149,7 +306,31 @@ impl PyFile {
|
||||
"/"
|
||||
}
|
||||
|
||||
/// The path the file was opened with.
|
||||
/// `'r'` for a file opened read-only (a local file or a URL), `'r+'`
|
||||
/// for one open for editing or writing, as h5py reports it.
|
||||
#[getter]
|
||||
fn mode(&self) -> PyResult<&'static str> {
|
||||
match &self.inner {
|
||||
Some(FileInner::Read(root)) if !root.handle.is_writable() => Ok("r"),
|
||||
Some(_) => Ok("r+"),
|
||||
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||
"file is closed",
|
||||
)),
|
||||
}
|
||||
}
|
||||
|
||||
/// Nothing to do: every edit is written and synced when it is made, and
|
||||
/// a file opened with 'w' is written on `close()`.
|
||||
fn flush(&self) {}
|
||||
|
||||
/// Deleting objects is not supported (h5py's `del f[name]`).
|
||||
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||
Err(PyNotImplementedError::new_err(format!(
|
||||
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
|
||||
)))
|
||||
}
|
||||
|
||||
/// The path (or URL) the file was opened with.
|
||||
#[getter]
|
||||
fn filename(&self) -> &str {
|
||||
&self.filename
|
||||
@@ -204,9 +385,9 @@ impl PyFile {
|
||||
/// Attribute access. In read mode, returns attributes of the root group.
|
||||
/// In write mode, returns a writable attrs handle.
|
||||
#[getter]
|
||||
fn attrs(&self) -> PyResult<PyAttrs> {
|
||||
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||
match self.inner.as_ref() {
|
||||
Some(FileInner::Read(root)) => root.attrs(),
|
||||
Some(FileInner::Read(root)) => root.attrs(py),
|
||||
Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))),
|
||||
None => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||
"file is closed",
|
||||
@@ -216,9 +397,10 @@ impl PyFile {
|
||||
|
||||
fn __repr__(&self) -> String {
|
||||
match &self.inner {
|
||||
Some(FileInner::Read(root)) => {
|
||||
format!("<HDF5 File (read, {} bytes)>", root.file.as_bytes().len())
|
||||
}
|
||||
Some(FileInner::Read(root)) => match root.handle.redacted_url() {
|
||||
Some(url) => format!("<HDF5 File (read, \"{url}\")>"),
|
||||
None => format!("<HDF5 File (read, \"{}\")>", self.filename),
|
||||
},
|
||||
Some(FileInner::Write(s)) => {
|
||||
format!("<HDF5 File (write, \"{}\")>", s.path.display())
|
||||
}
|
||||
@@ -226,8 +408,8 @@ impl PyFile {
|
||||
}
|
||||
}
|
||||
|
||||
fn __contains__(&self, key: &str) -> PyResult<bool> {
|
||||
Ok(self.read_file()?.contains(key))
|
||||
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||
self.read_file()?.contains(py, key)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -248,6 +430,13 @@ impl PyFile {
|
||||
fn write_state_mut(&mut self) -> PyResult<&mut WriteState> {
|
||||
match &mut self.inner {
|
||||
Some(FileInner::Write(s)) => Ok(s),
|
||||
Some(FileInner::Read(root)) if root.handle.is_writable() => {
|
||||
Err(PyNotImplementedError::new_err(
|
||||
"creating datasets or groups in an existing file is not supported by \
|
||||
clawhdf5's in-place editor (mode 'r+' changes values, shapes and \
|
||||
attributes)",
|
||||
))
|
||||
}
|
||||
Some(FileInner::Read(_)) => Err(PyErr::new::<pyo3::exceptions::PyIOError, _>(
|
||||
"cannot write to a file opened for reading",
|
||||
)),
|
||||
@@ -271,11 +460,6 @@ fn parse_compression(
|
||||
}
|
||||
}
|
||||
|
||||
fn root_group(file: Arc<clawhdf5_rs::File>) -> ReadGroup {
|
||||
let root = file.superblock().root_group_address;
|
||||
ReadGroup::new(file, String::new(), root)
|
||||
}
|
||||
|
||||
/// Build and write the HDF5 file from accumulated write state.
|
||||
fn finalize_write(state: WriteState) -> PyResult<()> {
|
||||
crate::no_panic(|| {
|
||||
@@ -309,6 +493,18 @@ fn finalize_write(state: WriteState) -> PyResult<()> {
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn urls_and_paths() {
|
||||
assert!(is_url("http://h/f.h5"));
|
||||
assert!(is_url("s3://bucket/key.h5"));
|
||||
assert!(is_url("git+https://x"));
|
||||
assert!(!is_url("data.h5"));
|
||||
assert!(!is_url("/tmp/a://b.h5"));
|
||||
assert!(!is_url("dir/x://y"));
|
||||
assert!(!is_url("1http://x"));
|
||||
assert!(!is_url("://x"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_gzip_compression() {
|
||||
assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6));
|
||||
|
||||
@@ -3,11 +3,12 @@
|
||||
use std::collections::HashMap;
|
||||
use std::sync::{Arc, Mutex, OnceLock};
|
||||
|
||||
use pyo3::exceptions::{PyIOError, PyKeyError, PyValueError};
|
||||
use pyo3::exceptions::{PyIOError, PyKeyError, PyNotImplementedError, PyOSError, PyValueError};
|
||||
use pyo3::prelude::*;
|
||||
use pyo3::types::PyList;
|
||||
|
||||
use crate::attrs::PyAttrs;
|
||||
use crate::handle::Handle;
|
||||
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node};
|
||||
|
||||
/// Shared state for a group being written.
|
||||
@@ -34,9 +35,9 @@ enum GroupInner {
|
||||
}
|
||||
|
||||
impl PyGroup {
|
||||
pub(crate) fn from_read(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
|
||||
pub(crate) fn from_read(handle: Arc<Handle>, path: String, addr: u64) -> Self {
|
||||
Self {
|
||||
inner: GroupInner::Read(ReadGroup::new(file, path, addr)),
|
||||
inner: GroupInner::Read(ReadGroup::new(handle, path, addr)),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -60,9 +61,10 @@ impl PyGroup {
|
||||
/// h5py). It keeps its own address and, once listed, its links, so looking
|
||||
/// up a child neither resolves the path from the root nor scans the group's
|
||||
/// links again: visiting every member of a large group is linear, not
|
||||
/// quadratic.
|
||||
/// quadratic. (Edits never add or remove links, so these stay valid in a
|
||||
/// file open for editing.)
|
||||
pub(crate) struct ReadGroup {
|
||||
pub file: Arc<clawhdf5_rs::File>,
|
||||
pub handle: Arc<Handle>,
|
||||
pub path: String,
|
||||
pub addr: u64,
|
||||
/// Link name -> object address (soft links resolved), filled on first use.
|
||||
@@ -72,9 +74,9 @@ pub(crate) struct ReadGroup {
|
||||
}
|
||||
|
||||
impl ReadGroup {
|
||||
pub(crate) fn new(file: Arc<clawhdf5_rs::File>, path: String, addr: u64) -> Self {
|
||||
pub(crate) fn new(handle: Arc<Handle>, path: String, addr: u64) -> Self {
|
||||
Self {
|
||||
file,
|
||||
handle,
|
||||
path,
|
||||
addr,
|
||||
links: OnceLock::new(),
|
||||
@@ -82,17 +84,14 @@ impl ReadGroup {
|
||||
}
|
||||
}
|
||||
|
||||
fn links(&self) -> PyResult<&HashMap<String, u64>> {
|
||||
fn links(&self, py: Python<'_>) -> PyResult<&HashMap<String, u64>> {
|
||||
if let Some(links) = self.links.get() {
|
||||
return Ok(links);
|
||||
}
|
||||
let entries = crate::no_panic(|| {
|
||||
clawhdf5_format::group_v2::resolve_group_children(
|
||||
self.file.as_bytes(),
|
||||
self.file.superblock(),
|
||||
self.addr,
|
||||
)
|
||||
.map_err(|e| PyValueError::new_err(format!("{}: {e}", node::name(&self.path))))
|
||||
let (addr, path) = (self.addr, &self.path);
|
||||
let entries = self.handle.with(py, |f| {
|
||||
clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr)
|
||||
.map_err(|e| node::format_err(path, e, PyValueError::new_err))
|
||||
})?;
|
||||
let map = entries
|
||||
.into_iter()
|
||||
@@ -102,7 +101,7 @@ impl ReadGroup {
|
||||
}
|
||||
|
||||
/// The path and address of `key` (a name, a relative or an absolute path).
|
||||
fn locate(&self, key: &str) -> PyResult<(String, u64)> {
|
||||
fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> {
|
||||
let path = node::join(&self.path, key);
|
||||
let rel = if self.path.is_empty() {
|
||||
Some(path.as_str())
|
||||
@@ -112,24 +111,24 @@ impl ReadGroup {
|
||||
path.strip_prefix(self.path.as_str())
|
||||
.and_then(|r| r.strip_prefix('/'))
|
||||
};
|
||||
let addr = match rel {
|
||||
// A direct child: the link table, when it has the name.
|
||||
Some(name) if !name.is_empty() && !name.contains('/') => {
|
||||
match self.links()?.get(name) {
|
||||
Some(&a) => a,
|
||||
None => node::resolve_from(&self.file, self.addr, name, &path)?,
|
||||
}
|
||||
}
|
||||
Some(rel) => node::resolve_from(&self.file, self.addr, rel, &path)?,
|
||||
None => node::address(&self.file, &path)?,
|
||||
};
|
||||
Ok((path, addr))
|
||||
// A direct child: the link table, when it has the name.
|
||||
if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/'))
|
||||
&& let Some(&a) = self.links(py)?.get(name)
|
||||
{
|
||||
return Ok((path, a));
|
||||
}
|
||||
let addr = self.addr;
|
||||
let found = self.handle.with(py, |f| match rel {
|
||||
Some(rel) => node::resolve_from(f, addr, rel, &path),
|
||||
None => node::address(f, &path),
|
||||
})?;
|
||||
Ok((path, found))
|
||||
}
|
||||
|
||||
/// `group[key]`.
|
||||
pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult<Py<PyAny>> {
|
||||
let (path, addr) = self.locate(key)?;
|
||||
node::open(py, &self.file, path, addr)
|
||||
let (path, addr) = self.locate(py, key)?;
|
||||
node::open(py, &self.handle, path, addr)
|
||||
}
|
||||
|
||||
/// `group.get(key, default)`.
|
||||
@@ -148,48 +147,57 @@ impl ReadGroup {
|
||||
}
|
||||
|
||||
/// Names of the group's datasets and subgroups, sorted (h5py's order).
|
||||
pub(crate) fn member_names(&self) -> PyResult<&[String]> {
|
||||
pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> {
|
||||
if let Some(m) = self.members.get() {
|
||||
return Ok(m);
|
||||
}
|
||||
let mut names = Vec::new();
|
||||
for (name, &addr) in self.links()? {
|
||||
let hdr = node::header_at(&self.file, addr, &node::join(&self.path, name))?;
|
||||
if matches!(
|
||||
node::kind(&hdr),
|
||||
Some(node::Kind::Dataset | node::Kind::Group)
|
||||
) {
|
||||
names.push(name.clone());
|
||||
let links = self.links(py)?;
|
||||
let path = &self.path;
|
||||
let mut names = self.handle.with(py, |f| {
|
||||
let mut names = Vec::new();
|
||||
for (name, &addr) in links {
|
||||
if matches!(
|
||||
node::kind_at(f, addr, &node::join(path, name))?,
|
||||
Some(node::Kind::Dataset | node::Kind::Group)
|
||||
) {
|
||||
names.push(name.clone());
|
||||
}
|
||||
}
|
||||
}
|
||||
Ok(names)
|
||||
})?;
|
||||
names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes()));
|
||||
Ok(self.members.get_or_init(|| names))
|
||||
}
|
||||
|
||||
pub(crate) fn contains(&self, key: &str) -> bool {
|
||||
self.locate(key)
|
||||
.and_then(|(path, addr)| node::header_at(&self.file, addr, &path))
|
||||
.ok()
|
||||
.and_then(|h| node::kind(&h))
|
||||
.is_some_and(|k| k != node::Kind::Datatype)
|
||||
/// `key in group`: whether `key` names a dataset or group. A failed
|
||||
/// read of the file (a network error) is raised, not `False`.
|
||||
pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||
let found = self
|
||||
.locate(py, key)
|
||||
.and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path)));
|
||||
match found {
|
||||
Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)),
|
||||
Err(e) if e.is_instance_of::<PyOSError>(py) => Err(e),
|
||||
Err(_) => Ok(false),
|
||||
}
|
||||
}
|
||||
|
||||
pub(crate) fn values(&self, py: Python<'_>) -> PyResult<Vec<Py<PyAny>>> {
|
||||
self.member_names()?
|
||||
self.member_names(py)?
|
||||
.iter()
|
||||
.map(|n| self.get_item(py, n))
|
||||
.collect()
|
||||
}
|
||||
|
||||
pub(crate) fn items(&self, py: Python<'_>) -> PyResult<Vec<(String, Py<PyAny>)>> {
|
||||
self.member_names()?
|
||||
self.member_names(py)?
|
||||
.iter()
|
||||
.map(|n| Ok((n.clone(), self.get_item(py, n)?)))
|
||||
.collect()
|
||||
}
|
||||
|
||||
pub(crate) fn attrs(&self) -> PyResult<PyAttrs> {
|
||||
PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path)
|
||||
pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||
PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -210,7 +218,7 @@ impl PyGroup {
|
||||
fn keys(&self, py: Python<'_>) -> PyResult<Py<PyAny>> {
|
||||
match &self.inner {
|
||||
GroupInner::Read(g) => {
|
||||
let list = PyList::new(py, g.member_names()?)?;
|
||||
let list = PyList::new(py, g.member_names(py)?)?;
|
||||
Ok(list.into_any().unbind())
|
||||
}
|
||||
GroupInner::Write(state) => {
|
||||
@@ -236,9 +244,9 @@ impl PyGroup {
|
||||
self.keys(py)?.call_method0(py, "__iter__")
|
||||
}
|
||||
|
||||
fn __len__(&self) -> PyResult<usize> {
|
||||
fn __len__(&self, py: Python<'_>) -> PyResult<usize> {
|
||||
match &self.inner {
|
||||
GroupInner::Read(g) => Ok(g.member_names()?.len()),
|
||||
GroupInner::Read(g) => Ok(g.member_names(py)?.len()),
|
||||
GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()),
|
||||
}
|
||||
}
|
||||
@@ -293,17 +301,28 @@ impl PyGroup {
|
||||
state.lock().unwrap().datasets.push(spec);
|
||||
Ok(())
|
||||
}
|
||||
GroupInner::Read { .. } => Err(PyIOError::new_err(
|
||||
GroupInner::Read(g) if g.handle.is_writable() => Err(PyNotImplementedError::new_err(
|
||||
"creating datasets or groups in an existing file is not supported by \
|
||||
clawhdf5's in-place editor (mode 'r+' changes values, shapes and attributes)",
|
||||
)),
|
||||
GroupInner::Read(_) => Err(PyIOError::new_err(
|
||||
"cannot create datasets on a read-only group",
|
||||
)),
|
||||
}
|
||||
}
|
||||
|
||||
/// Deleting objects is not supported (h5py's `del group[name]`).
|
||||
fn __delitem__(&self, key: &str) -> PyResult<()> {
|
||||
Err(PyNotImplementedError::new_err(format!(
|
||||
"cannot delete '{key}': deleting objects is not supported by clawhdf5"
|
||||
)))
|
||||
}
|
||||
|
||||
/// Attribute access.
|
||||
#[getter]
|
||||
fn attrs(&self) -> PyResult<PyAttrs> {
|
||||
fn attrs(&self, py: Python<'_>) -> PyResult<PyAttrs> {
|
||||
match &self.inner {
|
||||
GroupInner::Read(g) => g.attrs(),
|
||||
GroupInner::Read(g) => g.attrs(py),
|
||||
GroupInner::Write(state) => {
|
||||
let store = Arc::clone(&state.lock().unwrap().attrs);
|
||||
Ok(PyAttrs::from_write(store))
|
||||
@@ -311,10 +330,10 @@ impl PyGroup {
|
||||
}
|
||||
}
|
||||
|
||||
fn __repr__(&self) -> String {
|
||||
fn __repr__(&self, py: Python<'_>) -> String {
|
||||
match &self.inner {
|
||||
GroupInner::Read(g) => {
|
||||
let n = g.member_names().map_or(0, |m| m.len());
|
||||
let n = g.member_names(py).map_or(0, |m| m.len());
|
||||
format!("<HDF5 group \"{}\" ({n} members)>", node::name(&g.path))
|
||||
}
|
||||
GroupInner::Write(state) => {
|
||||
@@ -324,9 +343,9 @@ impl PyGroup {
|
||||
}
|
||||
}
|
||||
|
||||
fn __contains__(&self, key: &str) -> PyResult<bool> {
|
||||
fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult<bool> {
|
||||
match &self.inner {
|
||||
GroupInner::Read(g) => Ok(g.contains(key)),
|
||||
GroupInner::Read(g) => g.contains(py, key),
|
||||
GroupInner::Write(state) => {
|
||||
let guard = state.lock().unwrap();
|
||||
Ok(guard.datasets.iter().any(|d| d.name == key))
|
||||
@@ -357,31 +376,6 @@ pub(crate) fn finalize_write_group(
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn member_names_are_sorted() {
|
||||
let mut b = clawhdf5_rs::FileBuilder::new();
|
||||
b.create_dataset("zeta").with_f64_data(&[1.0]);
|
||||
b.create_dataset("alpha").with_f64_data(&[1.0]);
|
||||
let mut g = b.create_group("mid");
|
||||
g.create_dataset("x").with_f64_data(&[1.0]);
|
||||
let finished = g.finish();
|
||||
b.add_group(finished);
|
||||
let bytes = b.finish().unwrap();
|
||||
let file = Arc::new(clawhdf5_rs::File::from_bytes(bytes).unwrap());
|
||||
let root = file.superblock().root_group_address;
|
||||
let top = ReadGroup::new(Arc::clone(&file), String::new(), root);
|
||||
assert_eq!(top.member_names().unwrap(), ["alpha", "mid", "zeta"]);
|
||||
let (path, addr) = top.locate("mid").unwrap();
|
||||
assert_eq!(path, "mid");
|
||||
let mid = ReadGroup::new(Arc::clone(&file), path, addr);
|
||||
assert_eq!(mid.member_names().unwrap(), ["x"]);
|
||||
assert!(top.contains("mid/x"));
|
||||
assert!(mid.contains("/alpha"));
|
||||
assert!(mid.contains("x") && mid.contains("./x"));
|
||||
assert!(!top.contains("nope"));
|
||||
assert!(!mid.contains("alpha"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn finalize_group() {
|
||||
let state = WriteGroupState {
|
||||
|
||||
@@ -0,0 +1,230 @@
|
||||
//! The open file every object of a `File` shares.
|
||||
//!
|
||||
//! Every read goes through [`Handle::with`], which releases the GIL and
|
||||
//! parses through `File::storage()`, so the same code serves a local file
|
||||
//! (memory-mapped), a remote one (`clawhdf5-remote`: range requests through
|
||||
//! a block cache, so a network read never holds the GIL) and a file open
|
||||
//! for editing.
|
||||
//!
|
||||
//! A file opened with `'r+'` also holds a [`FileEditor`]. An edit takes the
|
||||
//! file's write lock, so no read runs while the file changes underneath it,
|
||||
//! and reopens the file afterwards, through the editor's own open file
|
||||
//! rather than its path: reads after an edit see the new bytes
|
||||
//! (a grown file, a new dataspace), never a stale mapping or chunk cache.
|
||||
//! Objects that cache something an edit can change compare
|
||||
//! [`Handle::generation`] with the value they cached it at.
|
||||
//!
|
||||
//! Lock discipline (no deadlock with the GIL): the file lock is only taken
|
||||
//! with the GIL released, and code that holds it never touches Python.
|
||||
|
||||
use std::sync::atomic::{AtomicU64, Ordering};
|
||||
use std::sync::{Arc, Mutex, PoisonError, RwLock};
|
||||
|
||||
use clawhdf5_rs::{File, FileEditor};
|
||||
use pyo3::exceptions::PyOSError;
|
||||
use pyo3::prelude::*;
|
||||
|
||||
use crate::{panic_text, to_py_err};
|
||||
|
||||
/// Where the file's bytes come from.
|
||||
pub(crate) enum Source {
|
||||
/// A local file (memory-mapped). Its path is not kept: nothing reopens
|
||||
/// it by path (see `open_editable`).
|
||||
Local,
|
||||
/// A URL, read through `clawhdf5-remote`'s block cache.
|
||||
Remote {
|
||||
url: String,
|
||||
storage: Arc<clawhdf5_remote::RemoteStorage>,
|
||||
},
|
||||
}
|
||||
|
||||
pub(crate) struct Handle {
|
||||
/// The file as last opened; `None` if reopening it after an edit failed
|
||||
/// (every read is then an error rather than a read of stale bytes).
|
||||
file: RwLock<Option<File>>,
|
||||
/// For `'r+'`: the editor, until the file is closed.
|
||||
editor: Option<Mutex<Option<FileEditor>>>,
|
||||
source: Source,
|
||||
/// Bumped by every edit.
|
||||
generation: AtomicU64,
|
||||
pub offset_size: u8,
|
||||
pub length_size: u8,
|
||||
pub root: u64,
|
||||
}
|
||||
|
||||
fn closed_after_failed_reopen() -> PyErr {
|
||||
PyOSError::new_err("the file could not be reopened after an edit; open it again")
|
||||
}
|
||||
|
||||
impl Handle {
|
||||
fn new(file: File, source: Source, editor: Option<FileEditor>) -> Arc<Self> {
|
||||
let sb = file.superblock();
|
||||
let (offset_size, length_size, root) =
|
||||
(sb.offset_size, sb.length_size, sb.root_group_address);
|
||||
Arc::new(Self {
|
||||
file: RwLock::new(Some(file)),
|
||||
editor: editor.map(|e| Mutex::new(Some(e))),
|
||||
source,
|
||||
generation: AtomicU64::new(0),
|
||||
offset_size,
|
||||
length_size,
|
||||
root,
|
||||
})
|
||||
}
|
||||
|
||||
/// A local file, read-only.
|
||||
pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
|
||||
let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?;
|
||||
Ok(Self::new(file, Source::Local, None))
|
||||
}
|
||||
|
||||
/// A local file, open for in-place editing (`'r+'`): the editor takes
|
||||
/// the file's exclusive lock and checks that it can edit the file, then
|
||||
/// the file is read through the editor's own file — never by path
|
||||
/// again, so a later `os.chdir` or a rename or replacement of the path
|
||||
/// cannot make reads (or the editor's plans) come from another file.
|
||||
pub(crate) fn open_editable(py: Python<'_>, path: &str) -> PyResult<Arc<Self>> {
|
||||
let (file, editor) = py.detach(|| {
|
||||
crate::no_panic(|| {
|
||||
let editor = FileEditor::open(path).map_err(to_py_err)?;
|
||||
let file = editor.reader().map_err(to_py_err)?;
|
||||
Ok((file, editor))
|
||||
})
|
||||
})?;
|
||||
Ok(Self::new(file, Source::Local, Some(editor)))
|
||||
}
|
||||
|
||||
/// A remote file (`http(s)://`, `s3://`, ...).
|
||||
pub(crate) fn open_url(
|
||||
py: Python<'_>,
|
||||
url: &str,
|
||||
options: &clawhdf5_remote::Options,
|
||||
) -> PyResult<Arc<Self>> {
|
||||
let (file, storage) = py.detach(|| {
|
||||
crate::no_panic(|| {
|
||||
let storage = clawhdf5_remote::storage_for_url(url, options).map_err(remote_err)?;
|
||||
let file = File::open_storage(storage.clone()).map_err(to_py_err)?;
|
||||
Ok((file, storage))
|
||||
})
|
||||
})?;
|
||||
Ok(Self::new(
|
||||
file,
|
||||
Source::Remote {
|
||||
url: url.to_string(),
|
||||
storage,
|
||||
},
|
||||
None,
|
||||
))
|
||||
}
|
||||
|
||||
/// Run `f` on the file with the GIL released (a remote read may wait
|
||||
/// on the network; other Python threads run meanwhile). `f` must not
|
||||
/// touch Python.
|
||||
pub(crate) fn with<R: Send>(
|
||||
&self,
|
||||
py: Python<'_>,
|
||||
f: impl FnOnce(&File) -> PyResult<R> + Send,
|
||||
) -> PyResult<R> {
|
||||
py.detach(|| self.with_detached(f))
|
||||
}
|
||||
|
||||
/// [`with`](Self::with) for code that already runs without the GIL.
|
||||
pub(crate) fn with_detached<R>(&self, f: impl FnOnce(&File) -> PyResult<R>) -> PyResult<R> {
|
||||
crate::no_panic(|| {
|
||||
let guard = self.file.read().unwrap_or_else(PoisonError::into_inner);
|
||||
let file = guard.as_ref().ok_or_else(closed_after_failed_reopen)?;
|
||||
f(file)
|
||||
})
|
||||
}
|
||||
|
||||
/// Edits so far: objects that cache something an edit can change (a
|
||||
/// dataset's shape, an object's attributes) re-read it when this moved.
|
||||
pub(crate) fn generation(&self) -> u64 {
|
||||
self.generation.load(Ordering::Acquire)
|
||||
}
|
||||
|
||||
/// Whether the file was opened for editing (`'r+'`), even if closed since.
|
||||
pub(crate) fn is_writable(&self) -> bool {
|
||||
self.editor.is_some()
|
||||
}
|
||||
|
||||
/// The remote file's block cache.
|
||||
pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> {
|
||||
match &self.source {
|
||||
Source::Remote { storage, .. } => Some(storage),
|
||||
Source::Local => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// The URL of a remote file, credentials and query values redacted.
|
||||
pub(crate) fn redacted_url(&self) -> Option<String> {
|
||||
match &self.source {
|
||||
Source::Remote { url, .. } => Some(clawhdf5_remote::redact_url(url)),
|
||||
Source::Local => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Release the editor, and with it the file's lock. Objects still
|
||||
/// open keep reading the file as it was last written; an edit through
|
||||
/// them is an error.
|
||||
pub(crate) fn close(&self) {
|
||||
if let Some(ed) = &self.editor {
|
||||
ed.lock().unwrap_or_else(PoisonError::into_inner).take();
|
||||
}
|
||||
}
|
||||
|
||||
/// Apply one edit with the GIL released. No read runs while it writes,
|
||||
/// and the file is reopened afterwards — also after a failed edit, since
|
||||
/// a commit that failed part-way may have changed the file.
|
||||
pub(crate) fn edit<R: Send>(
|
||||
&self,
|
||||
py: Python<'_>,
|
||||
f: impl FnOnce(&mut FileEditor) -> Result<R, clawhdf5_rs::Error> + Send,
|
||||
) -> PyResult<R> {
|
||||
let Some(editor) = &self.editor else {
|
||||
return Err(PyOSError::new_err(match self.source {
|
||||
Source::Remote { .. } => "remote files are read-only",
|
||||
Source::Local => "the file is open read-only; open it with mode 'r+' to change it",
|
||||
}));
|
||||
};
|
||||
if matches!(self.source, Source::Remote { .. }) {
|
||||
return Err(PyOSError::new_err("remote files are read-only"));
|
||||
}
|
||||
py.detach(|| {
|
||||
let mut ed = editor.lock().unwrap_or_else(PoisonError::into_inner);
|
||||
let ed = ed
|
||||
.as_mut()
|
||||
.ok_or_else(|| PyOSError::new_err("the file is closed"))?;
|
||||
let mut file = self.file.write().unwrap_or_else(PoisonError::into_inner);
|
||||
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| f(ed)));
|
||||
// Drop the old mapping and its chunk cache before reopening.
|
||||
*file = None;
|
||||
// Through the editor's file, not the path (see `open_editable`).
|
||||
let reopened = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| ed.reader()));
|
||||
self.generation.fetch_add(1, Ordering::AcqRel);
|
||||
match reopened {
|
||||
Ok(Ok(f)) => *file = Some(f),
|
||||
Ok(Err(e)) => return Err(to_py_err(e)),
|
||||
Err(_) => return Err(closed_after_failed_reopen()),
|
||||
}
|
||||
drop(file);
|
||||
match result {
|
||||
Ok(r) => r.map_err(to_py_err),
|
||||
Err(p) => Err(crate::InternalError::new_err(format!(
|
||||
"clawhdf5 internal error (please report it): {}",
|
||||
panic_text(&*p)
|
||||
))),
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// A `clawhdf5_remote::Error` as a Python exception: the network side
|
||||
/// (unreachable, a status, no range support, a changed file) is `OSError`,
|
||||
/// a file that is not HDF5 is what `to_py_err` makes of it.
|
||||
pub(crate) fn remote_err(e: clawhdf5_remote::Error) -> PyErr {
|
||||
match e {
|
||||
clawhdf5_remote::Error::Hdf5(e) => to_py_err(e),
|
||||
other => PyOSError::new_err(other.to_string()),
|
||||
}
|
||||
}
|
||||
@@ -7,13 +7,22 @@
|
||||
//!
|
||||
//! with clawhdf5.File('data.h5', 'r') as f:
|
||||
//! data = f['dataset_name'][:]
|
||||
//!
|
||||
//! with clawhdf5.File('http://host/data.h5') as f: # range requests
|
||||
//! block = f['dataset_name'][10:20]
|
||||
//!
|
||||
//! with clawhdf5.File('data.h5', 'r+') as f: # in-place edits
|
||||
//! f['dataset_name'][0] = 1.5
|
||||
//! f.attrs['note'] = 'edited'
|
||||
//! ```
|
||||
|
||||
mod attrs;
|
||||
mod convert;
|
||||
mod dataset;
|
||||
mod edit;
|
||||
mod file;
|
||||
mod group;
|
||||
mod handle;
|
||||
mod node;
|
||||
mod select;
|
||||
|
||||
@@ -63,7 +72,7 @@ fn _panic_for_test() -> PyResult<()> {
|
||||
/// Convert a `clawhdf5_rs::Error` into a `PyErr`.
|
||||
///
|
||||
/// Maps different error variants to more specific Python exception types:
|
||||
/// - I/O errors -> `PyIOError`
|
||||
/// - I/O errors, and failed reads of a remote file -> `PyIOError`/`PyOSError`
|
||||
/// - Format/parsing errors -> `PyValueError`
|
||||
/// - Missing dataset/path errors -> `PyKeyError`
|
||||
/// - Invalid arguments -> `PyValueError`
|
||||
@@ -73,6 +82,10 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr {
|
||||
use clawhdf5_rs::Error;
|
||||
match &e {
|
||||
Error::Io(_) => PyErr::new::<pyo3::exceptions::PyIOError, _>(e.to_string()),
|
||||
// A failed read of the storage: a network error on a remote file.
|
||||
Error::Format(clawhdf5_format::error::FormatError::Storage(_)) => {
|
||||
PyErr::new::<pyo3::exceptions::PyOSError, _>(e.to_string())
|
||||
}
|
||||
Error::Format(_) => PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()),
|
||||
Error::NotADataset(_) | Error::MissingMessage(_) => {
|
||||
PyErr::new::<pyo3::exceptions::PyKeyError, _>(e.to_string())
|
||||
|
||||
@@ -1,16 +1,24 @@
|
||||
//! Resolving paths to objects in a file opened for reading.
|
||||
//!
|
||||
//! Everything here parses through `File::storage()` (the `clawhdf5_format`
|
||||
//! `*_in` functions), never `File::as_bytes()`, so it works the same on a
|
||||
//! memory-mapped local file and on a remote one; and it runs inside
|
||||
//! `Handle::with`, without the GIL.
|
||||
|
||||
use std::sync::Arc;
|
||||
|
||||
use clawhdf5_format::attribute::AttributeMessage;
|
||||
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
|
||||
use clawhdf5_format::error::FormatError;
|
||||
use clawhdf5_format::message_type::MessageType;
|
||||
use clawhdf5_format::object_header::ObjectHeader;
|
||||
use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError};
|
||||
use clawhdf5_rs::File;
|
||||
use pyo3::exceptions::{PyKeyError, PyOSError, PyTypeError, PyValueError};
|
||||
use pyo3::prelude::*;
|
||||
|
||||
use crate::dataset::PyDataset;
|
||||
use crate::dataset::{DatasetMeta, PyDataset};
|
||||
use crate::group::PyGroup;
|
||||
use crate::handle::Handle;
|
||||
|
||||
/// Join `key` onto the group path `base` the way h5py does: an absolute key
|
||||
/// starts from the root, a relative one from `base`. Paths are kept without
|
||||
@@ -33,42 +41,46 @@ pub(crate) fn name(path: &str) -> String {
|
||||
format!("/{path}")
|
||||
}
|
||||
|
||||
/// A format error met at `path`: a failed read of the storage (a network
|
||||
/// error on a remote file) is an `OSError`, anything else `other(message)`.
|
||||
pub(crate) fn format_err(path: &str, e: FormatError, other: fn(String) -> PyErr) -> PyErr {
|
||||
let msg = format!("{}: {e}", name(path));
|
||||
match e {
|
||||
FormatError::Storage(_) => PyOSError::new_err(msg),
|
||||
_ => other(msg),
|
||||
}
|
||||
}
|
||||
|
||||
fn value_err(msg: String) -> PyErr {
|
||||
PyValueError::new_err(msg)
|
||||
}
|
||||
|
||||
/// The address of the object at `path`, resolved from the root group.
|
||||
pub(crate) fn address(file: &clawhdf5_rs::File, path: &str) -> PyResult<u64> {
|
||||
pub(crate) fn address(file: &File, path: &str) -> PyResult<u64> {
|
||||
resolve_from(file, file.superblock().root_group_address, path, path)
|
||||
}
|
||||
|
||||
/// The address of `rel` resolved from the group at `group` (`full` is the
|
||||
/// resulting path, for the error message).
|
||||
pub(crate) fn resolve_from(
|
||||
file: &clawhdf5_rs::File,
|
||||
group: u64,
|
||||
rel: &str,
|
||||
full: &str,
|
||||
) -> PyResult<u64> {
|
||||
pub(crate) fn resolve_from(file: &File, group: u64, rel: &str, full: &str) -> PyResult<u64> {
|
||||
if rel.is_empty() {
|
||||
return Ok(group);
|
||||
}
|
||||
crate::no_panic(|| {
|
||||
clawhdf5_format::group_v2::resolve_path_from(file.as_bytes(), file.superblock(), group, rel)
|
||||
.map_err(|e| {
|
||||
PyKeyError::new_err(format!(
|
||||
"Unable to open object (object '{}' doesn't exist): {e}",
|
||||
name(full)
|
||||
))
|
||||
})
|
||||
})
|
||||
clawhdf5_format::group_v2::resolve_path_from_in(file.storage(), file.superblock(), group, rel)
|
||||
.map_err(|e| match e {
|
||||
FormatError::Storage(_) => format_err(full, e, value_err),
|
||||
e => PyKeyError::new_err(format!(
|
||||
"Unable to open object (object '{}' doesn't exist): {e}",
|
||||
name(full)
|
||||
)),
|
||||
})
|
||||
}
|
||||
|
||||
/// The object header at `addr` (the object at `path`).
|
||||
pub(crate) fn header_at(file: &clawhdf5_rs::File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
|
||||
crate::no_panic(|| {
|
||||
let sb = file.superblock();
|
||||
let at = usize::try_from(addr)
|
||||
.map_err(|_| PyValueError::new_err(format!("{}: address out of range", name(path))))?;
|
||||
ObjectHeader::parse(file.as_bytes(), at, sb.offset_size, sb.length_size)
|
||||
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))
|
||||
})
|
||||
pub(crate) fn header_at(file: &File, addr: u64, path: &str) -> PyResult<ObjectHeader> {
|
||||
let sb = file.superblock();
|
||||
ObjectHeader::parse_in(file.storage(), addr, sb.offset_size, sb.length_size)
|
||||
.map_err(|e| format_err(path, e, value_err))
|
||||
}
|
||||
|
||||
/// What an object header describes.
|
||||
@@ -96,29 +108,50 @@ pub(crate) fn kind(hdr: &ObjectHeader) -> Option<Kind> {
|
||||
}
|
||||
}
|
||||
|
||||
/// The kind of the object at `addr`, from its header.
|
||||
pub(crate) fn kind_at(file: &File, addr: u64, path: &str) -> PyResult<Option<Kind>> {
|
||||
Ok(kind(&header_at(file, addr, path)?))
|
||||
}
|
||||
|
||||
/// What opening an object found, read without the GIL.
|
||||
enum Found {
|
||||
Dataset(DatasetMeta),
|
||||
Group,
|
||||
Datatype,
|
||||
Other,
|
||||
}
|
||||
|
||||
/// Open the object at `addr` (whose path is `path`) as a `Dataset` or
|
||||
/// `Group`. Both keep the address, so later reads resolve nothing.
|
||||
pub(crate) fn open(
|
||||
py: Python<'_>,
|
||||
file: &Arc<clawhdf5_rs::File>,
|
||||
handle: &Arc<Handle>,
|
||||
path: String,
|
||||
addr: u64,
|
||||
) -> PyResult<Py<PyAny>> {
|
||||
let hdr = header_at(file, addr, &path)?;
|
||||
match kind(&hdr) {
|
||||
Some(Kind::Dataset) => Ok(PyDataset::open(py, Arc::clone(file), path, addr, &hdr)?
|
||||
let found = handle.with(py, |f| {
|
||||
let hdr = header_at(f, addr, &path)?;
|
||||
Ok(match kind(&hdr) {
|
||||
Some(Kind::Dataset) => Found::Dataset(DatasetMeta::load(f, addr, &hdr, &path)?),
|
||||
Some(Kind::Group) => Found::Group,
|
||||
Some(Kind::Datatype) => Found::Datatype,
|
||||
None => Found::Other,
|
||||
})
|
||||
})?;
|
||||
match found {
|
||||
Found::Dataset(meta) => Ok(PyDataset::new(py, Arc::clone(handle), path, addr, meta)
|
||||
.into_pyobject(py)?
|
||||
.into_any()
|
||||
.unbind()),
|
||||
Some(Kind::Group) => Ok(PyGroup::from_read(Arc::clone(file), path, addr)
|
||||
Found::Group => Ok(PyGroup::from_read(Arc::clone(handle), path, addr)
|
||||
.into_pyobject(py)?
|
||||
.into_any()
|
||||
.unbind()),
|
||||
Some(Kind::Datatype) => Err(PyTypeError::new_err(format!(
|
||||
Found::Datatype => Err(PyTypeError::new_err(format!(
|
||||
"{}: committed (named) datatypes are not supported by clawhdf5",
|
||||
name(&path)
|
||||
))),
|
||||
None => Err(PyValueError::new_err(format!(
|
||||
Found::Other => Err(PyValueError::new_err(format!(
|
||||
"{}: not a dataset, group or datatype",
|
||||
name(&path)
|
||||
))),
|
||||
@@ -126,32 +159,26 @@ pub(crate) fn open(
|
||||
}
|
||||
|
||||
/// The dataspace message of an object header.
|
||||
pub(crate) fn dataspace(file: &clawhdf5_rs::File, hdr: &ObjectHeader) -> PyResult<Dataspace> {
|
||||
crate::no_panic(|| {
|
||||
let sb = file.superblock();
|
||||
let msg = hdr
|
||||
.messages
|
||||
.iter()
|
||||
.find(|m| m.msg_type == MessageType::Dataspace)
|
||||
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
|
||||
let data = clawhdf5_format::shared_message::message_data(
|
||||
file.as_bytes(),
|
||||
msg,
|
||||
sb.offset_size,
|
||||
sb.length_size,
|
||||
)
|
||||
.map_err(|e| PyValueError::new_err(e.to_string()))?;
|
||||
Dataspace::parse(&data, sb.length_size).map_err(|e| PyValueError::new_err(e.to_string()))
|
||||
})
|
||||
pub(crate) fn dataspace(file: &File, hdr: &ObjectHeader, path: &str) -> PyResult<Dataspace> {
|
||||
let sb = file.superblock();
|
||||
let msg = hdr
|
||||
.messages
|
||||
.iter()
|
||||
.find(|m| m.msg_type == MessageType::Dataspace)
|
||||
.ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?;
|
||||
let data = clawhdf5_format::shared_message::message_data_in(
|
||||
file.storage(),
|
||||
msg,
|
||||
sb.offset_size,
|
||||
sb.length_size,
|
||||
)
|
||||
.map_err(|e| format_err(path, e, value_err))?;
|
||||
Dataspace::parse(&data, sb.length_size).map_err(|e| format_err(path, e, value_err))
|
||||
}
|
||||
|
||||
/// The chunk shape of a chunked dataset (one entry per dataset dimension),
|
||||
/// or `None` for other layouts or a layout message that does not parse.
|
||||
pub(crate) fn chunk_shape(
|
||||
file: &clawhdf5_rs::File,
|
||||
hdr: &ObjectHeader,
|
||||
rank: usize,
|
||||
) -> Option<Vec<u64>> {
|
||||
pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Option<Vec<u64>> {
|
||||
let sb = file.superblock();
|
||||
let msg = hdr
|
||||
.messages
|
||||
@@ -179,24 +206,18 @@ pub(crate) fn is_null(space: &Dataspace) -> bool {
|
||||
/// The attributes of the object at `addr` (whose path is `path`), sorted by
|
||||
/// name (h5py's order). Attributes whose messages cannot be parsed are left
|
||||
/// out, as the facade's `attrs()` does.
|
||||
pub(crate) fn attributes(
|
||||
file: &clawhdf5_rs::File,
|
||||
addr: u64,
|
||||
path: &str,
|
||||
) -> PyResult<Vec<AttributeMessage>> {
|
||||
pub(crate) fn attributes(file: &File, addr: u64, path: &str) -> PyResult<Vec<AttributeMessage>> {
|
||||
let hdr = header_at(file, addr, path)?;
|
||||
crate::no_panic(|| {
|
||||
let sb = file.superblock();
|
||||
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant(
|
||||
file.as_bytes(),
|
||||
&hdr,
|
||||
sb.offset_size,
|
||||
sb.length_size,
|
||||
)
|
||||
.map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))?;
|
||||
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
|
||||
Ok(attrs)
|
||||
})
|
||||
let sb = file.superblock();
|
||||
let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
|
||||
file.storage(),
|
||||
&hdr,
|
||||
sb.offset_size,
|
||||
sb.length_size,
|
||||
)
|
||||
.map_err(|e| format_err(path, e, value_err))?;
|
||||
attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes()));
|
||||
Ok(attrs)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
|
||||
@@ -7,11 +7,12 @@
|
||||
//! (negative from the end) drop their axis, slices must have a positive
|
||||
//! step, one `Ellipsis` fills the unmentioned axes, a single increasing list
|
||||
//! of integers may index one axis, and strings name compound fields.
|
||||
//! Everything else (`None`/`np.newaxis`, boolean masks, several index lists)
|
||||
//! is refused with the error h5py gives.
|
||||
//! Everything else (`None`/`np.newaxis`, several index lists) is refused
|
||||
//! with the error h5py gives; boolean masks, which h5py supports, raise
|
||||
//! `NotImplementedError`.
|
||||
|
||||
use clawhdf5_format::selection::Selection;
|
||||
use pyo3::exceptions::{PyIndexError, PyTypeError, PyValueError};
|
||||
use pyo3::exceptions::{PyIndexError, PyNotImplementedError, PyTypeError, PyValueError};
|
||||
use pyo3::prelude::*;
|
||||
use pyo3::types::{PyEllipsis, PySlice, PyString, PyTuple};
|
||||
|
||||
@@ -254,6 +255,17 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
|
||||
}
|
||||
}
|
||||
|
||||
// A mask of the dataset's whole shape (`ds[ds[()] > 0]`).
|
||||
if let [a] = args.as_slice() {
|
||||
let np = key.py().import("numpy")?;
|
||||
if a.is_instance(&np.getattr("ndarray")?)?
|
||||
&& a.getattr("dtype")?.getattr("kind")?.extract::<String>()? == "b"
|
||||
&& a.getattr("ndim")?.extract::<usize>()? > 1
|
||||
&& a.getattr("shape")?.extract::<Vec<u64>>()? == dims
|
||||
{
|
||||
return Err(mask_unsupported());
|
||||
}
|
||||
}
|
||||
if args.iter().any(|a| a.is_none()) {
|
||||
return Err(PyTypeError::new_err(
|
||||
"Indexing with None (or np.newaxis) is not supported",
|
||||
@@ -332,6 +344,10 @@ pub(crate) fn parse(key: &Bound<'_, PyAny>, dims: &[u64]) -> PyResult<Plan> {
|
||||
})
|
||||
}
|
||||
|
||||
fn mask_unsupported() -> PyErr {
|
||||
PyNotImplementedError::new_err("boolean mask indexing is not supported by clawhdf5")
|
||||
}
|
||||
|
||||
fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
|
||||
if a.is_none() {
|
||||
return Err(PyTypeError::new_err(
|
||||
@@ -379,8 +395,15 @@ fn parse_axis(py: Python<'_>, a: &Bound<'_, PyAny>, n: u64) -> PyResult<Axis> {
|
||||
let arr = np.call_method1("asarray", (a,))?;
|
||||
let kind: String = arr.getattr("dtype")?.getattr("kind")?.extract()?;
|
||||
if kind == "b" {
|
||||
// A mask along this axis (h5py supports them; clawhdf5 does
|
||||
// not, for reads or writes: an unsupported operation). A mask
|
||||
// of any other shape is a wrong key, as in h5py.
|
||||
let shape: Vec<u64> = arr.getattr("shape")?.extract()?;
|
||||
if shape == [n] {
|
||||
return Err(mask_unsupported());
|
||||
}
|
||||
return Err(PyTypeError::new_err(
|
||||
"Boolean mask indexing is not supported by clawhdf5",
|
||||
"Boolean indexing array has incompatible shape",
|
||||
));
|
||||
}
|
||||
let ndim: usize = arr.getattr("ndim")?.extract()?;
|
||||
|
||||
@@ -1,6 +1,10 @@
|
||||
"""Shared fixtures for the clawhdf5 Python binding tests."""
|
||||
|
||||
import os
|
||||
import re
|
||||
import threading
|
||||
import time
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
import pytest
|
||||
|
||||
@@ -16,3 +20,134 @@ def h5py():
|
||||
pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable")
|
||||
pytest.skip("h5py not installed")
|
||||
return mod
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# An HTTP server for remote reads
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
|
||||
|
||||
|
||||
class RangeServer:
|
||||
"""A static file server on 127.0.0.1, in a thread of this process, that
|
||||
answers `Range: bytes=a-b` with 206 and `Content-Range` (the way S3 and
|
||||
common web servers do), sends an ETag and honours `If-Match`.
|
||||
|
||||
- `ranges=False`: ignores `Range` and answers 200 with the whole file,
|
||||
like a server without range support.
|
||||
- `down` (set by `close()`): hang up on every request.
|
||||
- `delay`: seconds to wait before answering each request after the
|
||||
first `delay_after` ones (a slow network).
|
||||
- `log`: every request as `(method, path, range header)`.
|
||||
"""
|
||||
|
||||
def __init__(self, root, ranges=True):
|
||||
self.root = str(root)
|
||||
self.ranges = ranges
|
||||
self.delay = 0.0
|
||||
self.delay_after = 0
|
||||
self.down = False
|
||||
self.log = []
|
||||
self._lock = threading.Lock()
|
||||
server = self
|
||||
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
protocol_version = "HTTP/1.1"
|
||||
|
||||
def log_message(self, *args): # quiet
|
||||
pass
|
||||
|
||||
def do_HEAD(self):
|
||||
self._serve(body=False)
|
||||
|
||||
def do_GET(self):
|
||||
self._serve(body=True)
|
||||
|
||||
def _serve(self, body):
|
||||
with server._lock:
|
||||
server.log.append((self.command, self.path, self.headers.get("Range")))
|
||||
n = len(server.log)
|
||||
if server.delay and n > server.delay_after:
|
||||
time.sleep(server.delay)
|
||||
if server.down:
|
||||
# Hang up without an answer (open keep-alive
|
||||
# connections outlive shutdown(), so close() sets this).
|
||||
self.close_connection = True
|
||||
return
|
||||
path = os.path.join(server.root, self.path.lstrip("/").split("?")[0])
|
||||
if not os.path.isfile(path):
|
||||
self.send_response(404)
|
||||
self.send_header("Content-Length", "0")
|
||||
self.end_headers()
|
||||
return
|
||||
with open(path, "rb") as fh:
|
||||
data = fh.read()
|
||||
st = os.stat(path)
|
||||
etag = f'"{st.st_mtime_ns:x}-{st.st_size:x}"'
|
||||
want = self.headers.get("If-Match")
|
||||
if want is not None and want != etag and want != "*":
|
||||
self.send_response(412)
|
||||
self.send_header("Content-Length", "0")
|
||||
self.end_headers()
|
||||
return
|
||||
rng = self.headers.get("Range") if server.ranges else None
|
||||
m = _RANGE.match(rng.strip()) if rng else None
|
||||
if m and (m.group(1) or m.group(2)):
|
||||
size = len(data)
|
||||
if m.group(1):
|
||||
start = int(m.group(1))
|
||||
end = int(m.group(2)) if m.group(2) else size - 1
|
||||
else:
|
||||
start = max(0, size - int(m.group(2)))
|
||||
end = size - 1
|
||||
if start >= size:
|
||||
self.send_response(416)
|
||||
self.send_header("Content-Range", f"bytes */{size}")
|
||||
self.send_header("Content-Length", "0")
|
||||
self.end_headers()
|
||||
return
|
||||
end = min(end, size - 1)
|
||||
part = data[start : end + 1]
|
||||
self.send_response(206)
|
||||
self.send_header("Content-Range", f"bytes {start}-{end}/{size}")
|
||||
else:
|
||||
part = data
|
||||
self.send_response(200)
|
||||
if server.ranges:
|
||||
self.send_header("Accept-Ranges", "bytes")
|
||||
self.send_header("ETag", etag)
|
||||
self.send_header("Content-Length", str(len(part)))
|
||||
self.send_header("Content-Type", "application/x-hdf5")
|
||||
self.end_headers()
|
||||
if body:
|
||||
try:
|
||||
self.wfile.write(part)
|
||||
except (BrokenPipeError, ConnectionResetError):
|
||||
pass
|
||||
|
||||
self.httpd = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
|
||||
self.httpd.daemon_threads = True
|
||||
self.port = self.httpd.server_address[1]
|
||||
self.thread = threading.Thread(target=self.httpd.serve_forever, daemon=True)
|
||||
self.thread.start()
|
||||
|
||||
def url(self, name):
|
||||
return f"http://127.0.0.1:{self.port}/{name}"
|
||||
|
||||
def requests(self):
|
||||
with self._lock:
|
||||
return len(self.log)
|
||||
|
||||
def close(self):
|
||||
self.down = True
|
||||
self.httpd.shutdown()
|
||||
self.httpd.server_close()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def range_server(tmp_path):
|
||||
"""A range-capable server over `tmp_path`."""
|
||||
server = RangeServer(tmp_path)
|
||||
yield server
|
||||
server.close()
|
||||
|
||||
@@ -0,0 +1,940 @@
|
||||
"""In-place editing: clawhdf5.File(path, 'r+') against h5py.
|
||||
|
||||
Every edit is applied twice, to two copies of the same file: once through
|
||||
h5py (libhdf5) and once through clawhdf5 (FileEditor). After every edit both
|
||||
files are read back with h5py and must hold the same shapes, values and
|
||||
attributes; clawhdf5's own view must agree; when h5py refuses an edit,
|
||||
clawhdf5 must refuse it too and leave its file as it was. Files are written
|
||||
by h5py (libver earliest and latest, so every chunk index kind) and by
|
||||
clawhdf5; `h5dump` must read every result."""
|
||||
|
||||
import io
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import threading
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
|
||||
import clawhdf5
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Files
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
ENUM = {"RED": 0, "GREEN": 1, "BLUE": 7}
|
||||
|
||||
|
||||
def _h5py_file(h5py, path, libver):
|
||||
rng = np.random.default_rng(1)
|
||||
with h5py.File(path, "w", libver=libver) as f:
|
||||
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
|
||||
f.create_dataset("be_i2", data=np.arange(24, dtype=">i2").reshape(4, 6))
|
||||
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
|
||||
f.create_dataset("f2", data=rng.standard_normal(12).astype("<f2"))
|
||||
f.create_dataset("f4_2d", data=rng.standard_normal((8, 9)).astype("<f4"))
|
||||
f.create_dataset("f8_3d", data=rng.standard_normal((4, 5, 6)))
|
||||
f.create_dataset("u8", data=np.arange(10, dtype="<u8"))
|
||||
f.create_dataset("c8", data=(np.arange(6) + 1j * np.arange(6)).astype("<c8"))
|
||||
f.create_dataset("bool", data=np.array([True, False, True, True]))
|
||||
f.create_dataset("enum", data=np.array([0, 1, 7, 0], dtype="i1"),
|
||||
dtype=h5py.enum_dtype(ENUM, basetype="i1"))
|
||||
f.create_dataset("s5", data=np.array([b"ab", b"cdefg", b""], dtype="S5"))
|
||||
cmp_dt = np.dtype([("id", "<i4"), ("x", "<f8"), ("tag", "S3")])
|
||||
f.create_dataset("cmp", data=np.array([(i, i / 2, b"t%d" % i) for i in range(5)], dtype=cmp_dt))
|
||||
f.create_dataset("scalar", data=np.float64(3.5))
|
||||
# Chunked: fixed maxshape (v4 fixed array under latest), one
|
||||
# unlimited dimension (extensible array), two (v2 B-tree), one chunk.
|
||||
f.create_dataset("chunk_fixed", data=np.arange(100, dtype="<i8").reshape(10, 10), chunks=(3, 4))
|
||||
f.create_dataset("chunk_ext", data=rng.standard_normal((12, 7)), chunks=(5, 7), maxshape=(None, 7))
|
||||
f.create_dataset("chunk_bt2", data=np.arange(30, dtype="<i4").reshape(5, 6), chunks=(2, 2),
|
||||
maxshape=(None, None))
|
||||
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20), chunks=(6, 6),
|
||||
compression="gzip", maxshape=(40, 40))
|
||||
f.create_dataset("chunk_single", data=np.arange(12, dtype="<u2").reshape(3, 4), chunks=(3, 4),
|
||||
maxshape=(3, 4))
|
||||
f.create_dataset("chunk_fill", shape=(8,), dtype="<i4", chunks=(3,), maxshape=(20,), fillvalue=-1)
|
||||
f.create_dataset("vlen", data=["a", "bb"], dtype=h5py.string_dtype())
|
||||
# Compact layout (low-level API).
|
||||
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
|
||||
dcpl.set_layout(h5py.h5d.COMPACT)
|
||||
space = h5py.h5s.create_simple((7,))
|
||||
dsid = h5py.h5d.create(f.id, b"compact", h5py.h5t.STD_I32LE, space, dcpl=dcpl)
|
||||
dsid.write(h5py.h5s.ALL, h5py.h5s.ALL, np.arange(7, dtype="<i4"))
|
||||
g = f.create_group("grp")
|
||||
g.create_dataset("leaf", data=np.arange(5.0))
|
||||
g.attrs["units"] = "m"
|
||||
f.attrs["version"] = np.int32(1)
|
||||
|
||||
|
||||
def _clawhdf5_file(path):
|
||||
with clawhdf5.File(str(path), "w") as f:
|
||||
f.create_dataset("i4", data=np.arange(60, dtype="<i4").reshape(6, 10))
|
||||
f.create_dataset("f8", data=np.linspace(0, 1, 30).reshape(5, 6))
|
||||
f.create_dataset("chunk_gzip", data=np.arange(400, dtype="<f4").reshape(20, 20),
|
||||
chunks=[6, 6], compression="gzip")
|
||||
f.create_dataset("u1", data=np.arange(16, dtype="u1"))
|
||||
g = f.create_group("grp")
|
||||
g.create_dataset("leaf", data=np.arange(5.0))
|
||||
g.attrs["units"] = "m"
|
||||
f.attrs["version"] = 1
|
||||
|
||||
|
||||
# h5py's libver: "earliest" (v1 B-tree chunk indexes), "v114" (the 1.10+
|
||||
# indexes: fixed and extensible arrays, v2 B-trees, single chunk) and
|
||||
# "latest" (HDF5 2.0's newest format, which h5dump 1.14 cannot read).
|
||||
SOURCES = ["h5py-earliest", "h5py-v114", "h5py-latest", "clawhdf5"]
|
||||
|
||||
|
||||
def _make(h5py, tmp_path, source):
|
||||
base = tmp_path / f"base-{source}.h5"
|
||||
if source == "clawhdf5":
|
||||
_clawhdf5_file(base)
|
||||
else:
|
||||
_h5py_file(h5py, str(base), source.split("-")[1])
|
||||
theirs = tmp_path / f"theirs-{source}.h5"
|
||||
ours = tmp_path / f"ours-{source}.h5"
|
||||
shutil.copy(base, theirs)
|
||||
shutil.copy(base, ours)
|
||||
return str(theirs), str(ours), str(base)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Comparing files through h5py
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _norm_attr(v):
|
||||
"""Attribute values comparable across the two writers: clawhdf5 stores
|
||||
`str` as fixed-length UTF-8 (h5py reads bytes), h5py as variable-length
|
||||
(h5py reads str)."""
|
||||
if isinstance(v, bytes):
|
||||
return ("str", v.decode("utf-8"))
|
||||
if isinstance(v, str):
|
||||
return ("str", v)
|
||||
arr = np.asarray(v)
|
||||
if arr.dtype.kind in "SO":
|
||||
return ("strs", [x.decode() if isinstance(x, bytes) else x for x in arr.ravel().tolist()], arr.shape)
|
||||
return (arr.dtype.str, arr.shape, arr.tobytes())
|
||||
|
||||
|
||||
def snapshot(h5py, path):
|
||||
"""What h5py sees in the file: every dataset's shape, dtype, bytes and
|
||||
attributes (read without locking: clawhdf5 may hold the file open)."""
|
||||
out = {}
|
||||
with h5py.File(path, "r", locking=False) as f:
|
||||
def visit(name, obj):
|
||||
attrs = {k: _norm_attr(obj.attrs[k]) for k in obj.attrs}
|
||||
if isinstance(obj, h5py.Dataset):
|
||||
if obj.dtype.kind == "O":
|
||||
data = [x for x in obj[...].ravel().tolist()]
|
||||
else:
|
||||
data = obj[()].tobytes() if obj.shape is not None else None
|
||||
out[name] = (obj.shape, obj.dtype.str, obj.maxshape, data, attrs)
|
||||
else:
|
||||
out[name] = ("group", attrs)
|
||||
visit("/", f)
|
||||
f.visititems(visit)
|
||||
return out
|
||||
|
||||
|
||||
def assert_same_files(h5py, theirs, ours, what):
|
||||
a, b = snapshot(h5py, theirs), snapshot(h5py, ours)
|
||||
assert a.keys() == b.keys(), what
|
||||
for k in a:
|
||||
assert a[k] == b[k], f"{what}: {k} differs\n h5py: {a[k]}\n clawhdf5: {b[k]}"
|
||||
|
||||
|
||||
def assert_ours_reads_like_h5py(h5py, f, path, what):
|
||||
"""clawhdf5's own view of the file it is editing matches h5py's."""
|
||||
with h5py.File(path, "r", locking=False) as t:
|
||||
for name in ["i4", "chunk_ext", "chunk_bt2", "chunk_gzip", "f8_3d", "cmp", "bool", "enum", "scalar"]:
|
||||
if name not in t:
|
||||
continue
|
||||
o = f[name]
|
||||
assert o.shape == t[name].shape, f"{what}: {name} shape"
|
||||
assert o.maxshape == t[name].maxshape, f"{what}: {name} maxshape"
|
||||
np.testing.assert_array_equal(o[()], t[name][()], err_msg=f"{what}: {name}")
|
||||
for obj in ["/", "grp"]:
|
||||
assert sorted(f[obj].attrs.keys()) == sorted(t[obj].attrs.keys()), what
|
||||
for k in t[obj].attrs:
|
||||
assert _norm_attr(f[obj].attrs[k]) == _norm_attr(t[obj].attrs[k]), f"{what}: {obj}.attrs[{k}]"
|
||||
|
||||
|
||||
def h5dump_reads(path, base=None):
|
||||
"""h5dump (libhdf5 1.14) reads every object and value of `path` — when it
|
||||
reads the unedited `base` (it cannot read HDF5 2.0's newest format)."""
|
||||
exe = shutil.which("h5dump")
|
||||
if exe is None:
|
||||
if os.environ.get("CLAWHDF5_REQUIRE_INTEROP") == "1":
|
||||
pytest.fail("h5dump is required (CLAWHDF5_REQUIRE_INTEROP=1)")
|
||||
return
|
||||
h5rs = os.environ.get("CLAWHDF5_H5RS")
|
||||
if h5rs:
|
||||
# clawhdf5's structural and checksum validator (scripts/ci-test.sh
|
||||
# points this at the h5rs it built).
|
||||
r = subprocess.run([h5rs, "check", path], capture_output=True, text=True)
|
||||
assert r.returncode == 0, (r.stdout + r.stderr)[-2000:]
|
||||
if base is not None and subprocess.run([exe, "-H", base], capture_output=True).returncode != 0:
|
||||
return
|
||||
r = subprocess.run([exe, path], capture_output=True, text=True)
|
||||
assert r.returncode == 0, r.stderr[-2000:]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Applying one edit both ways
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _native_conversion(h5py, value, ds_dtype):
|
||||
"""`value` as libhdf5 converts it to `ds_dtype` in native byte order.
|
||||
|
||||
libhdf5 2.0 (h5py 3.16) converts numbers differently when either side is
|
||||
not in native byte order (its "soft" conversions): a float in (-1, 0)
|
||||
becomes the integer type's minimum instead of 0, and an unsigned integer
|
||||
too large for the signed type of the same size wraps instead of
|
||||
saturating. clawhdf5 applies the native-order results to every byte
|
||||
order, so the reference is h5py converting into a native dataset of the
|
||||
same kind; the result then reaches the real dataset by a plain byte
|
||||
swap."""
|
||||
if not isinstance(value, np.ndarray):
|
||||
return value
|
||||
if value.dtype.kind not in "biuf" or ds_dtype.kind not in "biuf":
|
||||
return value
|
||||
if value.dtype.isnative and ds_dtype.isnative:
|
||||
return value
|
||||
with h5py.File(io.BytesIO(), "w") as tmp:
|
||||
d = tmp.create_dataset("t", shape=value.shape, dtype=ds_dtype.newbyteorder("="))
|
||||
d[...] = value.astype(value.dtype.newbyteorder("="))
|
||||
return np.asarray(d[()])
|
||||
|
||||
|
||||
def _apply(f, op, h5py=None):
|
||||
"""Apply `op` to `f`; with `h5py`, `f` is an h5py file and a numpy array
|
||||
value is first converted as libhdf5 converts in native byte order (see
|
||||
`_native_conversion`)."""
|
||||
kind = op[0]
|
||||
if kind == "set":
|
||||
_, name, key, value = op
|
||||
if h5py is not None:
|
||||
value = _native_conversion(h5py, value, f[name].dtype)
|
||||
f[name][key] = value
|
||||
elif kind == "resize":
|
||||
_, name, size, axis = op
|
||||
if axis is None:
|
||||
f[name].resize(size)
|
||||
else:
|
||||
f[name].resize(size, axis=axis)
|
||||
elif kind == "attr":
|
||||
_, obj, name, value = op
|
||||
f[obj].attrs[name] = value
|
||||
else:
|
||||
raise AssertionError(op)
|
||||
|
||||
|
||||
def edit_both(h5py, theirs, ours_path, ours, op):
|
||||
"""Apply `op` with h5py and with clawhdf5 (`ours`, open 'r+'); the two
|
||||
files must then read the same through h5py. If h5py refuses, clawhdf5
|
||||
must refuse and its file must be unchanged. Returns h5py's error."""
|
||||
before = snapshot(h5py, ours_path)
|
||||
try:
|
||||
with h5py.File(theirs, "r+") as t:
|
||||
_apply(t, op, h5py)
|
||||
except Exception as e: # noqa: BLE001 - h5py refuses: so must we
|
||||
try:
|
||||
_apply(ours, op)
|
||||
except Exception: # noqa: BLE001
|
||||
pass
|
||||
else:
|
||||
pytest.fail(f"{op!r:.300}: h5py refused ({type(e).__name__}: {e}), clawhdf5 did not")
|
||||
assert snapshot(h5py, ours_path) == before, f"{op!r}: clawhdf5 changed the file while failing"
|
||||
return e
|
||||
_apply(ours, op)
|
||||
assert_same_files(h5py, theirs, ours_path, repr(op)[:200])
|
||||
return None
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Tests
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.parametrize("source", SOURCES)
|
||||
def test_edit_sequence_matches_h5py(h5py, tmp_path, source):
|
||||
theirs, ours_path, base = _make(h5py, tmp_path, source)
|
||||
ops = [
|
||||
("set", "i4", 0, 99),
|
||||
("set", "i4", (slice(1, 5, 2), slice(None, None, 3)), np.array([[1.5, -2.5, 1e12, -1e12]])),
|
||||
("set", "i4", (slice(None), 2), np.arange(6, dtype="<i8") * 1000),
|
||||
("set", "i4", ([0, 2, 5], slice(4, 6)), np.array([[1, 2], [3, 4], [5, 6]])),
|
||||
("set", "i4", (Ellipsis, -1), np.int8(-7)),
|
||||
("set", "i4", (3, 3), 12345.9),
|
||||
("set", "u1", slice(2, 8), np.array([-5, 0, 300, 255, 256, 1], dtype="<i4")),
|
||||
("set", "u1", slice(0, 3), [1, 2, 3]),
|
||||
("set", "chunk_gzip", (slice(0, 20, 7), slice(3, 17)), 42.25),
|
||||
("set", "chunk_gzip", (slice(5, 11), slice(5, 11)), np.ones((6, 6), dtype="<f8") * np.pi),
|
||||
("set", "grp/leaf", slice(None), np.array([1, 2, 3, 4, 5], dtype="<u2")),
|
||||
("attr", "/", "version", np.int32(2)),
|
||||
("attr", "/", "count", 7),
|
||||
("attr", "grp", "scale", np.array([0.5, 0.25], dtype="<f4")),
|
||||
("attr", "grp", "matrix", np.arange(6, dtype=">i8").reshape(2, 3)),
|
||||
("attr", "i4", "flag", np.bool_(True)),
|
||||
("attr", "i4", "z", np.complex64(1 - 2j)),
|
||||
("attr", "i4", "raw", np.bytes_(b"abc")),
|
||||
("set", "i4", slice(0, 2), np.zeros((3, 10))), # shape mismatch: refused by both
|
||||
]
|
||||
if source != "clawhdf5":
|
||||
ops += [
|
||||
("set", "be_i2", (slice(None), slice(1, 3)), np.array([70000, -70000], dtype="<i8")),
|
||||
("set", "f2", slice(None, None, 4), np.array([1e6, -3.25, 0.1])),
|
||||
("set", "f4_2d", (2, slice(None)), np.linspace(-1, 1, 9)),
|
||||
("set", "f8_3d", (slice(1, 3), 2, slice(None, None, 2)), np.arange(3, dtype="<i2")),
|
||||
("set", "u8", slice(None), np.array([-1, 0, 2**63, 1e30, -1e30, 5.5, 2, 3, 4, 5])),
|
||||
("set", "c8", slice(1, 3), np.array([1 + 1j, 2 - 2j], dtype="<c16")),
|
||||
("set", "c8", 0, np.float64(1.0)), # h5py: no conversion path
|
||||
("set", "bool", slice(None), np.array([0, 3, 0, -1], dtype="<i4")),
|
||||
("set", "bool", 1, np.array(True)),
|
||||
("set", "enum", slice(0, 2), np.array([7, 1], dtype="<i4")),
|
||||
("set", "s5", 0, np.bytes_(b"xyzuvw")),
|
||||
("set", "s5", slice(1, 3), [b"q", b"rs"]),
|
||||
("set", "s5", 2, np.array("uni")), # h5py: no conversion from 'U'
|
||||
("set", "cmp", 2, np.array((9, 9.5, b"zz"), dtype=[("id", "<i4"), ("x", "<f8"), ("tag", "S3")])),
|
||||
("set", "cmp", slice(3, 5), [(1, 0.5, b"a"), (2, 1.5, b"b")]),
|
||||
("set", "scalar", (), 7.25),
|
||||
("set", "scalar", Ellipsis, np.float32(-1.5)),
|
||||
("set", "compact", slice(1, 6, 2), np.array([10, 20, 30])),
|
||||
("set", "chunk_fixed", (slice(2, 9), slice(1, 10, 4)), np.arange(21).reshape(7, 3)),
|
||||
("set", "chunk_single", (1, slice(None)), np.array([9, 8, 7, 6])),
|
||||
("resize", "chunk_ext", (20, 7), None),
|
||||
("set", "chunk_ext", slice(12, 20), np.full((8, 7), 2.5)),
|
||||
("resize", "chunk_ext", 9, 0),
|
||||
("resize", "chunk_ext", 16, 0),
|
||||
("resize", "chunk_bt2", (9, 11), None),
|
||||
("set", "chunk_bt2", (slice(4, 9), slice(5, 11)), np.arange(30).reshape(5, 6)),
|
||||
("resize", "chunk_bt2", (3, 3), None),
|
||||
("resize", "chunk_bt2", (7, 8), None),
|
||||
("resize", "chunk_gzip", (40, 25), None),
|
||||
("resize", "chunk_gzip", (41, 25), None), # beyond maxshape: refused
|
||||
("resize", "chunk_fill", 15, None), # h5py: a size without axis must be a tuple
|
||||
("resize", "chunk_fill", (15,), None),
|
||||
("set", "chunk_fill", slice(10, 12), [1, 2]),
|
||||
("resize", "chunk_fill", (4,), None),
|
||||
("resize", "chunk_fill", (20,), None),
|
||||
("resize", "i4", (7, 10), None), # not chunked: refused
|
||||
]
|
||||
with clawhdf5.File(ours_path, "r+") as ours:
|
||||
assert ours.mode == "r+"
|
||||
for i, op in enumerate(ops):
|
||||
edit_both(h5py, theirs, ours_path, ours, op)
|
||||
if i % 5 == 0:
|
||||
assert_ours_reads_like_h5py(h5py, ours, ours_path, repr(op))
|
||||
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
|
||||
h5dump_reads(ours_path, base)
|
||||
# Reopened, clawhdf5 reads what h5py reads.
|
||||
with clawhdf5.File(ours_path, "r") as f:
|
||||
assert f.mode == "r"
|
||||
assert_ours_reads_like_h5py(h5py, f, ours_path, "reopened")
|
||||
|
||||
|
||||
def _random_key(rng, shape):
|
||||
key = []
|
||||
for n in shape:
|
||||
r = rng.random()
|
||||
if n == 0 or r < 0.3:
|
||||
start = int(rng.integers(0, n + 1)) if n else 0
|
||||
stop = int(rng.integers(start, n + 1)) if n else 0
|
||||
step = int(rng.integers(1, 4))
|
||||
key.append(slice(start, stop, step))
|
||||
elif r < 0.55:
|
||||
key.append(int(rng.integers(-n, n)))
|
||||
elif r < 0.7:
|
||||
k = int(rng.integers(1, min(n, 4) + 1))
|
||||
key.append(sorted(rng.choice(n, size=k, replace=False).tolist()))
|
||||
else:
|
||||
key.append(slice(None))
|
||||
# Only one index list per key.
|
||||
lists = [i for i, k in enumerate(key) if isinstance(k, list)]
|
||||
for i in lists[1:]:
|
||||
key[i] = slice(None)
|
||||
return tuple(key)
|
||||
|
||||
|
||||
def _selection_shape(key, shape):
|
||||
out = []
|
||||
fancy = False
|
||||
for k, n in zip(key, shape):
|
||||
if isinstance(k, slice):
|
||||
out.append(len(range(*k.indices(n))))
|
||||
elif isinstance(k, list):
|
||||
out.append(len(k))
|
||||
fancy = True
|
||||
return tuple(out), fancy
|
||||
|
||||
|
||||
def _random_value(rng, sel_shape, fancy, dtype):
|
||||
r = rng.random()
|
||||
if r < 0.2 or not sel_shape:
|
||||
v = rng.standard_normal() * 1000
|
||||
return np.float64(v) if rng.random() < 0.5 else int(v)
|
||||
shape = list(sel_shape)
|
||||
if not fancy and r < 0.35 and shape:
|
||||
shape[0] = 1 # broadcast along the first axis
|
||||
kind = rng.choice(["same", "f8", "i8", "u1", "f4"])
|
||||
if kind == "same" and np.dtype(dtype).kind in "iuf":
|
||||
dt = np.dtype(dtype)
|
||||
else:
|
||||
dt = np.dtype(str(kind) if kind != "same" else "f8")
|
||||
base = rng.standard_normal(size=shape) * (10 ** rng.integers(0, 6))
|
||||
with np.errstate(all="ignore"):
|
||||
return base.astype(dt)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("source", SOURCES)
|
||||
@pytest.mark.parametrize("seed", [0, 1, 2, 3])
|
||||
def test_random_edits_match_h5py(h5py, tmp_path, source, seed):
|
||||
"""Random writes (slices, steps, integers, index lists, broadcasts,
|
||||
other dtypes and out-of-range values), resizes and attributes, each
|
||||
compared with h5py applying the same edit."""
|
||||
rng = np.random.default_rng(1000 * seed + SOURCES.index(source))
|
||||
theirs, ours_path, base = _make(h5py, tmp_path, source)
|
||||
with h5py.File(theirs, "r") as t:
|
||||
names = [n for n in ["i4", "u1", "f8", "f4_2d", "f8_3d", "be_i2", "chunk_fixed", "chunk_ext",
|
||||
"chunk_bt2", "chunk_gzip", "chunk_fill", "compact", "grp/leaf"] if n in t]
|
||||
resizable = [n for n in ["chunk_ext", "chunk_bt2", "chunk_gzip", "chunk_fill"] if n in names]
|
||||
refused = 0
|
||||
with clawhdf5.File(ours_path, "r+") as ours:
|
||||
for step in range(40):
|
||||
r = rng.random()
|
||||
if r < 0.15 and resizable:
|
||||
name = str(rng.choice(resizable))
|
||||
with h5py.File(theirs, "r") as t:
|
||||
maxshape = t[name].maxshape
|
||||
shape = t[name].shape
|
||||
new = tuple(int(rng.integers(0, (m if m is not None else s + 10) + 1)) for m, s in zip(maxshape, shape))
|
||||
op = ("resize", name, new, None)
|
||||
elif r < 0.25:
|
||||
obj = str(rng.choice(["/", "grp", names[0]]))
|
||||
choices = [np.int16(rng.integers(-100, 100)), rng.standard_normal(3),
|
||||
np.arange(int(rng.integers(1, 5)), dtype=">u4"), np.float32(0.5)]
|
||||
value = choices[int(rng.integers(0, len(choices)))]
|
||||
op = ("attr", obj, f"a{int(rng.integers(0, 4))}", value)
|
||||
else:
|
||||
name = str(rng.choice(names))
|
||||
with h5py.File(theirs, "r") as t:
|
||||
shape, dtype = t[name].shape, t[name].dtype
|
||||
key = _random_key(rng, shape)
|
||||
sel_shape, fancy = _selection_shape(key, shape)
|
||||
op = ("set", name, key, _random_value(rng, sel_shape, fancy, dtype))
|
||||
before = None
|
||||
if op[0] == "resize":
|
||||
with h5py.File(ours_path, "r", locking=False) as o:
|
||||
before, fill = o[op[1]][()], o[op[1]].fillvalue
|
||||
if edit_both(h5py, theirs, ours_path, ours, op) is not None:
|
||||
refused += 1
|
||||
elif before is not None:
|
||||
# Independently of h5py (which a wrong index layout fools
|
||||
# the same way): the kept elements keep their values.
|
||||
want = resized_model(before, op[2], fill)
|
||||
np.testing.assert_array_equal(ours[op[1]][()], want, err_msg=repr(op))
|
||||
with h5py.File(ours_path, "r", locking=False) as o:
|
||||
np.testing.assert_array_equal(o[op[1]][()], want, err_msg=repr(op))
|
||||
assert_ours_reads_like_h5py(h5py, ours, ours_path, "end")
|
||||
assert refused < 30
|
||||
h5dump_reads(ours_path, base)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("seed", range(10, 40))
|
||||
def test_random_edits_on_clawhdf5_files(h5py, tmp_path, seed):
|
||||
"""More random sequences on a clawhdf5-written file, whose resizes to
|
||||
zero extents once left files `h5rs check` could not read."""
|
||||
test_random_edits_match_h5py(h5py, tmp_path, "clawhdf5", seed)
|
||||
|
||||
|
||||
def resized_model(before, shape, fill):
|
||||
"""`before` resized to `shape` as HDF5 resizes: elements inside both
|
||||
extents keep their values, the others read as the fill value."""
|
||||
out = np.full(shape, fill, dtype=before.dtype)
|
||||
common = tuple(slice(0, min(a, b)) for a, b in zip(before.shape, shape))
|
||||
out[common] = before[common]
|
||||
return out
|
||||
|
||||
|
||||
RESIZES = [(15, 15), (3, 2), (20, 20), (1, 1), (1, 0), (0, 0), (7, 20), (20, 13), (20, 20)]
|
||||
|
||||
|
||||
def _check_resizes(h5py, path, name, orig, maxshape=(20, 20), base=None):
|
||||
"""Resize `name` through RESIZES in 'r+', each checked against a numpy
|
||||
model with clawhdf5 and h5py and the file with `h5rs check`."""
|
||||
model = orig
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
ds = f[name]
|
||||
for shape in RESIZES:
|
||||
ds.resize(shape)
|
||||
model = resized_model(model, shape, 0)
|
||||
np.testing.assert_array_equal(ds[()], model, err_msg=f"{name} {shape}")
|
||||
with h5py.File(path, "r", locking=False) as t:
|
||||
np.testing.assert_array_equal(t[name][()], model, err_msg=f"h5py: {name} {shape}")
|
||||
assert t[name].maxshape == maxshape
|
||||
h5rs = os.environ.get("CLAWHDF5_H5RS")
|
||||
if h5rs: # h5dump would wait for the editor's lock
|
||||
r = subprocess.run([h5rs, "check", "--data", path], capture_output=True, text=True)
|
||||
assert r.returncode == 0, f"{name} {shape}: " + (r.stdout + r.stderr)[-2000:]
|
||||
with pytest.raises(ValueError):
|
||||
ds.resize((maxshape[0] + 1, 20))
|
||||
# Written values survive a shrink.
|
||||
ds[...] = orig
|
||||
ds.resize((15, 15))
|
||||
with h5py.File(path, "r") as t:
|
||||
np.testing.assert_array_equal(t[name][()], orig[:15, :15])
|
||||
h5dump_reads(path, base)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("source", SOURCES)
|
||||
def test_resizes_keep_values(h5py, tmp_path, source):
|
||||
"""Shrinking, zero extents and growing back keep the values a numpy
|
||||
model keeps, on every source (on clawhdf5's own files a shrink once
|
||||
moved every chunk: h5py read the same wrong values)."""
|
||||
_, ours, base = _make(h5py, tmp_path, source)
|
||||
with clawhdf5.File(ours, "r+") as f:
|
||||
f["chunk_gzip"].resize((20, 20))
|
||||
maxshape = (20, 20) if source == "clawhdf5" else (40, 40)
|
||||
_check_resizes(h5py, ours, "chunk_gzip", np.arange(400, dtype="<f4").reshape(20, 20), maxshape, base)
|
||||
|
||||
|
||||
FIXTURES = os.path.join(os.path.dirname(__file__), "..", "..", "clawhdf5", "tests", "fixtures")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", ["d", "z"])
|
||||
def test_resizes_of_a_file_with_no_recorded_maxshape(h5py, tmp_path, name):
|
||||
"""A file clawhdf5 2.7.0 wrote records no maximum dimensions for chunked
|
||||
datasets; resizing it must keep the Fixed Array index's layout."""
|
||||
path = str(tmp_path / "old.h5")
|
||||
shutil.copy(os.path.join(FIXTURES, "chunked_no_maxshape_v2_7_0.h5"), path)
|
||||
with h5py.File(path, "r") as t:
|
||||
orig = t[name][()]
|
||||
_check_resizes(h5py, path, name, orig)
|
||||
|
||||
|
||||
NUMERIC = ["<i1", "<u1", "<i2", ">u2", "<i4", "<u4", ">i8", "<u8", "<f2", "<f4", ">f8"]
|
||||
|
||||
|
||||
def _quiet(f):
|
||||
with np.errstate(all="ignore"):
|
||||
return f()
|
||||
|
||||
|
||||
def _libhdf5_undefined(vals, target):
|
||||
"""The values whose conversion to `target` libhdf5 2.0 (h5py 3.16) gets
|
||||
wrong even in native byte order, where its C casts are undefined
|
||||
behaviour; clawhdf5 saturates them as libhdf5's range handling
|
||||
intends (docs/known-issues.md):
|
||||
|
||||
- half floats into unsigned integers: negatives wrap (-1 -> 65535) and
|
||||
+inf becomes 0; into signed integers, +-inf becomes the minimum;
|
||||
- a float equal to the integer maximum rounded up in the float's
|
||||
precision (float32(2**31 - 1) == 2**31 -> int32, float64(2**64 - 1)
|
||||
-> uint64) becomes the minimum (or 0);
|
||||
- a double between 65504 and 65520 into a half float becomes infinity
|
||||
(IEEE rounds it down to 65504, as numpy does).
|
||||
"""
|
||||
bad = np.zeros(vals.shape, dtype=bool)
|
||||
if vals.dtype.kind != "f":
|
||||
return bad
|
||||
if target.kind in "iu":
|
||||
if vals.dtype.itemsize == 2:
|
||||
bad |= np.isinf(vals)
|
||||
if target.kind == "u":
|
||||
bad |= vals <= -1
|
||||
top = _quiet(lambda: np.array(np.iinfo(target).max).astype(vals.dtype))
|
||||
if float(top) > np.iinfo(target).max:
|
||||
bad |= vals == top
|
||||
if target.kind == "f" and target.itemsize == 2 and vals.dtype.itemsize > 2:
|
||||
bad |= (np.abs(vals) > 65504) & (np.abs(vals) < 65520)
|
||||
return bad
|
||||
|
||||
|
||||
def test_numeric_conversions_match_h5py(h5py, tmp_path):
|
||||
"""Every numeric source dtype into every numeric dataset dtype, with
|
||||
values at and beyond the targets' limits, as libhdf5 converts them."""
|
||||
edge = np.array([0, 1, -1, -0.3, 2.5, -2.5, 3.7, -3.7, 127.9, -128.9, 200.5, 255.5, 256, -129,
|
||||
32767.5, 40000, 65504, 70000, -70000, 2**31 - 1, 2**31, -2**31 - 1,
|
||||
4e9, 1e15, -1e15, 1e19, 1e300, -1e300, np.inf, -np.inf])
|
||||
sources = {
|
||||
"f8": edge,
|
||||
"f4": _quiet(lambda: edge.astype("<f4")),
|
||||
"f2": np.array([0, 1, -1, 2.5, -3.5, 65504, -65504, np.inf, -np.inf, 100.5], dtype="<f2"),
|
||||
"i8": np.array([0, 1, -1, 127, 128, -129, 255, 256, 32768, -32769, 65536, 2**31, -2**31 - 1,
|
||||
2**32, 2**62, -2**63, 2**63 - 1], dtype="<i8"),
|
||||
"u8": np.array([0, 1, 127, 128, 255, 256, 65535, 65536, 2**31, 2**32, 2**63, 2**64 - 1], dtype="<u8"),
|
||||
"i1": np.array([-128, -1, 0, 1, 127], dtype="i1"),
|
||||
"u2": np.array([0, 255, 256, 65535], dtype=">u2"),
|
||||
"b": np.array([True, False, True]),
|
||||
}
|
||||
for target in NUMERIC:
|
||||
for sname, src in sources.items():
|
||||
vals = src[~_libhdf5_undefined(src, np.dtype(target))]
|
||||
path_t = str(tmp_path / f"t_{target[1:]}_{sname}.h5")
|
||||
path_o = str(tmp_path / f"o_{target[1:]}_{sname}.h5")
|
||||
with h5py.File(path_t, "w") as f:
|
||||
f.create_dataset("d", shape=vals.shape, dtype=target)
|
||||
shutil.copy(path_t, path_o)
|
||||
what = f"{vals.dtype} -> {target}"
|
||||
try:
|
||||
with h5py.File(path_t, "r+") as f:
|
||||
f["d"][...] = _native_conversion(h5py, vals, f["d"].dtype)
|
||||
except Exception: # noqa: BLE001
|
||||
with clawhdf5.File(path_o, "r+") as f, pytest.raises(Exception):
|
||||
f["d"][...] = vals
|
||||
continue
|
||||
with clawhdf5.File(path_o, "r+") as f:
|
||||
f["d"][...] = vals
|
||||
with h5py.File(path_t, "r") as a, h5py.File(path_o, "r") as b:
|
||||
assert a["d"][...].tobytes() == b["d"][...].tobytes(), (
|
||||
f"{what}: h5py {a['d'][...].tolist()} clawhdf5 {b['d'][...].tolist()}"
|
||||
)
|
||||
|
||||
|
||||
def test_nan_into_an_integer_dataset_is_refused(h5py, tmp_path):
|
||||
"""libhdf5 stores NaN as an arbitrary integer (0, the minimum or 2**63,
|
||||
depending on the type); clawhdf5 refuses and writes nothing."""
|
||||
path = str(tmp_path / "nan.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
with pytest.raises(ValueError, match="NaN"):
|
||||
f["d"][...] = np.array([1.0, np.nan, 2.0, 3.0])
|
||||
# A Python list goes through numpy, which refuses NaN too.
|
||||
with pytest.raises(ValueError):
|
||||
f["d"][0:2] = [np.nan, 1.0]
|
||||
with h5py.File(path, "r") as f:
|
||||
np.testing.assert_array_equal(f["d"][...], np.arange(4))
|
||||
|
||||
|
||||
def test_unsupported_edits_are_clear_errors(h5py, tmp_path):
|
||||
path = str(tmp_path / "u.h5")
|
||||
_h5py_file(h5py, path, "earliest")
|
||||
before = snapshot(h5py, path)
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
with pytest.raises(NotImplementedError, match="delet"):
|
||||
del f.attrs["version"]
|
||||
with pytest.raises(NotImplementedError, match="delet"):
|
||||
del f["i4"]
|
||||
with pytest.raises(NotImplementedError, match="delet"):
|
||||
del f["grp"]["leaf"]
|
||||
with pytest.raises(NotImplementedError):
|
||||
f.create_dataset("new", data=np.arange(3.0))
|
||||
with pytest.raises(NotImplementedError):
|
||||
f.create_group("newgrp")
|
||||
with pytest.raises(NotImplementedError):
|
||||
f["grp"].create_dataset("new", data=np.arange(3.0))
|
||||
with pytest.raises(NotImplementedError, match="variable-length"):
|
||||
f["vlen"][0] = "x"
|
||||
with pytest.raises(NotImplementedError, match="field"):
|
||||
f["cmp"]["id"] = np.arange(5)
|
||||
with pytest.raises(NotImplementedError):
|
||||
f.attrs["empty"] = clawhdf5.Empty("f8")
|
||||
with pytest.raises(TypeError, match="chunked"):
|
||||
f["i4"].resize((7, 10))
|
||||
with pytest.raises(ValueError):
|
||||
f["chunk_gzip"].resize((41, 20))
|
||||
with pytest.raises(ValueError, match="axis"):
|
||||
f["chunk_ext"].resize(3, axis=2)
|
||||
with pytest.raises(TypeError):
|
||||
f["i4"][0] = np.array(["a"] * 10)
|
||||
# h5py writes through boolean masks; clawhdf5 does not.
|
||||
with pytest.raises(NotImplementedError, match="mask"):
|
||||
f["u1"][np.arange(16) % 2 == 0] = 5
|
||||
with pytest.raises(NotImplementedError, match="mask"):
|
||||
f["i4"][f["i4"][()] > 30] = 0
|
||||
with pytest.raises(NotImplementedError, match="mask"):
|
||||
f["i4"][np.ones(6, dtype=bool), 2] = 0
|
||||
assert snapshot(h5py, path) == before
|
||||
h5dump_reads(path)
|
||||
|
||||
|
||||
def test_read_only_files_and_modes(h5py, tmp_path):
|
||||
path = str(tmp_path / "m.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("d", data=np.arange(4, dtype="<i4"), chunks=(2,), maxshape=(None,))
|
||||
with clawhdf5.File(path, "r") as f:
|
||||
with pytest.raises(OSError, match="r\\+"):
|
||||
f["d"][0] = 1
|
||||
with pytest.raises(OSError):
|
||||
f["d"].resize((8,))
|
||||
with pytest.raises(OSError):
|
||||
f.attrs["x"] = 1
|
||||
with pytest.raises(NotImplementedError, match="does not exist"):
|
||||
clawhdf5.File(str(tmp_path / "missing.h5"), "a")
|
||||
with pytest.raises(ValueError, match="mode"):
|
||||
clawhdf5.File(path, "rw")
|
||||
with clawhdf5.File(path, "a") as f:
|
||||
assert f.mode == "r+"
|
||||
f["d"][1] = 10
|
||||
with h5py.File(path, "r") as f:
|
||||
assert f["d"][1] == 10
|
||||
|
||||
|
||||
def test_the_file_is_locked_while_open_for_editing(h5py, tmp_path):
|
||||
path = str(tmp_path / "lock.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||
f = clawhdf5.File(path, "r+")
|
||||
with pytest.raises(OSError):
|
||||
clawhdf5.File(path, "r+")
|
||||
with pytest.raises(OSError):
|
||||
h5py.File(path, "r+")
|
||||
f.close()
|
||||
with h5py.File(path, "r+") as g:
|
||||
g["d"][0] = 5
|
||||
with clawhdf5.File(path, "r+") as g:
|
||||
g["d"][1] = 6
|
||||
with h5py.File(path, "r") as g:
|
||||
np.testing.assert_array_equal(g["d"][...], [5, 6, 2, 3])
|
||||
|
||||
|
||||
def test_objects_see_edits_made_through_others(h5py, tmp_path):
|
||||
"""A dataset or attrs object taken before an edit reports the file as
|
||||
it is after it: the new shape, the new attribute."""
|
||||
path = str(tmp_path / "live.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("d", data=np.arange(6.0), chunks=(4,), maxshape=(None,))
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
d1 = f["d"]
|
||||
d2 = f["d"]
|
||||
attrs = d1.attrs
|
||||
assert len(attrs) == 0 and "u" not in attrs
|
||||
d2.resize((10,))
|
||||
assert d1.shape == (10,) and d1.size == 10 and len(d1) == 10
|
||||
np.testing.assert_array_equal(d1[6:], np.zeros(4))
|
||||
d2.attrs["u"] = "m/s"
|
||||
assert "u" in attrs and attrs["u"] == b"m/s" and len(attrs) == 1
|
||||
f.attrs.create("shaped", np.arange(6), shape=(2, 3), dtype="<i2")
|
||||
f.attrs.modify("shaped2", [1.5, 2.5])
|
||||
with h5py.File(path, "r") as f:
|
||||
assert f["d"].shape == (10,)
|
||||
assert f["d"].attrs["u"] == b"m/s"
|
||||
assert f.attrs["shaped"].dtype == np.dtype("<i2") and f.attrs["shaped"].shape == (2, 3)
|
||||
np.testing.assert_array_equal(f.attrs["shaped2"], [1.5, 2.5])
|
||||
|
||||
|
||||
def test_attribute_types_as_h5py_reads_them(h5py, tmp_path):
|
||||
path = str(tmp_path / "attrs.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_group("g")
|
||||
values = {
|
||||
"i8": 5,
|
||||
"f8": 2.5,
|
||||
"i1": np.int8(-3),
|
||||
"u8": np.uint64(2**64 - 1),
|
||||
">f4": np.array([1.5, 2.5], dtype=">f4"),
|
||||
"f2": np.float16(0.5),
|
||||
"b": True,
|
||||
"barr": np.array([True, False]),
|
||||
"c16": np.complex128(1 + 2j),
|
||||
"bytes": b"raw",
|
||||
"sarr": np.array([b"a", b"bcd"]),
|
||||
"str": "héllo",
|
||||
"strs": ["x", "yz"],
|
||||
"2d": np.arange(12, dtype="<u2").reshape(3, 4),
|
||||
"empty": np.zeros((0,), dtype="<i4"),
|
||||
}
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
for k, v in values.items():
|
||||
f["g"].attrs[k] = v
|
||||
# Replace one, with another type and size.
|
||||
f["g"].attrs["i8"] = np.arange(100.0)
|
||||
with h5py.File(path, "r") as f:
|
||||
a = f["g"].attrs
|
||||
np.testing.assert_array_equal(a["i8"], np.arange(100.0))
|
||||
assert a["f8"] == 2.5 and a["f8"].dtype == np.float64
|
||||
assert a["i1"] == -3 and a["i1"].dtype == np.int8
|
||||
assert a["u8"] == 2**64 - 1 and a["u8"].dtype == np.uint64
|
||||
assert a[">f4"].dtype == np.dtype(">f4")
|
||||
assert a["f2"].dtype == np.float16
|
||||
assert a["b"] is np.True_ or a["b"] == True # noqa: E712
|
||||
assert a["barr"].dtype == np.bool_
|
||||
assert a["c16"] == 1 + 2j
|
||||
assert a["bytes"] == b"raw"
|
||||
assert list(a["sarr"]) == [b"a", b"bcd"]
|
||||
# str is stored as fixed-length UTF-8: h5py reads bytes.
|
||||
assert a["str"].decode("utf-8") == "héllo"
|
||||
assert [x.decode() for x in a["strs"]] == ["x", "yz"]
|
||||
assert a["2d"].shape == (3, 4) and a["2d"].dtype == np.dtype("<u2")
|
||||
assert a["empty"].shape == (0,)
|
||||
with clawhdf5.File(path, "r") as f:
|
||||
assert f["g"].attrs["c16"] == 1 + 2j
|
||||
assert f["g"].attrs["str"].decode("utf-8") == "héllo"
|
||||
h5dump_reads(path)
|
||||
|
||||
|
||||
def test_many_attributes_move_to_dense_storage(h5py, tmp_path):
|
||||
"""Past the compact limit (8 attributes under libver v114) the object's
|
||||
attributes move to dense storage; h5py reads all of them."""
|
||||
path = str(tmp_path / "dense.h5")
|
||||
with h5py.File(path, "w", libver="v114") as f:
|
||||
f.create_dataset("d", data=np.arange(3))
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
for i in range(20):
|
||||
f["d"].attrs[f"a{i:02d}"] = np.full(i + 1, i, dtype="<i2")
|
||||
with h5py.File(path, "r") as f:
|
||||
assert sorted(f["d"].attrs.keys()) == [f"a{i:02d}" for i in range(20)]
|
||||
for i in range(20):
|
||||
np.testing.assert_array_equal(f["d"].attrs[f"a{i:02d}"], np.full(i + 1, i))
|
||||
h5dump_reads(path)
|
||||
|
||||
|
||||
def test_reads_never_see_a_half_written_edit(h5py, tmp_path):
|
||||
"""Readers on other threads while one thread rewrites a dataset: every
|
||||
read returns one whole version (all elements equal), never a mix."""
|
||||
path = str(tmp_path / "race.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("d", data=np.zeros((64, 64)), chunks=(16, 16), compression="gzip")
|
||||
errors = []
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
stop = threading.Event()
|
||||
|
||||
def read():
|
||||
ds = f["d"]
|
||||
while not stop.is_set():
|
||||
a = ds[...]
|
||||
if not (a == a.flat[0]).all():
|
||||
errors.append(a)
|
||||
return
|
||||
|
||||
readers = [threading.Thread(target=read) for _ in range(3)]
|
||||
for t in readers:
|
||||
t.start()
|
||||
try:
|
||||
for k in range(1, 25):
|
||||
f["d"][...] = float(k)
|
||||
finally:
|
||||
stop.set()
|
||||
for t in readers:
|
||||
t.join()
|
||||
np.testing.assert_array_equal(f["d"][...], np.full((64, 64), 24.0))
|
||||
assert not errors, "a read saw a partly written dataset"
|
||||
|
||||
|
||||
def test_close_releases_the_file(h5py, tmp_path):
|
||||
path = str(tmp_path / "close.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("d", data=np.arange(4, dtype="<i4"))
|
||||
f = clawhdf5.File(path, "r+")
|
||||
ds = f["d"]
|
||||
f.close()
|
||||
# The handle still reads the file as last written, but cannot edit it.
|
||||
np.testing.assert_array_equal(ds[...], np.arange(4))
|
||||
with pytest.raises(OSError, match="closed"):
|
||||
ds[0] = 1
|
||||
with h5py.File(path, "r+") as g:
|
||||
g["d"][0] = 9
|
||||
|
||||
|
||||
def _two_files(h5py, tmp_path):
|
||||
"""d1/f.h5 and d2/f.h5: same name, different layouts (the review's
|
||||
repro)."""
|
||||
(tmp_path / "d1").mkdir()
|
||||
(tmp_path / "d2").mkdir()
|
||||
with h5py.File(tmp_path / "d1" / "f.h5", "w") as f:
|
||||
f.create_dataset("x", data=np.arange(10, dtype="<i4"))
|
||||
f.create_dataset("big", data=np.full(5000, 1.5))
|
||||
with h5py.File(tmp_path / "d2" / "f.h5", "w") as f:
|
||||
f.create_dataset("pad", data=np.full(3000, 2.5))
|
||||
f.create_dataset("x", data=np.arange(10, dtype="<i4") + 500)
|
||||
return tmp_path / "d1" / "f.h5", tmp_path / "d2" / "f.h5"
|
||||
|
||||
|
||||
def _held_file_edited(h5py, held, other, other_bytes):
|
||||
"""The edits went to `held`, planned from its own metadata; `other`
|
||||
was not touched."""
|
||||
assert other.read_bytes() == other_bytes, "the other file changed"
|
||||
with h5py.File(held, "r") as f:
|
||||
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
|
||||
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
|
||||
np.testing.assert_array_equal(f.attrs["note"], np.arange(50.0))
|
||||
h5dump_reads(str(held))
|
||||
|
||||
|
||||
def test_relative_path_and_chdir(h5py, tmp_path, monkeypatch):
|
||||
"""A file opened by a relative path keeps being the file edited and read
|
||||
after os.chdir (an edit was planned from the file the path named in the
|
||||
new directory and written into the one open, corrupting it)."""
|
||||
held, other = _two_files(h5py, tmp_path)
|
||||
other_bytes = other.read_bytes()
|
||||
monkeypatch.chdir(held.parent)
|
||||
f = clawhdf5.File("f.h5", "r+")
|
||||
ds = f["x"]
|
||||
monkeypatch.chdir(other.parent)
|
||||
ds[:] = np.full(10, 7, "<i4")
|
||||
np.testing.assert_array_equal(ds[:5], np.full(5, 7)) # read from the held file
|
||||
f.attrs["note"] = np.arange(50.0)
|
||||
np.testing.assert_array_equal(f["big"][()], np.full(5000, 1.5))
|
||||
assert "pad" not in f
|
||||
f.close()
|
||||
_held_file_edited(h5py, held, other, other_bytes)
|
||||
# 'w' writes where the path named when the file was opened.
|
||||
monkeypatch.chdir(held.parent)
|
||||
w = clawhdf5.File("new.h5", "w")
|
||||
monkeypatch.chdir(other.parent)
|
||||
w.create_dataset("d", data=np.arange(3))
|
||||
w.close()
|
||||
assert (held.parent / "new.h5").exists() and not (other.parent / "new.h5").exists()
|
||||
|
||||
|
||||
def test_path_replaced_between_edits(h5py, tmp_path):
|
||||
"""The path renamed away and another file put in its place between
|
||||
edits: the edits go to the file held open, never mixed with the other."""
|
||||
held, other = _two_files(h5py, tmp_path)
|
||||
path = str(tmp_path / "f.h5")
|
||||
os.replace(held, path)
|
||||
moved = tmp_path / "moved.h5"
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
ds = f["x"]
|
||||
ds[0] = 7
|
||||
os.replace(path, moved)
|
||||
shutil.copy(other, path)
|
||||
other_bytes = open(path, "rb").read()
|
||||
ds[:] = np.full(10, 7, "<i4")
|
||||
f.attrs["note"] = np.arange(50.0)
|
||||
np.testing.assert_array_equal(f["x"][()], np.full(10, 7))
|
||||
assert "pad" not in f
|
||||
from pathlib import Path
|
||||
_held_file_edited(h5py, moved, Path(path), other_bytes)
|
||||
|
||||
|
||||
def test_an_edit_releases_the_gil(h5py, tmp_path):
|
||||
"""Another Python thread keeps running while a large edit is written:
|
||||
had the edit held the GIL, the other thread would stall for the whole
|
||||
edit (Rust code never yields it)."""
|
||||
import time
|
||||
path = str(tmp_path / "big.h5")
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("d", shape=(1024, 1024), dtype="<f8", chunks=(64, 64), compression="gzip")
|
||||
value = np.random.default_rng(0).standard_normal((1024, 1024))
|
||||
stamps = []
|
||||
stop = threading.Event()
|
||||
|
||||
def spin():
|
||||
while not stop.is_set():
|
||||
stamps.append(time.perf_counter())
|
||||
|
||||
with clawhdf5.File(path, "r+") as f:
|
||||
ds = f["d"]
|
||||
t = threading.Thread(target=spin)
|
||||
t.start()
|
||||
try:
|
||||
time.sleep(0.05)
|
||||
t0 = time.perf_counter()
|
||||
ds[...] = value
|
||||
t1 = time.perf_counter()
|
||||
finally:
|
||||
stop.set()
|
||||
t.join()
|
||||
np.testing.assert_array_equal(ds[...], value)
|
||||
during = [x for x in stamps if t0 <= x <= t1]
|
||||
gaps = np.diff([t0] + during + [t1])
|
||||
assert t1 - t0 > 0.1, "the edit is too quick to tell"
|
||||
assert gaps.max() < 0.5 * (t1 - t0), (
|
||||
f"the other thread stalled for {gaps.max():.3f} s of a {t1 - t0:.3f} s edit")
|
||||
@@ -178,15 +178,39 @@ def _write_fixture(h5py, path):
|
||||
g.attrs["depth"] = np.int8(3)
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def pair(h5py, tmp_path_factory):
|
||||
path = str(tmp_path_factory.mktemp("h5") / "fixture.h5")
|
||||
@pytest.fixture(scope="module", params=["local", "r+", "http", "http-1k-blocks"])
|
||||
def pair(request, h5py, tmp_path_factory):
|
||||
"""The fixture file through h5py and through clawhdf5: opened locally
|
||||
(read-only, and for editing: a copy, since editing locks the file), and
|
||||
over HTTP range requests (a local server in this process) with the
|
||||
default 1 MiB blocks and with 1 KiB blocks, so every structure is read
|
||||
through many small ranges."""
|
||||
import shutil
|
||||
|
||||
from conftest import RangeServer
|
||||
|
||||
root = tmp_path_factory.mktemp("h5")
|
||||
path = str(root / "fixture.h5")
|
||||
_write_fixture(h5py, path)
|
||||
theirs = h5py.File(path, "r")
|
||||
ours = clawhdf5.File(path, "r")
|
||||
server = None
|
||||
if request.param == "local":
|
||||
ours = clawhdf5.File(path, "r")
|
||||
elif request.param == "r+":
|
||||
copy = str(root / "editable.h5")
|
||||
shutil.copy(path, copy)
|
||||
ours = clawhdf5.File(copy, "r+")
|
||||
else:
|
||||
server = RangeServer(root)
|
||||
if request.param == "http":
|
||||
ours = clawhdf5.File(server.url("fixture.h5"))
|
||||
else:
|
||||
ours = clawhdf5.File.open_url(server.url("fixture.h5"), block_size=1024)
|
||||
yield ours, theirs, path
|
||||
theirs.close()
|
||||
ours.close()
|
||||
if server is not None:
|
||||
server.close()
|
||||
|
||||
|
||||
def _all_datasets(h5py, f):
|
||||
@@ -416,7 +440,7 @@ def test_unsupported_types_are_errors_not_data(pair):
|
||||
|
||||
def test_boolean_masks_are_refused(pair):
|
||||
ours, _, _ = pair
|
||||
with pytest.raises(TypeError):
|
||||
with pytest.raises(NotImplementedError, match="mask"):
|
||||
ours["num/le_i4_1d"][np.ones(37, dtype=bool)]
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,244 @@
|
||||
"""Remote files: `clawhdf5.File(url)` / `File.open_url(url, ...)` read over
|
||||
HTTP range requests (clawhdf5-remote's block cache), against a server in this
|
||||
process (conftest.RangeServer). Values are compared with h5py reading the
|
||||
same file locally; the rest checks what the server saw (only the blocks a
|
||||
read needs are fetched), the failure modes (no range support, a missing
|
||||
file, a file that changes, a server that goes away: errors, never wrong
|
||||
data), and that the GIL is released while a read waits on the network."""
|
||||
|
||||
import os
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
|
||||
import clawhdf5
|
||||
from conftest import RangeServer
|
||||
|
||||
|
||||
def _write(h5py, path):
|
||||
rng = np.random.default_rng(7)
|
||||
with h5py.File(path, "w") as f:
|
||||
f.create_dataset("contig", data=rng.standard_normal((400, 300)))
|
||||
f.create_dataset(
|
||||
"chunked",
|
||||
data=rng.integers(0, 1000, size=(512, 512), dtype="<i4"),
|
||||
chunks=(64, 64),
|
||||
compression="gzip",
|
||||
)
|
||||
f.create_dataset("strings", data=["alpha", "beta", "gamma"], dtype=h5py.string_dtype())
|
||||
g = f.create_group("grp")
|
||||
g.create_dataset("small", data=np.arange(10, dtype="<u2"))
|
||||
g.attrs["units"] = "m/s"
|
||||
f.attrs["title"] = "remote test"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def remote_file(h5py, tmp_path, range_server):
|
||||
path = tmp_path / "remote.h5"
|
||||
_write(h5py, str(path))
|
||||
return path, range_server.url("remote.h5"), range_server
|
||||
|
||||
|
||||
def test_remote_reads_match_h5py(h5py, remote_file):
|
||||
path, url, server = remote_file
|
||||
with h5py.File(path, "r") as theirs, clawhdf5.File(url) as ours:
|
||||
assert ours.filename == url
|
||||
assert list(ours.keys()) == list(theirs.keys())
|
||||
assert "grp/small" in ours and "nope" not in ours
|
||||
for name in ["contig", "chunked", "grp/small"]:
|
||||
np.testing.assert_array_equal(ours[name][...], theirs[name][...])
|
||||
np.testing.assert_array_equal(ours[name][3:7], theirs[name][3:7])
|
||||
np.testing.assert_array_equal(ours["chunked"][100:130, 200:300:3], theirs["chunked"][100:130, 200:300:3])
|
||||
np.testing.assert_array_equal(ours["chunked"][[1, 70, 300], 5], theirs["chunked"][[1, 70, 300], 5])
|
||||
assert list(ours["strings"][...]) == list(theirs["strings"][...])
|
||||
assert ours["grp"].attrs["units"] == theirs["grp"].attrs["units"]
|
||||
assert ours.attrs["title"] == theirs.attrs["title"]
|
||||
assert ours["chunked"].maxshape == theirs["chunked"].maxshape
|
||||
assert server.requests() >= 2
|
||||
|
||||
|
||||
def test_a_small_read_fetches_only_its_blocks(h5py, remote_file):
|
||||
"""With 4 KiB blocks, opening and reading one chunk of a 1 MB chunked
|
||||
dataset costs a handful of requests and a few blocks, not the file."""
|
||||
path, url, server = remote_file
|
||||
size = os.path.getsize(path)
|
||||
f = clawhdf5.File.open_url(url, block_size=4096)
|
||||
opened = server.requests()
|
||||
assert opened == 1, server.log
|
||||
ds = f["chunked"]
|
||||
got = ds[0:10, 0:10]
|
||||
with h5py.File(path, "r") as theirs:
|
||||
np.testing.assert_array_equal(got, theirs["chunked"][0:10, 0:10])
|
||||
stats = f.remote_stats
|
||||
assert stats["bytes_fetched"] < size / 4, (stats, size)
|
||||
assert server.requests() - opened <= 12, server.log
|
||||
# A second read of the same region is served by the cache.
|
||||
before = server.requests()
|
||||
ds[0:10, 0:10]
|
||||
assert server.requests() == before
|
||||
assert f.remote_stats["hits"] > stats["hits"]
|
||||
assert clawhdf5.File(str(path), "r").remote_stats is None
|
||||
|
||||
|
||||
def test_server_without_range_support(h5py, tmp_path):
|
||||
"""A server that ignores Range answers 200 with the whole file: that is
|
||||
an OSError by default, and a whole download when allowed."""
|
||||
path = tmp_path / "remote.h5"
|
||||
_write(h5py, str(path))
|
||||
server = RangeServer(tmp_path, ranges=False)
|
||||
try:
|
||||
url = server.url("remote.h5")
|
||||
with pytest.raises(OSError, match="range"):
|
||||
clawhdf5.File(url)
|
||||
with clawhdf5.File.open_url(url, allow_full_download=True) as ours, h5py.File(path, "r") as theirs:
|
||||
np.testing.assert_array_equal(ours["chunked"][...], theirs["chunked"][...])
|
||||
np.testing.assert_array_equal(ours["contig"][5], theirs["contig"][5])
|
||||
with pytest.raises(OSError):
|
||||
clawhdf5.File.open_url(url, allow_full_download=True, max_full_download=1000)
|
||||
finally:
|
||||
server.close()
|
||||
|
||||
|
||||
def test_errors_are_oserrors(remote_file):
|
||||
_, url, server = remote_file
|
||||
with pytest.raises(OSError, match="404"):
|
||||
clawhdf5.File(server.url("missing.h5"))
|
||||
with pytest.raises(ValueError, match="read-only"):
|
||||
clawhdf5.File(url, "r+")
|
||||
with pytest.raises(ValueError, match="read-only"):
|
||||
clawhdf5.File(url, "w")
|
||||
with pytest.raises(OSError, match="unsupported URL"):
|
||||
clawhdf5.File("nosuchscheme://x/y.h5")
|
||||
with pytest.raises(ValueError):
|
||||
clawhdf5.File.open_url(url, block_size=0)
|
||||
with pytest.raises(TypeError):
|
||||
clawhdf5.File.open_url(url, no_such_option=1)
|
||||
|
||||
|
||||
def test_object_store_urls_need_their_features():
|
||||
"""The default wheel has no S3/GCS/Azure clients (aws-lc-rs builds C):
|
||||
such a URL is an OSError naming the build feature."""
|
||||
for url, feature in [("s3://bucket/k.h5", "s3"), ("gs://b/k.h5", "gcs"), ("az://c/k.h5", "azure")]:
|
||||
try:
|
||||
clawhdf5.File(url)
|
||||
except OSError as e:
|
||||
if "feature" in str(e):
|
||||
assert f"`{feature}`" in str(e), str(e)
|
||||
else:
|
||||
pytest.fail(f"{url} opened")
|
||||
|
||||
|
||||
def test_https_needs_the_https_feature():
|
||||
"""The default wheel has no TLS stack (rustls needs ring, which builds C):
|
||||
an https URL is an OSError that names the build feature."""
|
||||
with pytest.raises(OSError) as e:
|
||||
clawhdf5.File("https://127.0.0.1:1/x.h5")
|
||||
msg = str(e.value)
|
||||
# Built with `--features https` the error is the refused connection.
|
||||
assert "https" in msg or "connect" in msg.lower() or "refused" in msg.lower(), msg
|
||||
|
||||
|
||||
def test_a_changed_file_is_an_error_not_mixed_data(h5py, remote_file):
|
||||
path, url, _ = remote_file
|
||||
f = clawhdf5.File.open_url(url, block_size=1024)
|
||||
first = f["grp/small"][...]
|
||||
# Rewrite the file with other values: new ETag, same name.
|
||||
time.sleep(0.01)
|
||||
with h5py.File(path, "w") as g:
|
||||
g.create_dataset("contig", data=np.zeros((400, 300)))
|
||||
with pytest.raises(OSError, match="changed"):
|
||||
f["contig"][...]
|
||||
np.testing.assert_array_equal(first, np.arange(10, dtype="<u2"))
|
||||
|
||||
|
||||
def test_a_server_that_goes_away_is_an_error(h5py, tmp_path):
|
||||
path = tmp_path / "remote.h5"
|
||||
_write(h5py, str(path))
|
||||
server = RangeServer(tmp_path)
|
||||
f = clawhdf5.File.open_url(server.url("remote.h5"), block_size=1024, retries=0, timeout=2)
|
||||
ds = f["contig"]
|
||||
server.close()
|
||||
with pytest.raises(OSError):
|
||||
ds[...]
|
||||
|
||||
|
||||
def test_threads_read_one_remote_file(h5py, remote_file):
|
||||
path, url, _ = remote_file
|
||||
f = clawhdf5.File.open_url(url, block_size=2048)
|
||||
with h5py.File(path, "r") as theirs:
|
||||
expected = theirs["chunked"][...]
|
||||
errors = []
|
||||
|
||||
def work(i):
|
||||
try:
|
||||
rows = slice((i * 37) % 400, (i * 37) % 400 + 64)
|
||||
np.testing.assert_array_equal(f["chunked"][rows], expected[rows])
|
||||
except Exception as e: # noqa: BLE001
|
||||
errors.append(e)
|
||||
|
||||
threads = [threading.Thread(target=work, args=(i,)) for i in range(16)]
|
||||
for t in threads:
|
||||
t.start()
|
||||
for t in threads:
|
||||
t.join()
|
||||
assert not errors, errors[:3]
|
||||
|
||||
|
||||
def test_remote_reads_release_the_gil(h5py, remote_file):
|
||||
"""A read waiting on a slow server lets other Python threads run: a
|
||||
thread counting in a loop keeps counting (and never stalls for long)
|
||||
while the main thread reads through requests that each take 0.2 s."""
|
||||
_, url, server = remote_file
|
||||
f = clawhdf5.File.open_url(url, block_size=1024, max_parallel=1)
|
||||
ds = f["contig"]
|
||||
server.delay = 0.2
|
||||
server.delay_after = server.requests()
|
||||
old = sys.getswitchinterval()
|
||||
sys.setswitchinterval(0.001)
|
||||
stop = threading.Event()
|
||||
progress = {"n": 0, "worst": 0.0}
|
||||
|
||||
def spin():
|
||||
last = time.perf_counter()
|
||||
while not stop.is_set():
|
||||
now = time.perf_counter()
|
||||
progress["worst"] = max(progress["worst"], now - last)
|
||||
last = now
|
||||
progress["n"] += 1
|
||||
|
||||
t = threading.Thread(target=spin)
|
||||
try:
|
||||
t.start()
|
||||
time.sleep(0.02)
|
||||
t0 = time.perf_counter()
|
||||
before = server.requests()
|
||||
ds[0:2]
|
||||
took = time.perf_counter() - t0
|
||||
stop.set()
|
||||
t.join()
|
||||
finally:
|
||||
sys.setswitchinterval(old)
|
||||
assert server.requests() > before
|
||||
assert took >= 0.2, took
|
||||
assert progress["n"] > 1000, progress
|
||||
# Held across a 0.2 s request, the spinner would stall that long.
|
||||
assert progress["worst"] < 0.1, (progress, took)
|
||||
|
||||
|
||||
def test_a_clawhdf5_written_file_reads_the_same_remotely(tmp_path, range_server):
|
||||
path = tmp_path / "ours.h5"
|
||||
data = np.arange(3000, dtype="<f8").reshape(100, 30)
|
||||
with clawhdf5.File(str(path), "w") as f:
|
||||
f.create_dataset("d", data=data, chunks=[10, 30], compression="gzip")
|
||||
g = f.create_group("g")
|
||||
g.create_dataset("i", data=np.arange(5, dtype="<i4"))
|
||||
f.attrs["k"] = 3
|
||||
with clawhdf5.File(range_server.url("ours.h5")) as f, clawhdf5.File(str(path)) as local:
|
||||
np.testing.assert_array_equal(f["d"][...], data)
|
||||
np.testing.assert_array_equal(f["d"][5:9, ::4], local["d"][5:9, ::4])
|
||||
np.testing.assert_array_equal(f["g/i"][...], np.arange(5))
|
||||
assert f.attrs["k"] == local.attrs["k"]
|
||||
assert "file (read" in repr(f).lower() and "127.0.0.1" in repr(f)
|
||||
@@ -1509,3 +1509,68 @@ fn unencodable_filters_are_unsupported() {
|
||||
"a refused edit changed the file"
|
||||
);
|
||||
}
|
||||
|
||||
fn fixture(dir: &Path, name: &str) -> std::path::PathBuf {
|
||||
let path = dir.join(name);
|
||||
std::fs::copy(
|
||||
Path::new(env!("CARGO_MANIFEST_DIR"))
|
||||
.join("../clawhdf5/tests/fixtures")
|
||||
.join(name),
|
||||
&path,
|
||||
)
|
||||
.unwrap();
|
||||
path
|
||||
}
|
||||
|
||||
/// Zero extents on chunked datasets with no recorded maximum. The unfixed
|
||||
/// editor left `chunk_zero_extent_no_maxshape.h5` (a 2.7.0-written file
|
||||
/// resized to 1x0): a Fixed Array whose maximum, taken from the current
|
||||
/// dimensions, has no chunks along one dimension, so every stride before it
|
||||
/// is 0 — `h5rs check` panicked dividing by it and the next resize failed
|
||||
/// with an internal error. Such a file must check clean and resize on; a
|
||||
/// 2.7.0-written file taken through zero extents by the fixed editor must
|
||||
/// check clean at every step and read the fill value where it grew.
|
||||
#[test]
|
||||
fn zero_extent_resizes_without_a_recorded_maximum() {
|
||||
if !tools_ok() {
|
||||
return;
|
||||
}
|
||||
let dir = tmpdir();
|
||||
let path = fixture(dir.path(), "chunk_zero_extent_no_maxshape.h5");
|
||||
check_tools(&path, true);
|
||||
let mut ed = FileEditor::open(&path).unwrap();
|
||||
ed.resize("d", &[0, 0]).unwrap();
|
||||
ed.resize("z", &[0, 0]).unwrap();
|
||||
// Their maximum is now what the index was laid out by (1 x 0).
|
||||
ed.resize("d", &[1, 0]).unwrap();
|
||||
assert!(matches!(
|
||||
ed.resize("d", &[1, 1]),
|
||||
Err(Error::InvalidArgument(_))
|
||||
));
|
||||
drop(ed);
|
||||
check_tools(&path, true);
|
||||
|
||||
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
|
||||
for shape in [[15, 15], [3, 2], [1, 1], [1, 0], [0, 0], [0, 20], [20, 20]] {
|
||||
let mut ed = FileEditor::open(&path).unwrap();
|
||||
ed.resize("d", &shape).unwrap();
|
||||
ed.resize("z", &shape).unwrap();
|
||||
drop(ed);
|
||||
check_tools(&path, true);
|
||||
}
|
||||
let f = File::open(&path).unwrap();
|
||||
for name in ["d", "z"] {
|
||||
let d = f.dataset(name).unwrap();
|
||||
assert_eq!(d.shape().unwrap(), [20, 20]);
|
||||
assert!(d.read_f32().unwrap().iter().all(|&v| v == 0.0), "{name}");
|
||||
}
|
||||
assert_eq!(
|
||||
py(&format!(
|
||||
"import h5py\n\
|
||||
with h5py.File({:?}) as f:\n\
|
||||
\x20 print(int(abs(f['d'][()]).sum() + abs(f['z'][()]).sum()), f['d'].maxshape)",
|
||||
path.to_str().unwrap()
|
||||
)),
|
||||
"0 (20, 20)"
|
||||
);
|
||||
}
|
||||
|
||||
@@ -24,6 +24,8 @@ clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
|
||||
# Must match the wasm-bindgen CLI exactly; build.sh checks.
|
||||
wasm-bindgen = "0.2.129"
|
||||
js-sys = "0.3.106"
|
||||
# Promises for openUrl and RemoteFile (pure Rust over js-sys).
|
||||
wasm-bindgen-futures = "0.4.79"
|
||||
|
||||
[dev-dependencies]
|
||||
serde_json = "1"
|
||||
|
||||
@@ -0,0 +1,235 @@
|
||||
// HTTP for clawhdf5-wasm's openUrl (see src/lib.rs and src/lazy.rs).
|
||||
//
|
||||
// The Rust side decides which byte ranges a read needs; this file fetches
|
||||
// them with `fetch` and `Range` headers and checks every answer, so a server
|
||||
// that ignores the range, answers with other bytes, or serves a file that
|
||||
// changed since it was opened is an error, never data. wasm-bindgen copies
|
||||
// it into the package (pkg/snippets/...).
|
||||
|
||||
const DEFAULT_MAX_DOWNLOAD = 512 * 1024 * 1024;
|
||||
const DEFAULT_PARALLEL = 6;
|
||||
|
||||
function fetcher(opts) {
|
||||
const f = opts?.fetch ?? globalThis.fetch;
|
||||
if (typeof f !== "function") {
|
||||
throw new Error("openUrl: no fetch() in this environment (pass opts.fetch)");
|
||||
}
|
||||
return f;
|
||||
}
|
||||
|
||||
// The caller's headers (`opts.headers`: a Headers, [name, value] pairs or a
|
||||
// plain object, as fetch takes them; names come out in lower case) with
|
||||
// `extra` over them. A Range of the caller's is dropped: this file asks for
|
||||
// the ranges.
|
||||
function init(opts, extra, method = "GET", signal = undefined) {
|
||||
const headers = {};
|
||||
if (opts?.headers != null) {
|
||||
for (const [name, value] of new Headers(opts.headers)) {
|
||||
if (name !== "range") headers[name] = value;
|
||||
}
|
||||
}
|
||||
Object.assign(headers, extra);
|
||||
return { method, headers, credentials: opts?.credentials, signal };
|
||||
}
|
||||
|
||||
// `opts.parallel`: range requests in flight at once.
|
||||
function parallelism(opts) {
|
||||
const p = opts?.parallel ?? DEFAULT_PARALLEL;
|
||||
if (!Number.isSafeInteger(p) || p < 1) {
|
||||
throw new Error(`openUrl: parallel must be a positive integer, got ${String(p)}`);
|
||||
}
|
||||
return p;
|
||||
}
|
||||
|
||||
// "bytes a-b/total" -> { start, end (exclusive), total | null }; null when
|
||||
// the page cannot see the header (cross-origin, not exposed).
|
||||
function contentRange(resp, url) {
|
||||
const v = resp.headers.get("Content-Range");
|
||||
if (v === null) return null;
|
||||
const m = /^bytes (\d+)-(\d+)\/(\d+|\*)$/.exec(v.trim());
|
||||
if (!m) throw new Error(`${url}: the server sent an unusable Content-Range: ${v}`);
|
||||
return { start: Number(m[1]), end: Number(m[2]) + 1, total: m[3] === "*" ? null : Number(m[3]) };
|
||||
}
|
||||
|
||||
// What pins the file: its ETag, else its Last-Modified (null if neither is
|
||||
// visible to this page).
|
||||
function validatorOf(resp) {
|
||||
return resp.headers.get("ETag") ?? resp.headers.get("Last-Modified");
|
||||
}
|
||||
|
||||
async function discard(resp) {
|
||||
try {
|
||||
await resp.body?.cancel();
|
||||
} catch {
|
||||
// Nothing to release.
|
||||
}
|
||||
}
|
||||
|
||||
// The body, refusing more than `limit` bytes as they arrive: it is piped
|
||||
// through a TransformStream that errors the moment the count passes the
|
||||
// limit, which cancels the body and so aborts the request. Whatever the
|
||||
// server declares or sends, the page never holds more than `limit` bytes of
|
||||
// it. (`tooBig(n)` makes the error; n is the count so far.) A reader loop
|
||||
// would do the same, but it stalls on small bodies in headless Chromium
|
||||
// under --virtual-time-budget, which the page test uses; a pipe does not.
|
||||
async function readCapped(resp, limit, tooBig) {
|
||||
const declared = resp.headers.get("Content-Length");
|
||||
if (declared !== null && Number(declared) > limit) {
|
||||
await discard(resp);
|
||||
throw tooBig(declared);
|
||||
}
|
||||
if (!resp.body) {
|
||||
// No stream to read from (some fetch implementations): all at once.
|
||||
const all = new Uint8Array(await resp.arrayBuffer());
|
||||
if (all.length > limit) throw tooBig(all.length);
|
||||
return all;
|
||||
}
|
||||
let n = 0;
|
||||
let over = null;
|
||||
const capped = resp.body.pipeThrough(new TransformStream({
|
||||
transform(chunk, ctl) {
|
||||
n += chunk.length;
|
||||
if (n > limit) {
|
||||
over = tooBig(`over ${limit}`);
|
||||
ctl.error(over);
|
||||
return;
|
||||
}
|
||||
ctl.enqueue(chunk);
|
||||
},
|
||||
}));
|
||||
try {
|
||||
return new Uint8Array(await new Response(capped).arrayBuffer());
|
||||
} catch (e) {
|
||||
throw over ?? e;
|
||||
}
|
||||
}
|
||||
|
||||
// The whole body of a 200 answer (a server without range support), at
|
||||
// most `limit` (maxDownload) bytes.
|
||||
function readAll(resp, limit, url) {
|
||||
return readCapped(resp, limit, (n) =>
|
||||
new Error(`${url} is ${n} bytes, more than maxDownload (${limit}); ` +
|
||||
"the server does not support range requests, so the whole file would have to be downloaded"));
|
||||
}
|
||||
|
||||
// The body of a 206 answer, which may not be longer than the `limit` bytes
|
||||
// asked for at `start` (the caller checks the exact length).
|
||||
function readLimited(resp, limit, url, start) {
|
||||
return readCapped(resp, limit, () =>
|
||||
new Error(`${url}: asked for ${limit} bytes at offset ${start}, the server sent more`));
|
||||
}
|
||||
|
||||
/**
|
||||
* Ask for the file's first `firstLen` bytes. A server that honours the
|
||||
* range (206) gives `{ length, first, validator, requests }`; one that
|
||||
* answers 200 sends the whole file, which is kept (`{ whole, requests }`)
|
||||
* when `opts.fallback` is "download" (the default) and the file is at most
|
||||
* `opts.maxDownload` bytes, and is an error otherwise.
|
||||
*/
|
||||
export async function probe(url, firstLen, opts) {
|
||||
const f = fetcher(opts);
|
||||
parallelism(opts);
|
||||
const resp = await f(url, init(opts, { Range: `bytes=0-${firstLen - 1}` }));
|
||||
if (resp.status === 206) {
|
||||
const cr = contentRange(resp, url);
|
||||
if (cr && cr.start !== 0) {
|
||||
await discard(resp);
|
||||
throw new Error(`${url}: asked for bytes from 0, the server sent bytes from ${cr.start}`);
|
||||
}
|
||||
const first = await readLimited(resp, firstLen, url, 0);
|
||||
let length = cr?.total ?? null;
|
||||
let requests = 1;
|
||||
if (length === null) {
|
||||
// Content-Range is not readable here: a cross-origin server that does
|
||||
// not list it in Access-Control-Expose-Headers. Content-Length of a
|
||||
// HEAD request is always readable.
|
||||
const head = await f(url, init(opts, {}, "HEAD"));
|
||||
requests++;
|
||||
const cl = head.headers.get("Content-Length");
|
||||
if (!head.ok || cl === null) {
|
||||
throw new Error(`${url}: cannot learn the file's size (a cross-origin server must send ` +
|
||||
"Access-Control-Expose-Headers: Content-Range, or answer HEAD with Content-Length)");
|
||||
}
|
||||
length = Number(cl);
|
||||
}
|
||||
if (!Number.isSafeInteger(length) || length < 0) {
|
||||
throw new Error(`${url}: the server gave a file size of ${length} bytes; openUrl reads files ` +
|
||||
"of up to 2^53 - 1 bytes (the largest offset a JavaScript number holds exactly)");
|
||||
}
|
||||
if (first.length !== Math.min(firstLen, length)) {
|
||||
throw new Error(`${url}: asked for the first ${firstLen} bytes of ${length}, got ${first.length}`);
|
||||
}
|
||||
return { length, first, validator: validatorOf(resp), requests };
|
||||
}
|
||||
if (resp.status === 200) {
|
||||
if ((opts?.fallback ?? "download") !== "download") {
|
||||
await discard(resp);
|
||||
throw new Error(`${url}: the server does not support HTTP range requests (it answered 200 ` +
|
||||
"to a Range request); open it with { fallback: \"download\" } to download the whole file");
|
||||
}
|
||||
const whole = await readAll(resp, opts?.maxDownload ?? DEFAULT_MAX_DOWNLOAD, url);
|
||||
return { whole, requests: 1 };
|
||||
}
|
||||
await discard(resp);
|
||||
throw new Error(`${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
|
||||
}
|
||||
|
||||
/**
|
||||
* Fetch `ranges` ([start0, end0, start1, end1, ...], ends exclusive) of a
|
||||
* file opened by `probe`, at most `opts.parallel` (default 6) at a time.
|
||||
* Every answer must be a 206 with exactly the bytes asked for, from the same
|
||||
* file (validator and length). When one request fails, the others in
|
||||
* flight are aborted and no more are made; that failure is the error.
|
||||
*/
|
||||
export async function fetchRanges(url, ranges, opts, validator, length) {
|
||||
const f = fetcher(opts);
|
||||
const parallel = parallelism(opts);
|
||||
const n = ranges.length / 2;
|
||||
const out = new Array(n);
|
||||
const abort = new AbortController();
|
||||
let next = 0;
|
||||
async function one(i) {
|
||||
const start = ranges[2 * i];
|
||||
const end = ranges[2 * i + 1];
|
||||
const resp = await f(url, init(opts, { Range: `bytes=${start}-${end - 1}` }, "GET", abort.signal));
|
||||
if (resp.status !== 206) {
|
||||
await discard(resp);
|
||||
throw new Error(resp.status === 200
|
||||
? `${url}: the server stopped honouring range requests`
|
||||
: `${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
|
||||
}
|
||||
const cr = contentRange(resp, url);
|
||||
const v = validatorOf(resp);
|
||||
if ((validator != null && v !== null && v !== validator) ||
|
||||
(cr?.total != null && cr.total !== length)) {
|
||||
await discard(resp);
|
||||
throw new Error(`${url} changed on the server since it was opened`);
|
||||
}
|
||||
if (cr && (cr.start !== start || cr.end !== end)) {
|
||||
await discard(resp);
|
||||
throw new Error(`${url}: asked for bytes ${start}-${end - 1}, the server sent ${cr.start}-${cr.end - 1}`);
|
||||
}
|
||||
const body = await readLimited(resp, end - start, url, start);
|
||||
if (body.length !== end - start) {
|
||||
throw new Error(`${url}: asked for ${end - start} bytes at offset ${start}, got ${body.length}`);
|
||||
}
|
||||
out[i] = body;
|
||||
}
|
||||
async function worker() {
|
||||
while (next < n && !abort.signal.aborted) {
|
||||
try {
|
||||
await one(next++);
|
||||
} catch (e) {
|
||||
// The first failure stops the rest: requests in flight are aborted
|
||||
// (their AbortErrors are not reported) and no new ones start.
|
||||
if (!abort.signal.aborted) {
|
||||
abort.abort();
|
||||
throw e;
|
||||
}
|
||||
return;
|
||||
}
|
||||
}
|
||||
}
|
||||
await Promise.all(Array.from({ length: Math.min(parallel, n) }, worker));
|
||||
return out;
|
||||
}
|
||||
@@ -6,11 +6,25 @@
|
||||
//! with no typed-array mapping (compound, reference, opaque, ...) is refused
|
||||
//! with a message naming it, never returned as reinterpreted bytes.
|
||||
|
||||
use std::sync::Arc;
|
||||
|
||||
use clawhdf5::{AttrValue, File, Selection};
|
||||
use clawhdf5_format::data_read;
|
||||
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
||||
use clawhdf5_format::message_type::MessageType;
|
||||
use clawhdf5_format::object_header::ObjectHeader;
|
||||
use clawhdf5_format::storage::Storage;
|
||||
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
||||
|
||||
/// The most memory one read may use while it decodes: the stored bytes,
|
||||
/// the values at 64 bits (integers are widened first) and the values
|
||||
/// returned. A larger read fails with an error naming `readHyperslab`,
|
||||
/// before anything is read: on wasm32 a buffer past 2 GiB cannot be
|
||||
/// allocated at all, and failing to allocate aborts the module (every open
|
||||
/// file on the page with it). 1 GiB leaves room in wasm32's 4 GiB for the
|
||||
/// file's cached blocks and the JavaScript copy of the result.
|
||||
pub const MAX_READ_BYTES: u64 = 1 << 30;
|
||||
|
||||
/// Errors are reported to JavaScript as messages.
|
||||
pub type Result<T> = std::result::Result<T, String>;
|
||||
|
||||
@@ -119,7 +133,9 @@ pub struct Attr {
|
||||
pub value: AttrValue,
|
||||
}
|
||||
|
||||
/// An open file, held in memory.
|
||||
/// An open file: held in memory ([`Reader::open`]) or read through a
|
||||
/// [`Storage`] ([`Reader::open_storage`], such as a
|
||||
/// [`LazyStorage`](crate::lazy::LazyStorage)).
|
||||
pub struct Reader {
|
||||
file: File,
|
||||
}
|
||||
@@ -132,6 +148,14 @@ impl Reader {
|
||||
})
|
||||
}
|
||||
|
||||
/// Open a file read through `storage` (the file's bytes from offset 0,
|
||||
/// user block included, as [`File::open_storage`] takes them).
|
||||
pub fn open_storage(storage: Arc<dyn Storage + Send + Sync>) -> Result<Self> {
|
||||
Ok(Self {
|
||||
file: File::open_storage(storage).map_err(err)?,
|
||||
})
|
||||
}
|
||||
|
||||
/// Whether `path` names a group or a dataset.
|
||||
pub fn kind(&self, path: &str) -> Result<Kind> {
|
||||
match self.file.dataset(path) {
|
||||
@@ -144,31 +168,54 @@ impl Reader {
|
||||
/// The groups, then the datasets, in the group at `path` (`/` is the
|
||||
/// root). Soft links are listed as their targets; external and dangling
|
||||
/// links, and named datatypes, are left out.
|
||||
///
|
||||
/// What [`Group::groups`](clawhdf5::Group::groups) and `datasets` list,
|
||||
/// but every child's object header is read before an error ends the
|
||||
/// listing (the first error, in listing order, is the one returned, as
|
||||
/// there). Over a [`LazyStorage`](crate::lazy::LazyStorage) that makes
|
||||
/// one pass ask for all the headers it is missing at once, instead of
|
||||
/// one pass, and one round trip, per header.
|
||||
pub fn list(&self, path: &str) -> Result<Vec<Child>> {
|
||||
if self.kind(path)? != Kind::Group {
|
||||
return Err(format!("not a group: {path}"));
|
||||
}
|
||||
let group = self.file.group(path).map_err(err)?;
|
||||
let mut out: Vec<Child> = group
|
||||
.groups()
|
||||
.map_err(err)?
|
||||
.into_iter()
|
||||
.map(|name| Child {
|
||||
name,
|
||||
kind: Kind::Group,
|
||||
})
|
||||
.collect();
|
||||
out.extend(
|
||||
group
|
||||
.datasets()
|
||||
.map_err(err)?
|
||||
.into_iter()
|
||||
.map(|name| Child {
|
||||
name,
|
||||
kind: Kind::Dataset,
|
||||
}),
|
||||
);
|
||||
Ok(out)
|
||||
let entries = group.entries().map_err(err)?;
|
||||
let sb = self.file.superblock();
|
||||
let storage = self.file.storage();
|
||||
let mut groups = Vec::new();
|
||||
let mut datasets = Vec::new();
|
||||
let mut first_error = None;
|
||||
for (name, address) in entries {
|
||||
match ObjectHeader::parse_in(storage, address, sb.offset_size, sb.length_size) {
|
||||
Ok(header) => {
|
||||
let has = |t: MessageType| header.messages.iter().any(|m| m.msg_type == t);
|
||||
if has(MessageType::LinkInfo)
|
||||
|| has(MessageType::Link)
|
||||
|| has(MessageType::SymbolTable)
|
||||
{
|
||||
groups.push(Child {
|
||||
name: name.clone(),
|
||||
kind: Kind::Group,
|
||||
});
|
||||
}
|
||||
if has(MessageType::DataLayout) {
|
||||
datasets.push(Child {
|
||||
name,
|
||||
kind: Kind::Dataset,
|
||||
});
|
||||
}
|
||||
}
|
||||
Err(e) => {
|
||||
first_error.get_or_insert(e);
|
||||
}
|
||||
}
|
||||
}
|
||||
if let Some(e) = first_error {
|
||||
return Err(err(clawhdf5::Error::from(e)));
|
||||
}
|
||||
groups.extend(datasets);
|
||||
Ok(groups)
|
||||
}
|
||||
|
||||
/// Shape, max shape and datatype of the dataset at `path`.
|
||||
@@ -228,13 +275,27 @@ impl Reader {
|
||||
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
|
||||
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
|
||||
}
|
||||
let raw = ds.read_selection(&selection).map_err(err)?;
|
||||
let data = self.decode(&raw, &dt)?;
|
||||
out_shape.extend(element_shape(&dt));
|
||||
let expected = out_shape
|
||||
.iter()
|
||||
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
|
||||
.ok_or("selection size overflows")?;
|
||||
let cost = expected.saturating_mul(bytes_per_value(&dt));
|
||||
if cost > MAX_READ_BYTES {
|
||||
return Err(format!(
|
||||
"reading {path}{} would take about {} MiB of memory, more than the {} MiB \
|
||||
one read may use; read it in parts (readHyperslab)",
|
||||
if slab.is_some() {
|
||||
" (this selection)"
|
||||
} else {
|
||||
" whole"
|
||||
},
|
||||
cost >> 20,
|
||||
MAX_READ_BYTES >> 20
|
||||
));
|
||||
}
|
||||
let raw = ds.read_selection(&selection).map_err(err)?;
|
||||
let data = self.decode(&raw, &dt)?;
|
||||
if data.len() as u64 != expected {
|
||||
return Err(format!(
|
||||
"read {} values for shape {out_shape:?} ({expected} expected)",
|
||||
@@ -281,11 +342,14 @@ impl Reader {
|
||||
// string ends at its first NUL and a heap object of the
|
||||
// wrong size is an error, as in libhdf5 and h5py.
|
||||
let sb = self.file.superblock();
|
||||
Data::Strings(
|
||||
VlResolver::new(self.file.as_bytes(), sb.offset_size, sb.length_size)
|
||||
.strings(raw)
|
||||
.map_err(err)?,
|
||||
)
|
||||
let strings = match self.file.contiguous_bytes() {
|
||||
Some(bytes) => {
|
||||
VlResolver::new(bytes, sb.offset_size, sb.length_size).strings(raw)
|
||||
}
|
||||
None => VlResolver::new_in(self.file.storage(), sb.offset_size, sb.length_size)
|
||||
.strings(raw),
|
||||
};
|
||||
Data::Strings(strings.map_err(err)?)
|
||||
}
|
||||
Datatype::Enumeration { .. } if !is_array => {
|
||||
Data::Strings(data_read::read_enum_names(raw, dt).map_err(err)?)
|
||||
@@ -300,6 +364,27 @@ impl Reader {
|
||||
}
|
||||
}
|
||||
|
||||
/// Memory one value of type `dt` takes while [`Reader::read`] decodes it
|
||||
/// (an array type's elements count as values): its stored bytes, plus what
|
||||
/// [`Reader::decode`] builds from them. A string counts its `String` (24
|
||||
/// bytes on 64-bit targets, less on wasm32) and, for a fixed-length one,
|
||||
/// its text; a variable-length string's text lives in the heap and is
|
||||
/// bounded by the storage's own read limit.
|
||||
fn bytes_per_value(dt: &Datatype) -> u64 {
|
||||
let base = array_base(dt);
|
||||
let stored = u64::from(base.type_size());
|
||||
stored
|
||||
+ match base {
|
||||
Datatype::FloatingPoint { size, .. } if *size <= 4 => 4,
|
||||
Datatype::FloatingPoint { .. } => 8,
|
||||
// Widened to 64 bits, then narrowed to a new vector.
|
||||
Datatype::FixedPoint { .. } => 8 + stored,
|
||||
Datatype::String { .. } => 24 + stored,
|
||||
Datatype::VariableLength { .. } | Datatype::Enumeration { .. } => 24,
|
||||
_ => 0,
|
||||
}
|
||||
}
|
||||
|
||||
/// Narrow integers read at 64 bits to the dataset's own width. The source is
|
||||
/// that width, so this cannot fail on correct input; it is checked anyway.
|
||||
fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T>> {
|
||||
|
||||
@@ -0,0 +1,815 @@
|
||||
//! Reading a file that is not all here, when no read may wait for the
|
||||
//! network: the restartable "NeedBytes" mode of `docs/design/range-reads.md`
|
||||
//! (milestone M4).
|
||||
//!
|
||||
//! A browser's main thread cannot block on `fetch`, and the parsers are
|
||||
//! synchronous. So an operation (open, list a group, read a dataset) runs
|
||||
//! as a *pass* over a [`LazyStorage`] that holds the blocks fetched so far:
|
||||
//!
|
||||
//! 1. [`LazyStorage::attempt`] runs the operation. A read whose blocks are
|
||||
//! all present is served; a read that misses records the missing blocks
|
||||
//! and fails with a storage error.
|
||||
//! 2. If the pass missed anything, its result is thrown away — whatever it
|
||||
//! is, since a parser may have caught the error and carried on (a
|
||||
//! listing skips a link it cannot resolve) — and the caller gets the
|
||||
//! byte ranges to fetch ([`Step::Need`]).
|
||||
//! 3. The caller fetches them (asynchronously, with HTTP `Range` requests),
|
||||
//! hands them over with [`LazyStorage::supply`] and runs the operation
|
||||
//! again.
|
||||
//!
|
||||
//! A pass is pure over the storage: the facade only caches what completed
|
||||
//! reads decoded (its chunk cache), so re-running it is safe. Every pass
|
||||
//! that does not finish asks for at least one block not yet present, and no
|
||||
//! block is evicted while an operation is in flight
|
||||
//! ([`LazyStorage::operation`]), so an operation finishes after at most one
|
||||
//! pass per block it needs. In practice it is one pass per *wave* of
|
||||
//! misses: a chunked read asks for all the chunks of a batch at once.
|
||||
//!
|
||||
//! Blocks are kept in an LRU cache with a byte budget, trimmed only when no
|
||||
//! operation is in flight. Blocks fetched for bulk reads (raw data: a
|
||||
//! `read_ranges` call, or a read longer than a block) go first, so reading
|
||||
//! a large dataset does not evict the metadata.
|
||||
|
||||
use std::borrow::Cow;
|
||||
use std::collections::{BTreeSet, HashMap};
|
||||
use std::ops::Range;
|
||||
use std::sync::{Arc, Mutex, MutexGuard};
|
||||
|
||||
use clawhdf5_format::error::FormatError;
|
||||
use clawhdf5_format::storage::Storage;
|
||||
|
||||
/// Default block size: 1 MiB, as `clawhdf5-remote`'s block cache (the size
|
||||
/// `docs/design/range-reads.md` §2 measured).
|
||||
pub const DEFAULT_BLOCK_SIZE: u64 = 1 << 20;
|
||||
|
||||
/// Default of [`LazyConfig::max_fetch`]: 512 MiB, the same as `openUrl`'s
|
||||
/// `maxDownload` for a server without range support.
|
||||
pub const DEFAULT_MAX_FETCH: u64 = 512 << 20;
|
||||
|
||||
/// The message of the error a read that misses returns. It never reaches
|
||||
/// the caller of [`LazyStorage::attempt`]: a pass that missed is re-run.
|
||||
pub const NEED_BYTES: &str = "bytes not fetched yet (restartable read)";
|
||||
|
||||
/// Settings of a [`LazyStorage`].
|
||||
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||
pub struct LazyConfig {
|
||||
/// Size of a block in bytes (at least 512); fetches are whole, aligned
|
||||
/// blocks (the file's last block is shorter).
|
||||
pub block_size: u64,
|
||||
/// Byte budget of cached blocks between operations. An operation keeps
|
||||
/// every block it needs until it finishes, whatever the budget.
|
||||
pub capacity: u64,
|
||||
/// Largest single range asked for, in bytes (whole blocks, at least
|
||||
/// one); longer runs are split so they can be fetched in parallel.
|
||||
pub max_request: u64,
|
||||
/// Most bytes one operation may fetch (at least one block), and so the
|
||||
/// longest single read: a read longer than this fails at once, before
|
||||
/// anything is fetched, and so does an operation whose passes would
|
||||
/// fetch more. The file's length comes from the server, so without
|
||||
/// this a hostile file (a heap "collection" claiming 2 GiB) makes the
|
||||
/// reader fetch and hold whatever it names; on wasm32 a buffer past
|
||||
/// 2 GiB cannot even be allocated.
|
||||
pub max_fetch: u64,
|
||||
}
|
||||
|
||||
impl Default for LazyConfig {
|
||||
fn default() -> Self {
|
||||
LazyConfig {
|
||||
block_size: DEFAULT_BLOCK_SIZE,
|
||||
capacity: 64 << 20,
|
||||
max_request: 8 << 20,
|
||||
max_fetch: DEFAULT_MAX_FETCH,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// What a [`LazyStorage`] has done so far.
|
||||
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
|
||||
pub struct LazyStats {
|
||||
/// Passes run by [`LazyStorage::attempt`].
|
||||
pub passes: u64,
|
||||
/// Ranges handed to [`LazyStorage::supply`]: one HTTP request each.
|
||||
pub requests: u64,
|
||||
/// Bytes handed to [`LazyStorage::supply`].
|
||||
pub bytes_fetched: u64,
|
||||
/// Blocks evicted to stay within the budget.
|
||||
pub evictions: u64,
|
||||
/// Bytes cached now.
|
||||
pub cached_bytes: u64,
|
||||
}
|
||||
|
||||
/// The outcome of one pass.
|
||||
#[derive(Debug)]
|
||||
pub enum Step<T> {
|
||||
/// The pass read only bytes that were present: its result stands.
|
||||
Done(T),
|
||||
/// The pass missed: fetch these byte ranges (sorted, disjoint, block
|
||||
/// aligned), [`supply`](LazyStorage::supply) them and run it again.
|
||||
Need(Vec<Range<u64>>),
|
||||
}
|
||||
|
||||
struct Block {
|
||||
data: Arc<[u8]>,
|
||||
/// Eviction order: bulk blocks (`false`) before metadata (`true`),
|
||||
/// then least recently used first.
|
||||
key: (bool, u64),
|
||||
}
|
||||
|
||||
#[derive(Default)]
|
||||
struct State {
|
||||
blocks: HashMap<u64, Block>,
|
||||
/// `(metadata?, tick, block index)`, in eviction order.
|
||||
order: BTreeSet<(bool, u64, u64)>,
|
||||
tick: u64,
|
||||
bytes: u64,
|
||||
/// Blocks the current pass missed, and whether a small read wanted
|
||||
/// them (metadata).
|
||||
missing: HashMap<u64, bool>,
|
||||
/// Blocks a bulk read missed that have not been supplied yet: kept
|
||||
/// as bulk when they arrive.
|
||||
bulk_pending: BTreeSet<u64>,
|
||||
/// Operations in flight: no eviction while any is.
|
||||
active: u32,
|
||||
stats: LazyStats,
|
||||
}
|
||||
|
||||
/// A [`Storage`] over the blocks of a file fetched so far; a read of
|
||||
/// anything else fails and is recorded, so the pass can be re-run once the
|
||||
/// bytes arrive. See the [module documentation](self).
|
||||
pub struct LazyStorage {
|
||||
len: u64,
|
||||
config: LazyConfig,
|
||||
state: Mutex<State>,
|
||||
}
|
||||
|
||||
fn lock(m: &Mutex<State>) -> MutexGuard<'_, State> {
|
||||
m.lock().unwrap_or_else(std::sync::PoisonError::into_inner)
|
||||
}
|
||||
|
||||
/// Keeps an operation's blocks cached until it is dropped; see
|
||||
/// [`LazyStorage::operation`].
|
||||
pub struct Operation<'a> {
|
||||
storage: &'a LazyStorage,
|
||||
/// Bytes fetched for this operation so far.
|
||||
fetched: std::cell::Cell<u64>,
|
||||
}
|
||||
|
||||
impl Operation<'_> {
|
||||
/// Count `ranges` against the operation's budget
|
||||
/// ([`LazyConfig::max_fetch`]) before they are fetched: an error, and
|
||||
/// nothing counted, if they would take it past the budget.
|
||||
pub fn charge(&self, ranges: &[Range<u64>]) -> Result<(), String> {
|
||||
let max = self.storage.config.max_fetch;
|
||||
let total = ranges.iter().fold(self.fetched.get(), |n, r| {
|
||||
n.saturating_add(r.end.saturating_sub(r.start))
|
||||
});
|
||||
if total > max {
|
||||
return Err(format!(
|
||||
"this call would fetch more than {max} bytes of the file (the maxFetch limit); \
|
||||
read less at a time (readHyperslab) or raise maxFetch"
|
||||
));
|
||||
}
|
||||
self.fetched.set(total);
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
impl Drop for Operation<'_> {
|
||||
fn drop(&mut self) {
|
||||
let mut st = lock(&self.storage.state);
|
||||
st.active = st.active.saturating_sub(1);
|
||||
if st.active == 0 {
|
||||
self.storage.evict(&mut st);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl LazyStorage {
|
||||
/// An empty cache for a file of `len` bytes.
|
||||
pub fn new(len: u64, mut config: LazyConfig) -> Self {
|
||||
config.block_size = config.block_size.max(512);
|
||||
config.max_request = (config.max_request / config.block_size).max(1) * config.block_size;
|
||||
config.max_fetch = config.max_fetch.max(config.block_size);
|
||||
LazyStorage {
|
||||
len,
|
||||
config,
|
||||
state: Mutex::new(State::default()),
|
||||
}
|
||||
}
|
||||
|
||||
/// The settings in use (after rounding).
|
||||
pub fn config(&self) -> &LazyConfig {
|
||||
&self.config
|
||||
}
|
||||
|
||||
/// Counters since the storage was made.
|
||||
pub fn stats(&self) -> LazyStats {
|
||||
let st = lock(&self.state);
|
||||
LazyStats {
|
||||
cached_bytes: st.bytes,
|
||||
..st.stats
|
||||
}
|
||||
}
|
||||
|
||||
/// Mark an operation in flight until the guard is dropped: no block is
|
||||
/// evicted meanwhile, so re-running its passes always makes progress.
|
||||
/// Hold it across every pass of one operation.
|
||||
pub fn operation(&self) -> Operation<'_> {
|
||||
lock(&self.state).active += 1;
|
||||
Operation {
|
||||
storage: self,
|
||||
fetched: std::cell::Cell::new(0),
|
||||
}
|
||||
}
|
||||
|
||||
/// Run one pass of `f` over this storage. `Done` when `f` read nothing
|
||||
/// that is missing; otherwise `Need` with the ranges to fetch, and `f`'s
|
||||
/// result is dropped (it may be an error caused by the miss, or a
|
||||
/// result built around one).
|
||||
pub fn attempt<T>(&self, f: impl FnOnce() -> T) -> Step<T> {
|
||||
{
|
||||
let mut st = lock(&self.state);
|
||||
st.missing.clear();
|
||||
st.stats.passes += 1;
|
||||
}
|
||||
let out = f();
|
||||
let missing = std::mem::take(&mut lock(&self.state).missing);
|
||||
if missing.is_empty() {
|
||||
return Step::Done(out);
|
||||
}
|
||||
drop(out);
|
||||
Step::Need(self.runs(missing))
|
||||
}
|
||||
|
||||
/// The bytes of the file at `offset`, fetched for a range a pass asked
|
||||
/// for. `offset` must be block aligned and the bytes whole blocks (the
|
||||
/// last block of the file may be short) inside the file, or this is an
|
||||
/// error and nothing is kept. Blocks already present are left alone.
|
||||
pub fn supply(&self, offset: u64, bytes: &[u8]) -> Result<(), String> {
|
||||
let bs = self.config.block_size;
|
||||
let end = offset
|
||||
.checked_add(bytes.len() as u64)
|
||||
.filter(|&e| e <= self.len)
|
||||
.ok_or_else(|| {
|
||||
format!(
|
||||
"{} bytes at offset {offset} run past the end of the {}-byte file",
|
||||
bytes.len(),
|
||||
self.len
|
||||
)
|
||||
})?;
|
||||
if !offset.is_multiple_of(bs) || (!end.is_multiple_of(bs) && end != self.len) {
|
||||
return Err(format!(
|
||||
"{} bytes at offset {offset} are not whole {bs}-byte blocks",
|
||||
bytes.len()
|
||||
));
|
||||
}
|
||||
let mut st = lock(&self.state);
|
||||
st.stats.requests += 1;
|
||||
st.stats.bytes_fetched += bytes.len() as u64;
|
||||
let mut start = offset;
|
||||
while start < end {
|
||||
let i = start / bs;
|
||||
let stop = (start + bs).min(end);
|
||||
if !st.blocks.contains_key(&i) {
|
||||
let rel = (start - offset) as usize..(stop - offset) as usize;
|
||||
let metadata = !st.bulk_pending.remove(&i);
|
||||
self.keep(&mut st, i, Arc::from(&bytes[rel]), metadata);
|
||||
}
|
||||
start = stop;
|
||||
}
|
||||
if st.active == 0 {
|
||||
self.evict(&mut st);
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// [`supply`](Self::supply) the bytes fetched for `range`, one of the
|
||||
/// ranges a [`Step::Need`] asked for: anything but exactly its length
|
||||
/// (a server that answered with more or less) is an error.
|
||||
pub fn supply_range(&self, range: &Range<u64>, bytes: &[u8]) -> Result<(), String> {
|
||||
let want = range.end.saturating_sub(range.start);
|
||||
if bytes.len() as u64 != want {
|
||||
return Err(format!(
|
||||
"asked for {want} bytes at offset {}, got {}",
|
||||
range.start,
|
||||
bytes.len()
|
||||
));
|
||||
}
|
||||
self.supply(range.start, bytes)
|
||||
}
|
||||
|
||||
/// Run `f` to completion, fetching what its passes miss with `fetch`
|
||||
/// (a byte range to its bytes). The blocking driver, for native code
|
||||
/// and tests; the browser's is the same loop with an `await` between
|
||||
/// passes.
|
||||
pub fn run_blocking<T>(
|
||||
&self,
|
||||
mut f: impl FnMut() -> T,
|
||||
mut fetch: impl FnMut(Range<u64>) -> Result<Vec<u8>, String>,
|
||||
) -> Result<T, String> {
|
||||
let op = self.operation();
|
||||
loop {
|
||||
match self.attempt(&mut f) {
|
||||
Step::Done(v) => return Ok(v),
|
||||
Step::Need(ranges) => {
|
||||
op.charge(&ranges)?;
|
||||
for r in ranges {
|
||||
let bytes = fetch(r.clone())?;
|
||||
self.supply_range(&r, &bytes)?;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Cache block `i`.
|
||||
fn keep(&self, st: &mut State, i: u64, data: Arc<[u8]>, metadata: bool) {
|
||||
st.tick += 1;
|
||||
let key = (metadata, st.tick);
|
||||
st.bytes += <[u8]>::len(&data) as u64;
|
||||
st.order.insert((key.0, key.1, i));
|
||||
if let Some(old) = st.blocks.insert(i, Block { data, key }) {
|
||||
st.order.remove(&(old.key.0, old.key.1, i));
|
||||
st.bytes -= <[u8]>::len(&old.data) as u64;
|
||||
}
|
||||
}
|
||||
|
||||
fn evict(&self, st: &mut State) {
|
||||
while st.bytes > self.config.capacity {
|
||||
let Some((_, _, i)) = st.order.pop_first() else {
|
||||
break;
|
||||
};
|
||||
if let Some(b) = st.blocks.remove(&i) {
|
||||
st.bytes -= <[u8]>::len(&b.data) as u64;
|
||||
st.stats.evictions += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Byte ranges covering the missing blocks: runs of consecutive
|
||||
/// blocks, a one-block hole between two runs filled so they merge
|
||||
/// (unless the hole is cached: it would be fetched again), each at
|
||||
/// most `max_request` long.
|
||||
fn runs(&self, missing: HashMap<u64, bool>) -> Vec<Range<u64>> {
|
||||
let bs = self.config.block_size;
|
||||
let mut wanted: Vec<u64> = missing.keys().copied().collect();
|
||||
wanted.sort_unstable();
|
||||
let mut st = lock(&self.state);
|
||||
// Remember which blocks only bulk reads asked for: they are kept
|
||||
// as bulk once supplied.
|
||||
for (&i, &metadata) in &missing {
|
||||
if metadata {
|
||||
st.bulk_pending.remove(&i);
|
||||
} else {
|
||||
st.bulk_pending.insert(i);
|
||||
}
|
||||
}
|
||||
let per_request = self.config.max_request / bs;
|
||||
let mut runs: Vec<(u64, u64)> = Vec::new();
|
||||
for i in wanted {
|
||||
match runs.last_mut() {
|
||||
Some((first, last))
|
||||
if (i == *last + 1
|
||||
|| (i == *last + 2 && !st.blocks.contains_key(&(i - 1))))
|
||||
&& i - *first < per_request =>
|
||||
{
|
||||
*last = i
|
||||
}
|
||||
_ => runs.push((i, i)),
|
||||
}
|
||||
}
|
||||
drop(st);
|
||||
runs.into_iter()
|
||||
.map(|(a, b)| a * bs..((b + 1) * bs).min(self.len))
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Block indices covering `[offset, offset + len)`, clamped to the file.
|
||||
fn span(&self, offset: u64, len: u64) -> Option<Range<u64>> {
|
||||
let end = offset.saturating_add(len).min(self.len);
|
||||
if offset >= end {
|
||||
return None;
|
||||
}
|
||||
let bs = self.config.block_size;
|
||||
Some(offset / bs..(end - 1) / bs + 1)
|
||||
}
|
||||
|
||||
/// The blocks of `spans` if all are present (touching them), else
|
||||
/// record the missing ones and fail.
|
||||
fn blocks(
|
||||
&self,
|
||||
spans: &[Range<u64>],
|
||||
metadata: bool,
|
||||
) -> Result<HashMap<u64, Arc<[u8]>>, FormatError> {
|
||||
let mut st = lock(&self.state);
|
||||
let mut have = HashMap::new();
|
||||
let mut missed = false;
|
||||
for span in spans {
|
||||
for i in span.clone() {
|
||||
if have.contains_key(&i) {
|
||||
continue;
|
||||
}
|
||||
match st.blocks.get(&i) {
|
||||
Some(b) => {
|
||||
have.insert(i, b.data.clone());
|
||||
}
|
||||
None => {
|
||||
missed = true;
|
||||
let m = st.missing.entry(i).or_insert(metadata);
|
||||
*m |= metadata;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
if missed {
|
||||
return Err(FormatError::Storage(NEED_BYTES.into()));
|
||||
}
|
||||
// Touch: most recently used last; a small read promotes a bulk
|
||||
// block to metadata.
|
||||
for &i in have.keys() {
|
||||
st.tick += 1;
|
||||
let tick = st.tick;
|
||||
let Some(b) = st.blocks.get_mut(&i) else {
|
||||
continue;
|
||||
};
|
||||
let old = b.key;
|
||||
b.key = (old.0 || metadata, tick);
|
||||
let new = b.key;
|
||||
st.order.remove(&(old.0, old.1, i));
|
||||
st.order.insert((new.0, new.1, i));
|
||||
}
|
||||
Ok(have)
|
||||
}
|
||||
|
||||
/// Refuse a read of `n` bytes longer than an operation may fetch
|
||||
/// ([`LazyConfig::max_fetch`]), before its blocks are asked for.
|
||||
fn check_len(&self, n: u64) -> Result<(), FormatError> {
|
||||
let max = self.config.max_fetch;
|
||||
if n > max {
|
||||
return Err(FormatError::Storage(format!(
|
||||
"a read of {n} bytes is more than one call may fetch ({max} bytes, the maxFetch limit)"
|
||||
)));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// The bytes `offset..end` from `blocks`, which hold every block of
|
||||
/// that span. The buffer is reserved fallibly: a length the address
|
||||
/// space cannot hold (past `isize::MAX` on wasm32) is an error, never
|
||||
/// an abort.
|
||||
fn assemble(
|
||||
&self,
|
||||
offset: u64,
|
||||
end: u64,
|
||||
blocks: &HashMap<u64, Arc<[u8]>>,
|
||||
) -> Result<Vec<u8>, FormatError> {
|
||||
let bs = self.config.block_size;
|
||||
let n = end - offset;
|
||||
let too_long =
|
||||
|| FormatError::Storage(format!("cannot hold a read of {n} bytes in memory"));
|
||||
let mut out = Vec::new();
|
||||
out.try_reserve_exact(usize::try_from(n).map_err(|_| too_long())?)
|
||||
.map_err(|_| too_long())?;
|
||||
let mut pos = offset;
|
||||
while pos < end {
|
||||
let i = pos / bs;
|
||||
let block = &blocks[&i];
|
||||
let from = (pos - i * bs) as usize;
|
||||
let to = ((end - i * bs) as usize).min(<[u8]>::len(block));
|
||||
out.extend_from_slice(&block[from..to]);
|
||||
pos = i * bs + to as u64;
|
||||
}
|
||||
Ok(out)
|
||||
}
|
||||
}
|
||||
|
||||
impl Storage for LazyStorage {
|
||||
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
|
||||
let Some(span) = self.span(offset, len as u64) else {
|
||||
return Ok(Cow::Owned(Vec::new()));
|
||||
};
|
||||
let end = offset.saturating_add(len as u64).min(self.len);
|
||||
self.check_len(end - offset)?;
|
||||
let metadata = len as u64 <= self.config.block_size;
|
||||
let blocks = self.blocks(std::slice::from_ref(&span), metadata)?;
|
||||
Ok(Cow::Owned(self.assemble(offset, end, &blocks)?))
|
||||
}
|
||||
|
||||
fn len(&self) -> u64 {
|
||||
self.len
|
||||
}
|
||||
|
||||
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
|
||||
let mut spans = Vec::with_capacity(ranges.len());
|
||||
let mut total = 0u64;
|
||||
for r in ranges {
|
||||
if r.end < r.start {
|
||||
return Err(FormatError::Storage(
|
||||
"read range ends before it starts".into(),
|
||||
));
|
||||
}
|
||||
total = total.saturating_add(r.end.min(self.len).saturating_sub(r.start));
|
||||
spans.extend(self.span(r.start, r.end - r.start));
|
||||
}
|
||||
self.check_len(total)?;
|
||||
let blocks = self.blocks(&spans, false)?;
|
||||
ranges
|
||||
.iter()
|
||||
.map(|r| {
|
||||
let end = r.end.min(self.len);
|
||||
if r.start >= end {
|
||||
Ok(Cow::Owned(Vec::new()))
|
||||
} else {
|
||||
self.assemble(r.start, end, &blocks).map(Cow::Owned)
|
||||
}
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
#[allow(clippy::single_range_in_vec_init)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn file(n: usize) -> Vec<u8> {
|
||||
(0..n).map(|i| (i * 7 + i / 251) as u8).collect()
|
||||
}
|
||||
|
||||
fn config(block: u64, capacity: u64) -> LazyConfig {
|
||||
LazyConfig {
|
||||
block_size: block,
|
||||
capacity,
|
||||
max_request: 4 * block,
|
||||
max_fetch: DEFAULT_MAX_FETCH,
|
||||
}
|
||||
}
|
||||
|
||||
/// Supply every range of `need` from `data`.
|
||||
fn serve(s: &LazyStorage, data: &[u8], need: &[Range<u64>]) {
|
||||
for r in need {
|
||||
s.supply_range(r, &data[r.start as usize..r.end as usize])
|
||||
.unwrap();
|
||||
}
|
||||
}
|
||||
|
||||
fn owned(r: Result<Cow<'_, [u8]>, FormatError>) -> Result<Vec<u8>, FormatError> {
|
||||
r.map(Cow::into_owned)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_miss_asks_for_whole_blocks_then_the_rerun_reads_them() {
|
||||
let data = file(10_000);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||
let Step::Need(need) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
|
||||
panic!("nothing is cached yet");
|
||||
};
|
||||
assert_eq!(need, vec![1024..3072]);
|
||||
serve(&s, &data, &need);
|
||||
let Step::Done(got) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
|
||||
panic!("the blocks were supplied");
|
||||
};
|
||||
assert_eq!(got.unwrap(), &data[1500..2500]);
|
||||
// Past the end: short, then empty, as for a slice.
|
||||
serve(&s, &data, &[9216..10_000]);
|
||||
let Step::Done(tail) = s.attempt(|| owned(s.read_at(9_990, 100))) else {
|
||||
panic!("the last block was supplied");
|
||||
};
|
||||
assert_eq!(tail.unwrap(), &data[9_990..]);
|
||||
assert!(matches!(
|
||||
s.attempt(|| s.read_at(20_000, 10).map(|c| c.len())),
|
||||
Step::Done(Ok(0))
|
||||
));
|
||||
let st = s.stats();
|
||||
assert_eq!(
|
||||
(st.passes, st.requests, st.bytes_fetched),
|
||||
(4, 2, 2048 + 784)
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_pass_that_swallowed_the_miss_is_still_rerun() {
|
||||
// A parser that catches the error and returns something anyway
|
||||
// (a listing skipping a link it cannot resolve) must not have its
|
||||
// result used.
|
||||
let data = file(4096);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||
let step = s.attempt(|| s.read_at(0, 4).map(|b| b.len()).unwrap_or(0));
|
||||
assert!(
|
||||
matches!(step, Step::Need(ref n) if n == &vec![0..1024]),
|
||||
"{step:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_ranges_asks_for_every_missing_block_in_one_pass() {
|
||||
let data = file(64 * 1024);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||
let ranges = [
|
||||
100..200,
|
||||
5000..5100,
|
||||
5200..5300,
|
||||
30_000..33_000,
|
||||
60_000..60_010,
|
||||
];
|
||||
let read = || {
|
||||
s.read_ranges(&ranges)
|
||||
.map(|v| v.into_iter().map(Cow::into_owned).collect::<Vec<_>>())
|
||||
};
|
||||
let Step::Need(need) = s.attempt(read) else {
|
||||
panic!("nothing is cached yet");
|
||||
};
|
||||
// 5000..5300 is blocks 4 and 5; 30_000..33_000 is blocks 29..=32.
|
||||
assert_eq!(
|
||||
need,
|
||||
vec![
|
||||
0..1024,
|
||||
4096..6144,
|
||||
29 * 1024..33 * 1024,
|
||||
58 * 1024..59 * 1024
|
||||
]
|
||||
);
|
||||
serve(&s, &data, &need);
|
||||
let Step::Done(got) = s.attempt(read) else {
|
||||
panic!("one pass fetched everything");
|
||||
};
|
||||
for (r, g) in ranges.iter().zip(&got.unwrap()) {
|
||||
assert_eq!(g, &data[r.start as usize..r.end as usize]);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn runs_merge_one_block_holes_and_split_long_runs() {
|
||||
let data = file(32 * 1024);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||
// Blocks 0 and 2 (a hole of one: merged), 5..=14 (split in fours).
|
||||
let Step::Need(need) = s.attempt(|| {
|
||||
let _ = s.read_at(0, 10);
|
||||
let _ = s.read_at(2048, 10);
|
||||
s.read_ranges(&[5120..15 * 1024]).map(|_| ())
|
||||
}) else {
|
||||
panic!("nothing is cached yet");
|
||||
};
|
||||
assert_eq!(
|
||||
need,
|
||||
vec![0..3072, 5120..9216, 9216..13_312, 13_312..15_360]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_cached_hole_is_not_fetched_again() {
|
||||
let data = file(8 * 1024);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||
serve(&s, &data, &[1024..2048]);
|
||||
// Blocks 0 and 2 missing, 1 cached: two requests, not 0..3072.
|
||||
let Step::Need(need) = s.attempt(|| {
|
||||
let _ = s.read_at(0, 10);
|
||||
s.read_at(2048, 10).map(|_| ())
|
||||
}) else {
|
||||
panic!("blocks 0 and 2 are missing");
|
||||
};
|
||||
assert_eq!(need, vec![0..1024, 2048..3072]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn supply_refuses_what_was_not_asked_for() {
|
||||
let data = file(10_000);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||
assert!(s.supply(1, &data[1..1025]).unwrap_err().contains("whole"));
|
||||
assert!(s.supply(0, &data[..1000]).unwrap_err().contains("whole"));
|
||||
assert!(
|
||||
s.supply(9216, &[0u8; 1024])
|
||||
.unwrap_err()
|
||||
.contains("past the end")
|
||||
);
|
||||
for wrong in [&data[..1023], &data[..2048]] {
|
||||
assert!(
|
||||
s.supply_range(&(0..1024), wrong)
|
||||
.unwrap_err()
|
||||
.contains("asked for 1024 bytes")
|
||||
);
|
||||
}
|
||||
assert_eq!(s.stats().cached_bytes, 0);
|
||||
// The file's last, short block is whole.
|
||||
s.supply(9216, &data[9216..]).unwrap();
|
||||
assert_eq!(s.stats().cached_bytes, 784);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_operation_keeps_its_blocks_whatever_the_budget() {
|
||||
// A budget of one block, an operation that needs eight: without
|
||||
// the operation guard every pass would evict what the last one
|
||||
// fetched and never finish.
|
||||
let data = file(8 * 1024);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1024));
|
||||
let got = s
|
||||
.run_blocking(
|
||||
|| {
|
||||
(0..8)
|
||||
.map(|i| owned(s.read_at(i * 1024 + 3, 10)))
|
||||
.collect::<Result<Vec<_>, _>>()
|
||||
},
|
||||
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
|
||||
)
|
||||
.unwrap()
|
||||
.unwrap();
|
||||
for (i, g) in got.iter().enumerate() {
|
||||
assert_eq!(g, &data[i * 1024 + 3..i * 1024 + 13]);
|
||||
}
|
||||
// Trimmed to the budget once the operation is over.
|
||||
let st = s.stats();
|
||||
assert_eq!(st.cached_bytes, 1024);
|
||||
assert_eq!(st.evictions, 7);
|
||||
assert_eq!(st.passes, 9, "one pass per block, then the one that ends");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn bulk_blocks_are_evicted_before_metadata() {
|
||||
let data = file(16 * 1024);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 4 * 1024));
|
||||
let fetch = |r: Range<u64>| Ok(data[r.start as usize..r.end as usize].to_vec());
|
||||
// Metadata: a small read of block 0.
|
||||
s.run_blocking(|| s.read_at(0, 16).map(|_| ()), fetch)
|
||||
.unwrap()
|
||||
.unwrap();
|
||||
// Bulk: raw data over blocks 4..12, more than the budget.
|
||||
s.run_blocking(|| s.read_ranges(&[4096..12 * 1024]).map(|_| ()), fetch)
|
||||
.unwrap()
|
||||
.unwrap();
|
||||
// The metadata block survived: reading it again is a hit.
|
||||
let before = s.stats();
|
||||
assert!(before.cached_bytes <= 4 * 1024);
|
||||
assert!(matches!(
|
||||
s.attempt(|| s.read_at(0, 16).map(|_| ())),
|
||||
Step::Done(Ok(()))
|
||||
));
|
||||
assert_eq!(s.stats().requests, before.requests);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_read_longer_than_max_fetch_fails_without_fetching() {
|
||||
// A hostile file names a 2 GiB heap collection in a "file" the
|
||||
// server claims is 1 TiB: the read is refused before any block is
|
||||
// asked for (on wasm32 its buffer could not even be allocated).
|
||||
let s = LazyStorage::new(1 << 40, LazyConfig::default());
|
||||
let step = s.attempt(|| s.read_at(4096, (1usize << 31) + 4096).map(|b| b.len()));
|
||||
match step {
|
||||
Step::Done(Err(e)) => assert!(e.to_string().contains("maxFetch"), "{e}"),
|
||||
other => panic!("expected a refusal, got {other:?}"),
|
||||
}
|
||||
let step = s.attempt(|| s.read_ranges(&[0..(600 << 20)]).map(|v| v.len()));
|
||||
assert!(matches!(step, Step::Done(Err(_))), "{step:?}");
|
||||
assert_eq!(s.stats().requests, 0);
|
||||
// At the limit it is an ordinary miss.
|
||||
let s = LazyStorage::new(1 << 40, config(1024, 1 << 20));
|
||||
let step = s.attempt(|| s.read_at(0, DEFAULT_MAX_FETCH as usize).map(|b| b.len()));
|
||||
assert!(matches!(step, Step::Need(_)), "{step:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_operation_stops_at_its_fetch_budget() {
|
||||
// Many small reads, none over the limit, that together would fetch
|
||||
// more than the budget: the operation fails before fetching past it.
|
||||
let data = file(64 * 1024);
|
||||
let mut c = config(1024, 1 << 20);
|
||||
c.max_fetch = 8 * 1024;
|
||||
let s = LazyStorage::new(data.len() as u64, c);
|
||||
let e = s
|
||||
.run_blocking(
|
||||
|| {
|
||||
(0..64)
|
||||
.map(|i| owned(s.read_at(i * 1024, 8)))
|
||||
.collect::<Result<Vec<_>, _>>()
|
||||
},
|
||||
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
|
||||
)
|
||||
.unwrap_err();
|
||||
assert!(e.contains("maxFetch"), "{e}");
|
||||
assert!(s.stats().bytes_fetched <= 8 * 1024, "{:?}", s.stats());
|
||||
// Within the budget it completes, and the budget is per operation.
|
||||
for _ in 0..3 {
|
||||
s.run_blocking(
|
||||
|| owned(s.read_at(10 * 1024, 3000)),
|
||||
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
|
||||
)
|
||||
.unwrap()
|
||||
.unwrap();
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_failed_fetch_is_an_error_not_data() {
|
||||
let data = file(4096);
|
||||
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||
let e = s
|
||||
.run_blocking(|| owned(s.read_at(0, 8)), |_| Err("HTTP 500".into()))
|
||||
.unwrap_err();
|
||||
assert_eq!(e, "HTTP 500");
|
||||
let e = s
|
||||
.run_blocking(|| owned(s.read_at(0, 8)), |_| Ok(vec![0; 10]))
|
||||
.unwrap_err();
|
||||
assert!(e.contains("asked for 1024 bytes"), "{e}");
|
||||
assert_eq!(s.stats().cached_bytes, 0);
|
||||
drop(data);
|
||||
}
|
||||
}
|
||||
+476
-62
@@ -1,7 +1,7 @@
|
||||
//! clawhdf5's HDF5 reader for JavaScript, via `wasm-bindgen`.
|
||||
//!
|
||||
//! ```js
|
||||
//! import init, { open } from "./pkg/clawhdf5_wasm.js";
|
||||
//! import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
|
||||
//! await init();
|
||||
//! const file = open(new Uint8Array(await blob.arrayBuffer()));
|
||||
//! file.list("/"); // [{ name, kind: "group" | "dataset" }]
|
||||
@@ -10,6 +10,12 @@
|
||||
//! file.read("/x"); // { shape, dtype, data: Float64Array | ... | string[] }
|
||||
//! file.readHyperslab("/x", [0, 0], [10, 10]); // stride, block optional
|
||||
//! file.free();
|
||||
//!
|
||||
//! // A file on a web server, read by HTTP range requests as needed: the
|
||||
//! // same methods, returning promises.
|
||||
//! const remote = await openUrl("https://example.org/data.h5");
|
||||
//! await remote.list("/");
|
||||
//! remote.stats(); // { requests, bytesFetched, size, ... }
|
||||
//! ```
|
||||
//!
|
||||
//! Numeric data comes back in the typed array of the stored width
|
||||
@@ -18,15 +24,23 @@
|
||||
//! Anything else is a thrown `Error` naming the datatype. Only the reader is
|
||||
//! exposed: nothing here writes files.
|
||||
//!
|
||||
//! The logic lives in [`core`], which is plain Rust and tested natively.
|
||||
//! The logic lives in [`core`] and [`lazy`], which are plain Rust and tested
|
||||
//! natively; `js/remote.js` does the HTTP.
|
||||
|
||||
pub mod core;
|
||||
pub mod lazy;
|
||||
|
||||
use std::ops::Range;
|
||||
use std::rc::Rc;
|
||||
use std::sync::Arc;
|
||||
|
||||
use clawhdf5::AttrValue;
|
||||
use js_sys::{Array, Object, Reflect};
|
||||
use js_sys::{Array, Object, Promise, Reflect, Uint8Array};
|
||||
use wasm_bindgen::prelude::*;
|
||||
use wasm_bindgen_futures::future_to_promise;
|
||||
|
||||
use crate::core::{Data, Hyperslab, Reader};
|
||||
use crate::core::{Attr, Child, Data, DatasetInfo, Hyperslab, Reader};
|
||||
use crate::lazy::{LazyConfig, LazyStorage, Step};
|
||||
|
||||
/// JavaScript numbers are exact up to 2^53.
|
||||
const MAX_SAFE_INTEGER: f64 = 9_007_199_254_740_991.0;
|
||||
@@ -35,11 +49,30 @@ fn js_err(msg: String) -> JsError {
|
||||
JsError::new(&msg)
|
||||
}
|
||||
|
||||
/// A JavaScript exception (from `fetch`, or `js/remote.js`) as a `JsError`
|
||||
/// with its message.
|
||||
fn js_exception(e: JsValue) -> JsError {
|
||||
let msg = e
|
||||
.dyn_ref::<js_sys::Error>()
|
||||
.map(|e| String::from(e.message()))
|
||||
.or_else(|| e.as_string())
|
||||
.unwrap_or_else(|| format!("{e:?}"));
|
||||
JsError::new(&msg)
|
||||
}
|
||||
|
||||
fn set(obj: &Object, key: &str, value: impl Into<JsValue>) {
|
||||
// Defining a property on a fresh plain object cannot fail.
|
||||
Reflect::set(obj, &JsValue::from_str(key), &value.into()).unwrap_throw();
|
||||
}
|
||||
|
||||
fn get(obj: &JsValue, key: &str) -> JsValue {
|
||||
if obj.is_object() {
|
||||
Reflect::get(obj, &JsValue::from_str(key)).unwrap_or(JsValue::UNDEFINED)
|
||||
} else {
|
||||
JsValue::UNDEFINED
|
||||
}
|
||||
}
|
||||
|
||||
fn shape_to_js(shape: &[u64]) -> Array {
|
||||
shape.iter().map(|&d| JsValue::from_f64(d as f64)).collect()
|
||||
}
|
||||
@@ -58,6 +91,20 @@ fn indices_from_js(what: &str, v: &[f64]) -> Result<Vec<u64>, JsError> {
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn slab_from_js(
|
||||
start: &[f64],
|
||||
count: &[f64],
|
||||
stride: Option<Vec<f64>>,
|
||||
block: Option<Vec<f64>>,
|
||||
) -> Result<Hyperslab, JsError> {
|
||||
Ok(Hyperslab {
|
||||
start: indices_from_js("start", start)?,
|
||||
count: indices_from_js("count", count)?,
|
||||
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
|
||||
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
|
||||
})
|
||||
}
|
||||
|
||||
fn data_to_js(data: Data) -> JsValue {
|
||||
match data {
|
||||
Data::F32(v) => js_sys::Float32Array::from(&v[..]).into(),
|
||||
@@ -110,6 +157,72 @@ fn attr_to_js(value: AttrValue) -> (JsValue, Option<String>) {
|
||||
}
|
||||
}
|
||||
|
||||
fn list_to_js(children: Vec<Child>) -> Array {
|
||||
children
|
||||
.into_iter()
|
||||
.map(|c| {
|
||||
let o = Object::new();
|
||||
set(&o, "name", c.name);
|
||||
set(&o, "kind", c.kind.as_str());
|
||||
JsValue::from(o)
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn info_to_js(i: DatasetInfo) -> Object {
|
||||
let o = Object::new();
|
||||
set(&o, "shape", shape_to_js(&i.shape));
|
||||
let max: JsValue = match i.maxshape {
|
||||
None => JsValue::NULL,
|
||||
Some(dims) => dims
|
||||
.into_iter()
|
||||
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
|
||||
.collect::<Array>()
|
||||
.into(),
|
||||
};
|
||||
set(&o, "maxshape", max);
|
||||
set(&o, "dtype", i.dtype);
|
||||
set(&o, "elementShape", shape_to_js(&i.element_shape));
|
||||
o
|
||||
}
|
||||
|
||||
fn attrs_to_js(attrs: Vec<Attr>) -> Array {
|
||||
attrs
|
||||
.into_iter()
|
||||
.map(|a| {
|
||||
let o = Object::new();
|
||||
set(&o, "name", a.name);
|
||||
let (value, dtype) = attr_to_js(a.value);
|
||||
set(&o, "value", value);
|
||||
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
|
||||
JsValue::from(o)
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn errors_to_js(errors: Vec<String>) -> Array {
|
||||
errors.into_iter().map(JsValue::from).collect()
|
||||
}
|
||||
|
||||
/// A read's values with the dataset's datatype.
|
||||
fn values_to_js((dtype, a): (String, core::Array)) -> Object {
|
||||
let o = Object::new();
|
||||
set(&o, "shape", shape_to_js(&a.shape));
|
||||
set(&o, "dtype", dtype);
|
||||
set(&o, "data", data_to_js(a.data));
|
||||
o
|
||||
}
|
||||
|
||||
/// Read the dataset at `path` (whole, or `slab`) with its datatype.
|
||||
fn read_values(
|
||||
r: &Reader,
|
||||
path: &str,
|
||||
slab: Option<&Hyperslab>,
|
||||
) -> core::Result<(String, core::Array)> {
|
||||
let dtype = r.info(path)?.dtype;
|
||||
Ok((dtype, r.read(path, slab)?))
|
||||
}
|
||||
|
||||
/// An open HDF5 (or NetCDF-4) file.
|
||||
#[wasm_bindgen]
|
||||
pub struct H5File {
|
||||
@@ -145,38 +258,13 @@ impl H5File {
|
||||
|
||||
/// The group's members: `[{ name, kind }]`, groups first.
|
||||
pub fn list(&self, path: &str) -> Result<Array, JsError> {
|
||||
Ok(self
|
||||
.inner
|
||||
.list(path)
|
||||
.map_err(js_err)?
|
||||
.into_iter()
|
||||
.map(|c| {
|
||||
let o = Object::new();
|
||||
set(&o, "name", c.name);
|
||||
set(&o, "kind", c.kind.as_str());
|
||||
JsValue::from(o)
|
||||
})
|
||||
.collect())
|
||||
Ok(list_to_js(self.inner.list(path).map_err(js_err)?))
|
||||
}
|
||||
|
||||
/// `{ shape, maxshape, dtype, elementShape }`. `maxshape` is `null`
|
||||
/// when not recorded, with `null` for each unlimited dimension.
|
||||
pub fn info(&self, path: &str) -> Result<Object, JsError> {
|
||||
let i = self.inner.info(path).map_err(js_err)?;
|
||||
let o = Object::new();
|
||||
set(&o, "shape", shape_to_js(&i.shape));
|
||||
let max: JsValue = match i.maxshape {
|
||||
None => JsValue::NULL,
|
||||
Some(dims) => dims
|
||||
.into_iter()
|
||||
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
|
||||
.collect::<Array>()
|
||||
.into(),
|
||||
};
|
||||
set(&o, "maxshape", max);
|
||||
set(&o, "dtype", i.dtype);
|
||||
set(&o, "elementShape", shape_to_js(&i.element_shape));
|
||||
Ok(o)
|
||||
Ok(info_to_js(self.inner.info(path).map_err(js_err)?))
|
||||
}
|
||||
|
||||
/// `[{ name, value, dtype }]`, sorted by name. Scalars are `number`
|
||||
@@ -185,31 +273,21 @@ impl H5File {
|
||||
/// and its `dtype`; one that could not be read at all is reported by
|
||||
/// [`attrErrors`](Self::attr_errors).
|
||||
pub fn attrs(&self, path: &str) -> Result<Array, JsError> {
|
||||
let (attrs, _) = self.inner.attrs(path).map_err(js_err)?;
|
||||
Ok(attrs
|
||||
.into_iter()
|
||||
.map(|a| {
|
||||
let o = Object::new();
|
||||
set(&o, "name", a.name);
|
||||
let (value, dtype) = attr_to_js(a.value);
|
||||
set(&o, "value", value);
|
||||
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
|
||||
JsValue::from(o)
|
||||
})
|
||||
.collect())
|
||||
Ok(attrs_to_js(self.inner.attrs(path).map_err(js_err)?.0))
|
||||
}
|
||||
|
||||
/// Messages for attributes that could not be read.
|
||||
#[wasm_bindgen(js_name = attrErrors)]
|
||||
pub fn attr_errors(&self, path: &str) -> Result<Array, JsError> {
|
||||
let (_, errors) = self.inner.attrs(path).map_err(js_err)?;
|
||||
Ok(errors.into_iter().map(JsValue::from).collect())
|
||||
Ok(errors_to_js(self.inner.attrs(path).map_err(js_err)?.1))
|
||||
}
|
||||
|
||||
/// The whole dataset: `{ shape, dtype, data }`, `data` in row-major
|
||||
/// order.
|
||||
pub fn read(&self, path: &str) -> Result<Object, JsError> {
|
||||
self.read_impl(path, None)
|
||||
Ok(values_to_js(
|
||||
read_values(&self.inner, path, None).map_err(js_err)?,
|
||||
))
|
||||
}
|
||||
|
||||
/// A regular hyperslab (`H5Sselect_hyperslab`): `stride` and `block`
|
||||
@@ -223,22 +301,358 @@ impl H5File {
|
||||
stride: Option<Vec<f64>>,
|
||||
block: Option<Vec<f64>>,
|
||||
) -> Result<Object, JsError> {
|
||||
let slab = Hyperslab {
|
||||
start: indices_from_js("start", &start)?,
|
||||
count: indices_from_js("count", &count)?,
|
||||
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
|
||||
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
|
||||
};
|
||||
self.read_impl(path, Some(&slab))
|
||||
}
|
||||
|
||||
fn read_impl(&self, path: &str, slab: Option<&Hyperslab>) -> Result<Object, JsError> {
|
||||
let dtype = self.inner.info(path).map_err(js_err)?.dtype;
|
||||
let a = self.inner.read(path, slab).map_err(js_err)?;
|
||||
let slab = slab_from_js(&start, &count, stride, block)?;
|
||||
Ok(values_to_js(
|
||||
read_values(&self.inner, path, Some(&slab)).map_err(js_err)?,
|
||||
))
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Remote files: openUrl.
|
||||
|
||||
#[wasm_bindgen(module = "/js/remote.js")]
|
||||
extern "C" {
|
||||
#[wasm_bindgen(catch)]
|
||||
async fn probe(url: &str, first_len: f64, opts: &JsValue) -> Result<JsValue, JsValue>;
|
||||
|
||||
#[wasm_bindgen(catch, js_name = fetchRanges)]
|
||||
async fn fetch_ranges(
|
||||
url: &str,
|
||||
ranges: Vec<f64>,
|
||||
opts: &JsValue,
|
||||
validator: &JsValue,
|
||||
length: f64,
|
||||
) -> Result<JsValue, JsValue>;
|
||||
}
|
||||
|
||||
/// Where a remote file's bytes come from.
|
||||
struct Http {
|
||||
url: String,
|
||||
opts: JsValue,
|
||||
/// ETag or Last-Modified at open (`null` if the server sent neither).
|
||||
validator: JsValue,
|
||||
length: u64,
|
||||
/// Requests the probe made (1, or 2 with a HEAD for the length).
|
||||
probe_requests: u64,
|
||||
}
|
||||
|
||||
enum Source {
|
||||
/// Read by range requests through a restartable cache.
|
||||
Lazy {
|
||||
http: Http,
|
||||
storage: Arc<LazyStorage>,
|
||||
reader: Reader,
|
||||
},
|
||||
/// The server ignored `Range`: the whole file, downloaded at open.
|
||||
Whole {
|
||||
reader: Reader,
|
||||
size: u64,
|
||||
requests: u64,
|
||||
},
|
||||
}
|
||||
|
||||
impl Http {
|
||||
/// Fetch `ranges` and hand them to `storage`.
|
||||
async fn fetch(&self, storage: &LazyStorage, ranges: &[Range<u64>]) -> Result<(), JsError> {
|
||||
let flat: Vec<f64> = ranges
|
||||
.iter()
|
||||
.flat_map(|r| [r.start as f64, r.end as f64])
|
||||
.collect();
|
||||
let got = fetch_ranges(
|
||||
&self.url,
|
||||
flat,
|
||||
&self.opts,
|
||||
&self.validator,
|
||||
self.length as f64,
|
||||
)
|
||||
.await
|
||||
.map_err(js_exception)?;
|
||||
let got = Array::from(&got);
|
||||
if got.length() as usize != ranges.len() {
|
||||
return Err(js_err(format!(
|
||||
"fetchRanges returned {} ranges for {}",
|
||||
got.length(),
|
||||
ranges.len()
|
||||
)));
|
||||
}
|
||||
for (r, bytes) in ranges.iter().zip(got.iter()) {
|
||||
let bytes = Uint8Array::new(&bytes).to_vec();
|
||||
storage.supply_range(r, &bytes).map_err(js_err)?;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Run `f` over `storage` until it has every byte it reads.
|
||||
async fn drive<T>(
|
||||
&self,
|
||||
storage: &LazyStorage,
|
||||
mut f: impl FnMut() -> T,
|
||||
) -> Result<T, JsError> {
|
||||
let op = storage.operation();
|
||||
loop {
|
||||
match storage.attempt(&mut f) {
|
||||
Step::Done(v) => return Ok(v),
|
||||
Step::Need(ranges) => {
|
||||
op.charge(&ranges).map_err(js_err)?;
|
||||
self.fetch(storage, &ranges).await?
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl Source {
|
||||
async fn run<T>(&self, op: impl Fn(&Reader) -> core::Result<T>) -> Result<T, JsError> {
|
||||
match self {
|
||||
Source::Whole { reader, .. } => op(reader).map_err(js_err),
|
||||
Source::Lazy {
|
||||
http,
|
||||
storage,
|
||||
reader,
|
||||
} => http.drive(storage, || op(reader)).await?.map_err(js_err),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A non-negative integer option, or `None` when not given.
|
||||
fn int_opt(opts: &JsValue, key: &str, min: f64, max: f64) -> Result<Option<u64>, JsError> {
|
||||
let v = get(opts, key);
|
||||
if v.is_undefined() || v.is_null() {
|
||||
return Ok(None);
|
||||
}
|
||||
match v.as_f64() {
|
||||
Some(x) if x.fract() == 0.0 && (min..=max).contains(&x) => Ok(Some(x as u64)),
|
||||
_ => Err(js_err(format!(
|
||||
"openUrl: {key} must be an integer from {min} to {max}"
|
||||
))),
|
||||
}
|
||||
}
|
||||
|
||||
/// Most bytes `maxFetch` and `maxDownload` may allow: 1 GiB. wasm32 has
|
||||
/// 4 GiB of memory and no buffer past 2 GiB, and what is fetched is held
|
||||
/// while it is decoded.
|
||||
const MAX_FETCH_LIMIT: u64 = 1 << 30;
|
||||
|
||||
/// The largest file `openUrl` reads by ranges: on wasm32, 4 GiB - 1 bytes.
|
||||
/// The format code turns file offsets into `usize` to use them (with a
|
||||
/// clean error past it, see scripts/check-32bit-casts.sh), so on a 32-bit
|
||||
/// target nothing at 4 GiB or beyond can be read; a larger file is refused
|
||||
/// at open rather than failing on whichever read reaches past 4 GiB. On
|
||||
/// 64-bit targets it is 2^53 - 1, the largest offset a JavaScript number
|
||||
/// holds exactly.
|
||||
const MAX_REMOTE_LENGTH: u64 = if (usize::MAX as u64) < MAX_SAFE_INTEGER as u64 {
|
||||
usize::MAX as u64
|
||||
} else {
|
||||
MAX_SAFE_INTEGER as u64
|
||||
};
|
||||
|
||||
fn config_from(opts: &JsValue) -> Result<LazyConfig, JsError> {
|
||||
let mut c = LazyConfig::default();
|
||||
if let Some(b) = int_opt(opts, "blockSize", 512.0, (64u64 << 20) as f64)? {
|
||||
c.block_size = b;
|
||||
}
|
||||
if let Some(n) = int_opt(opts, "cacheSize", 0.0, MAX_SAFE_INTEGER)? {
|
||||
c.capacity = n;
|
||||
}
|
||||
if let Some(n) = int_opt(opts, "maxFetch", 512.0, MAX_FETCH_LIMIT as f64)? {
|
||||
c.max_fetch = n;
|
||||
}
|
||||
// Read by remote.js; checked here so a value wasm32 cannot hold is an
|
||||
// option error rather than a download that cannot be kept.
|
||||
int_opt(opts, "maxDownload", 0.0, MAX_FETCH_LIMIT as f64)?;
|
||||
// Also read by remote.js (which checks it too, for direct callers).
|
||||
int_opt(opts, "parallel", 1.0, 1024.0)?;
|
||||
Ok(c)
|
||||
}
|
||||
|
||||
/// Open the HDF5 file at `url` without downloading it: its bytes are
|
||||
/// fetched with HTTP `Range` requests as the methods of the returned
|
||||
/// [`RemoteFile`] need them, through a block cache.
|
||||
///
|
||||
/// `opts` (all optional):
|
||||
/// - `blockSize` — bytes per request block, 512 to 64 MiB (default 1 MiB);
|
||||
/// - `cacheSize` — bytes of blocks kept between calls (default 64 MiB);
|
||||
/// - `maxFetch` — most bytes one call may fetch, and so the longest single
|
||||
/// read, up to 1 GiB (default 512 MiB): a call that would fetch more
|
||||
/// fails before fetching it;
|
||||
/// - `fallback` — `"download"` (default) reads the whole file when the
|
||||
/// server ignores `Range` (answers 200), up to `maxDownload` bytes
|
||||
/// (default 512 MiB, at most 1 GiB); `"error"` refuses such a server;
|
||||
/// - `headers`, `credentials` — passed to every `fetch` (`headers` as
|
||||
/// `fetch` takes them: a `Headers`, `[name, value]` pairs or an object);
|
||||
/// - `parallel` — range requests in flight at once, 1 to 1024 (default
|
||||
/// 6); when one fails the others are aborted;
|
||||
/// - `fetch` — a `fetch`-compatible function to use instead of the global.
|
||||
///
|
||||
/// Cross-origin servers must allow CORS and expose `Content-Range` (or
|
||||
/// answer `HEAD` with `Content-Length`). A file may be up to 4 GiB - 1
|
||||
/// bytes long (wasm32 offsets); a longer one is refused at open. A whole-dataset `read` that would use more than 1 GiB of
|
||||
/// memory ([`core::MAX_READ_BYTES`]) is refused: read it in parts with
|
||||
/// `readHyperslab`.
|
||||
#[wasm_bindgen(js_name = openUrl)]
|
||||
pub async fn open_url(url: String, opts: JsValue) -> Result<RemoteFile, JsError> {
|
||||
let config = config_from(&opts)?;
|
||||
let p = probe(&url, config.block_size as f64, &opts)
|
||||
.await
|
||||
.map_err(js_exception)?;
|
||||
let requests = get(&p, "requests").as_f64().unwrap_or(1.0) as u64;
|
||||
let whole = get(&p, "whole");
|
||||
if !whole.is_undefined() {
|
||||
let bytes = Uint8Array::new(&whole).to_vec();
|
||||
let size = bytes.len() as u64;
|
||||
let reader = Reader::open(bytes).map_err(js_err)?;
|
||||
return Ok(RemoteFile {
|
||||
inner: Rc::new(Source::Whole {
|
||||
reader,
|
||||
size,
|
||||
requests,
|
||||
}),
|
||||
});
|
||||
}
|
||||
let length = get(&p, "length")
|
||||
.as_f64()
|
||||
.filter(|x| x.fract() == 0.0 && (0.0..=MAX_SAFE_INTEGER).contains(x))
|
||||
.ok_or_else(|| js_err(format!("{url}: the server gave no usable file size")))?;
|
||||
if length as u64 > MAX_REMOTE_LENGTH {
|
||||
return Err(js_err(format!(
|
||||
"{url} is {length} bytes; openUrl reads files of up to {MAX_REMOTE_LENGTH} bytes \
|
||||
(4 GiB - 1: the WebAssembly reader addresses a file with 32-bit offsets)"
|
||||
)));
|
||||
}
|
||||
let http = Http {
|
||||
url,
|
||||
opts,
|
||||
validator: get(&p, "validator"),
|
||||
length: length as u64,
|
||||
probe_requests: requests,
|
||||
};
|
||||
let storage = Arc::new(LazyStorage::new(http.length, config));
|
||||
let first = Uint8Array::new(&get(&p, "first")).to_vec();
|
||||
storage.supply(0, &first).map_err(js_err)?;
|
||||
let s = storage.clone();
|
||||
let reader = http
|
||||
.drive(&storage, || Reader::open_storage(s.clone()))
|
||||
.await?
|
||||
.map_err(js_err)?;
|
||||
Ok(RemoteFile {
|
||||
inner: Rc::new(Source::Lazy {
|
||||
http,
|
||||
storage,
|
||||
reader,
|
||||
}),
|
||||
})
|
||||
}
|
||||
|
||||
/// A file opened with [`openUrl`](open_url): the methods of [`H5File`],
|
||||
/// each returning a `Promise` (it may have to fetch bytes first).
|
||||
#[wasm_bindgen]
|
||||
pub struct RemoteFile {
|
||||
inner: Rc<Source>,
|
||||
}
|
||||
|
||||
impl RemoteFile {
|
||||
/// Run `op` (fetching what it needs) and convert its result.
|
||||
fn call<T: 'static>(
|
||||
&self,
|
||||
op: impl Fn(&Reader) -> core::Result<T> + 'static,
|
||||
to_js: impl FnOnce(T) -> JsValue + 'static,
|
||||
) -> Promise {
|
||||
let inner = self.inner.clone();
|
||||
future_to_promise(async move {
|
||||
let v = inner.run(op).await.map_err(JsValue::from)?;
|
||||
Ok(to_js(v))
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
#[wasm_bindgen]
|
||||
impl RemoteFile {
|
||||
/// `"group"` or `"dataset"`.
|
||||
#[wasm_bindgen(unchecked_return_type = "Promise<string>")]
|
||||
pub fn kind(&self, path: String) -> Promise {
|
||||
self.call(move |r| r.kind(&path), |k| k.as_str().into())
|
||||
}
|
||||
|
||||
/// The group's members: `[{ name, kind }]`, groups first.
|
||||
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
|
||||
pub fn list(&self, path: String) -> Promise {
|
||||
self.call(move |r| r.list(&path), |c| list_to_js(c).into())
|
||||
}
|
||||
|
||||
/// `{ shape, maxshape, dtype, elementShape }`, as [`H5File::info`].
|
||||
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
|
||||
pub fn info(&self, path: String) -> Promise {
|
||||
self.call(move |r| r.info(&path), |i| info_to_js(i).into())
|
||||
}
|
||||
|
||||
/// `[{ name, value, dtype }]`, as [`H5File::attrs`].
|
||||
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
|
||||
pub fn attrs(&self, path: String) -> Promise {
|
||||
self.call(move |r| r.attrs(&path), |a| attrs_to_js(a.0).into())
|
||||
}
|
||||
|
||||
/// Messages for attributes that could not be read.
|
||||
#[wasm_bindgen(js_name = attrErrors, unchecked_return_type = "Promise<Array<string>>")]
|
||||
pub fn attr_errors(&self, path: String) -> Promise {
|
||||
self.call(move |r| r.attrs(&path), |a| errors_to_js(a.1).into())
|
||||
}
|
||||
|
||||
/// The whole dataset: `{ shape, dtype, data }`, as [`H5File::read`].
|
||||
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
|
||||
pub fn read(&self, path: String) -> Promise {
|
||||
self.call(
|
||||
move |r| read_values(r, &path, None),
|
||||
|v| values_to_js(v).into(),
|
||||
)
|
||||
}
|
||||
|
||||
/// A regular hyperslab, as [`H5File::read_hyperslab`]. Only the chunks
|
||||
/// (or the contiguous runs) the selection touches are fetched.
|
||||
#[wasm_bindgen(js_name = readHyperslab, unchecked_return_type = "Promise<any>")]
|
||||
pub fn read_hyperslab(
|
||||
&self,
|
||||
path: String,
|
||||
start: Vec<f64>,
|
||||
count: Vec<f64>,
|
||||
stride: Option<Vec<f64>>,
|
||||
block: Option<Vec<f64>>,
|
||||
) -> Result<Promise, JsError> {
|
||||
let slab = slab_from_js(&start, &count, stride, block)?;
|
||||
Ok(self.call(
|
||||
move |r| read_values(r, &path, Some(&slab)),
|
||||
|v| values_to_js(v).into(),
|
||||
))
|
||||
}
|
||||
|
||||
/// What reading this file has cost so far: `{ lazy, size, requests,
|
||||
/// bytesFetched, cachedBytes, passes }`. `lazy` is false when the
|
||||
/// server ignored `Range` and the file was downloaded whole.
|
||||
pub fn stats(&self) -> Object {
|
||||
let o = Object::new();
|
||||
set(&o, "shape", shape_to_js(&a.shape));
|
||||
set(&o, "dtype", dtype);
|
||||
set(&o, "data", data_to_js(a.data));
|
||||
Ok(o)
|
||||
match &*self.inner {
|
||||
Source::Lazy { http, storage, .. } => {
|
||||
let st = storage.stats();
|
||||
set(&o, "lazy", true);
|
||||
set(&o, "size", http.length as f64);
|
||||
set(
|
||||
&o,
|
||||
"requests",
|
||||
(st.requests.saturating_sub(1) + http.probe_requests) as f64,
|
||||
);
|
||||
set(&o, "bytesFetched", st.bytes_fetched as f64);
|
||||
set(&o, "cachedBytes", st.cached_bytes as f64);
|
||||
set(&o, "passes", st.passes as f64);
|
||||
}
|
||||
Source::Whole { size, requests, .. } => {
|
||||
set(&o, "lazy", false);
|
||||
set(&o, "size", *size as f64);
|
||||
set(&o, "requests", *requests as f64);
|
||||
set(&o, "bytesFetched", *size as f64);
|
||||
set(&o, "cachedBytes", *size as f64);
|
||||
set(&o, "passes", 0.0);
|
||||
}
|
||||
}
|
||||
o
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,578 @@
|
||||
//! The restartable ("NeedBytes") reader against the in-memory one: every
|
||||
//! file must list, describe and read the same through a [`LazyStorage`]
|
||||
//! that starts empty and is fed only the ranges its passes ask for, as the
|
||||
//! browser's `openUrl` feeds it from HTTP range requests.
|
||||
//!
|
||||
//! - Files written here with `FileBuilder`, at several block sizes (512 B
|
||||
//! blocks make almost every structure read a miss).
|
||||
//! - The h5py/netCDF4 fixture of `examples/wasm-viewer/test/make_fixture.py`
|
||||
//! (skipped without h5py, unless `CLAWHDF5_REQUIRE_INTEROP=1`;
|
||||
//! `CLAWHDF5_PYTHON` names the interpreter).
|
||||
//! - `CLAWHDF5_WASM_CORPUS=dir[:dir...]`: every HDF5 file under those
|
||||
//! directories up to 64 MiB (e.g. `conformance/.cache/corpus`).
|
||||
//!
|
||||
//! Also the request budget: listing and reading one small dataset of a large
|
||||
//! file fetches a few blocks, not the file.
|
||||
|
||||
use std::ops::Range;
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::process::Command;
|
||||
use std::sync::Arc;
|
||||
|
||||
use clawhdf5::{AttrValue, FileBuilder};
|
||||
use clawhdf5_format::storage::CountingStorage;
|
||||
use clawhdf5_wasm::core::{Hyperslab, Kind, Reader};
|
||||
use clawhdf5_wasm::lazy::{LazyConfig, LazyStorage};
|
||||
|
||||
/// The API the JavaScript side calls, one operation at a time.
|
||||
trait Api {
|
||||
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T;
|
||||
}
|
||||
|
||||
struct Local(Reader);
|
||||
|
||||
impl Api for Local {
|
||||
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
|
||||
op(&self.0)
|
||||
}
|
||||
}
|
||||
|
||||
/// A lazily read file and the "server" it fetches from.
|
||||
struct Lazy {
|
||||
data: Arc<Vec<u8>>,
|
||||
storage: Arc<LazyStorage>,
|
||||
reader: Reader,
|
||||
}
|
||||
|
||||
fn fetch(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
|
||||
Ok(data[r.start as usize..r.end as usize].to_vec())
|
||||
}
|
||||
|
||||
impl Lazy {
|
||||
/// Open as `openUrl` does: the first block comes with the probe that
|
||||
/// learns the length, then the open is run until it has its bytes.
|
||||
fn open(data: Vec<u8>, config: LazyConfig) -> Result<Lazy, String> {
|
||||
let data = Arc::new(data);
|
||||
let storage = Arc::new(LazyStorage::new(data.len() as u64, config));
|
||||
let first = (storage.config().block_size as usize).min(data.len());
|
||||
storage.supply(0, &data[..first])?;
|
||||
let s = storage.clone();
|
||||
let reader =
|
||||
storage.run_blocking(|| Reader::open_storage(s.clone()), |r| fetch(&data, r))??;
|
||||
Ok(Lazy {
|
||||
data,
|
||||
storage,
|
||||
reader,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
impl Api for Lazy {
|
||||
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
|
||||
self.storage
|
||||
.run_blocking(|| op(&self.reader), |r| fetch(&self.data, r))
|
||||
.expect("serving from memory cannot fail")
|
||||
}
|
||||
}
|
||||
|
||||
/// Everything the viewer can show of a file, as text: each object's kind,
|
||||
/// listing, attributes (and attribute errors), dataset info, whole value
|
||||
/// and a hyperslab — or the error each gives.
|
||||
fn transcript(api: &impl Api) -> Vec<String> {
|
||||
let mut out = Vec::new();
|
||||
let mut todo = vec![("/".to_string(), 0usize)];
|
||||
while let Some((path, depth)) = todo.pop() {
|
||||
if out.len() > 4000 {
|
||||
out.push("... (truncated)".into());
|
||||
break;
|
||||
}
|
||||
let kind = api.call(|r| r.kind(&path));
|
||||
out.push(format!("{path}: {kind:?}"));
|
||||
out.push(format!("{path} attrs: {:?}", api.call(|r| r.attrs(&path))));
|
||||
match kind {
|
||||
Ok(Kind::Group) => {
|
||||
let list = api.call(|r| r.list(&path));
|
||||
out.push(format!("{path} list: {list:?}"));
|
||||
if let Ok(children) = list
|
||||
&& depth < 12
|
||||
{
|
||||
for c in children.into_iter().rev() {
|
||||
let child = if path == "/" {
|
||||
format!("/{}", c.name)
|
||||
} else {
|
||||
format!("{path}/{}", c.name)
|
||||
};
|
||||
todo.push((child, depth + 1));
|
||||
}
|
||||
}
|
||||
}
|
||||
Ok(Kind::Dataset) => {
|
||||
let info = api.call(|r| r.info(&path));
|
||||
out.push(format!("{path} info: {info:?}"));
|
||||
let Ok(info) = info else { continue };
|
||||
let n = info
|
||||
.shape
|
||||
.iter()
|
||||
.chain(&info.element_shape)
|
||||
.try_fold(1u64, |a, &d| a.checked_mul(d));
|
||||
if n.is_none_or(|n| n > 4_000_000) {
|
||||
out.push(format!("{path}: not read ({n:?} values)"));
|
||||
continue;
|
||||
}
|
||||
out.push(format!(
|
||||
"{path} read: {:?}",
|
||||
api.call(|r| r.read(&path, None))
|
||||
));
|
||||
if !info.shape.is_empty() && info.shape.iter().all(|&d| d > 1) {
|
||||
let slab = Hyperslab {
|
||||
start: info.shape.iter().map(|_| 1).collect(),
|
||||
count: info.shape.iter().map(|&d| d / 2).collect(),
|
||||
stride: None,
|
||||
block: None,
|
||||
};
|
||||
let part = api.call(|r| r.read(&path, Some(&slab)));
|
||||
out.push(format!("{path} slab: {part:?}"));
|
||||
}
|
||||
}
|
||||
Err(_) => {}
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// The lazy transcript of `data` at `block` bytes per block equals the
|
||||
/// transcript of the same file through a range storage that has every byte
|
||||
/// (`CountingStorage`: the facade's `Storage` path, the one the lazy reader
|
||||
/// takes), and agrees with the in-memory one: the same values, and an error
|
||||
/// wherever it has one (a malformed file can fail at a different check,
|
||||
/// with a different message, when read by ranges). Returns what the lazy
|
||||
/// reader fetched and its transcript.
|
||||
fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec<String>) {
|
||||
let ctx = format!("{name} (blocks of {block} B)");
|
||||
let ranged = Reader::open_storage(Arc::new(CountingStorage::new(data.to_vec())));
|
||||
let local = Reader::open(data.to_vec());
|
||||
let lazy = Lazy::open(data.to_vec(), config(block));
|
||||
let (ranged, local, lazy) = match (ranged, local, lazy) {
|
||||
(Ok(r), Ok(l), Ok(z)) => (r, l, z),
|
||||
(Err(r), Err(_), Err(z)) => {
|
||||
assert_eq!(z, r, "{ctx}: open error");
|
||||
return (0, 0, Vec::new());
|
||||
}
|
||||
(r, l, z) => panic!(
|
||||
"{ctx}: opens differently: ranged {:?}, in memory {:?}, lazily {:?}",
|
||||
r.err(),
|
||||
l.err(),
|
||||
z.err()
|
||||
),
|
||||
};
|
||||
let got = transcript(&lazy);
|
||||
let want = transcript(&Local(ranged));
|
||||
for (i, (w, g)) in want.iter().zip(&got).enumerate() {
|
||||
assert_eq!(g, w, "{ctx}, line {i}");
|
||||
}
|
||||
assert_eq!(got.len(), want.len(), "{ctx}: transcript length");
|
||||
let local = transcript(&Local(local));
|
||||
for (i, (l, g)) in local.iter().zip(&got).enumerate() {
|
||||
let both_errors = match (l.split_once("Err("), g.split_once("Err(")) {
|
||||
(Some((a, _)), Some((b, _))) => a == b,
|
||||
_ => false,
|
||||
};
|
||||
assert!(
|
||||
l == g || both_errors,
|
||||
"{ctx}, line {i}: in memory\n {l}\nlazily\n {g}"
|
||||
);
|
||||
}
|
||||
assert_eq!(got.len(), local.len(), "{ctx}: transcript length");
|
||||
let st = lazy.storage.stats();
|
||||
(st.requests, st.bytes_fetched, got)
|
||||
}
|
||||
|
||||
fn config(block: u64) -> LazyConfig {
|
||||
LazyConfig {
|
||||
block_size: block,
|
||||
// A small budget, so eviction between operations is exercised.
|
||||
capacity: 16 * block,
|
||||
max_request: 8 * block,
|
||||
..LazyConfig::default()
|
||||
}
|
||||
}
|
||||
|
||||
fn builder_file() -> Vec<u8> {
|
||||
let mut b = FileBuilder::new();
|
||||
b.create_dataset("grid")
|
||||
.with_f64_data(&(0..20_000).map(f64::from).collect::<Vec<_>>())
|
||||
.with_shape(&[100, 200])
|
||||
.with_chunks(&[10, 25])
|
||||
.with_deflate(4);
|
||||
b.create_dataset("contiguous")
|
||||
.with_i32_data(&(0..50_000).collect::<Vec<_>>());
|
||||
b.create_dataset("bytes").with_u8_data(&[1, 2, 250]);
|
||||
let mut g = b.create_group("sensors");
|
||||
for i in 0..40 {
|
||||
g.create_dataset(&format!("t{i}"))
|
||||
.with_f32_data(&[i as f32, 1.5, -2.25]);
|
||||
}
|
||||
g.set_attr("location", AttrValue::String("lab".into()));
|
||||
b.add_group(g.finish());
|
||||
b.set_attr("version", AttrValue::I64(3));
|
||||
b.set_attr("scale", AttrValue::F64Array(vec![0.5, 2.0]));
|
||||
b.finish().unwrap()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn builder_files_read_the_same_at_every_block_size() {
|
||||
let data = builder_file();
|
||||
for block in [512, 4096, 1 << 20] {
|
||||
let (requests, _, lines) = check_equal("builder", &data, block);
|
||||
assert!(requests > 0);
|
||||
// The transcript covers every object, values included.
|
||||
assert!(lines.iter().any(|l| l.starts_with("/grid read: Ok")));
|
||||
assert!(lines.iter().any(|l| l.starts_with("/grid slab: Ok")));
|
||||
assert!(lines.iter().any(|l| l.starts_with("/sensors/t39 read: Ok")));
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn garbage_fails_to_open_as_in_memory() {
|
||||
check_equal("zeros", &[0u8; 5000], 512);
|
||||
check_equal("empty", &[], 512);
|
||||
let mut cut = builder_file();
|
||||
cut.truncate(cut.len() / 3);
|
||||
check_equal("truncated", &cut, 512);
|
||||
}
|
||||
|
||||
/// Listing a large file and reading one small dataset fetches a few blocks,
|
||||
/// not the file.
|
||||
#[test]
|
||||
fn a_small_read_of_a_large_file_fetches_a_few_blocks() {
|
||||
let mut b = FileBuilder::new();
|
||||
b.create_dataset("small").with_f64_data(&[1.0, 2.0, 3.0]);
|
||||
// 48 MB of raw data, written after the small dataset's metadata.
|
||||
b.create_dataset("big")
|
||||
.with_f64_data(&(0..6_000_000).map(f64::from).collect::<Vec<_>>());
|
||||
let mut g = b.create_group("group");
|
||||
g.create_dataset("inner").with_i32_data(&[7, 8]);
|
||||
b.add_group(g.finish());
|
||||
let data = b.finish().unwrap();
|
||||
let lazy = Lazy::open(data.clone(), LazyConfig::default()).unwrap();
|
||||
let list = lazy.call(|r| r.list("/")).unwrap();
|
||||
assert_eq!(list.len(), 3);
|
||||
assert_eq!(
|
||||
format!("{:?}", lazy.call(|r| r.read("/small", None)).unwrap().data),
|
||||
"F64([1.0, 2.0, 3.0])"
|
||||
);
|
||||
assert_eq!(
|
||||
format!(
|
||||
"{:?}",
|
||||
lazy.call(|r| r.read("/group/inner", None)).unwrap().data
|
||||
),
|
||||
"I32([7, 8])"
|
||||
);
|
||||
// A window of the big dataset reads only its block(s).
|
||||
let slab = Hyperslab {
|
||||
start: vec![3_000_000],
|
||||
count: vec![4],
|
||||
stride: None,
|
||||
block: None,
|
||||
};
|
||||
assert_eq!(
|
||||
format!(
|
||||
"{:?}",
|
||||
lazy.call(|r| r.read("/big", Some(&slab))).unwrap().data
|
||||
),
|
||||
"F64([3000000.0, 3000001.0, 3000002.0, 3000003.0])"
|
||||
);
|
||||
let st = lazy.storage.stats();
|
||||
eprintln!("{} bytes: {st:?}", data.len());
|
||||
assert!(st.requests <= 6, "{st:?}");
|
||||
assert!(st.bytes_fetched <= 6 << 20, "{st:?}");
|
||||
assert!(st.bytes_fetched * 8 < data.len() as u64, "{st:?}");
|
||||
}
|
||||
|
||||
fn python() -> String {
|
||||
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||
}
|
||||
|
||||
fn python_available() -> bool {
|
||||
Command::new(python())
|
||||
.args(["-c", "import h5py, netCDF4, numpy"])
|
||||
.output()
|
||||
.is_ok_and(|o| o.status.success())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn h5py_and_netcdf4_files_read_the_same_lazily() {
|
||||
if !python_available() {
|
||||
assert!(
|
||||
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
|
||||
python()
|
||||
);
|
||||
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
|
||||
return;
|
||||
}
|
||||
let dir = fixture_dir();
|
||||
for name in ["fixture.h5", "fixture.nc"] {
|
||||
let data = std::fs::read(dir.path().join(name)).unwrap();
|
||||
for block in [512, 64 * 1024] {
|
||||
let (_, _, lines) = check_equal(name, &data, block);
|
||||
assert!(lines.iter().filter(|l| l.contains(" read: Ok")).count() >= 2);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The bytes of `data` at `r`, zero past its end: a server that claims
|
||||
/// the file is longer than it is.
|
||||
fn fetch_padded(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
|
||||
let mut out = vec![0u8; (r.end - r.start) as usize];
|
||||
let len = data.len() as u64;
|
||||
if r.start < len {
|
||||
let end = r.end.min(len);
|
||||
out[..(end - r.start) as usize].copy_from_slice(&data[r.start as usize..end as usize]);
|
||||
}
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
/// Sizes a hostile server or a large dataset can name are errors, never
|
||||
/// allocations that abort the wasm module: make_fixture.py's limits.h5 and
|
||||
/// hostile_vl.h5 (see write_limits there).
|
||||
#[test]
|
||||
fn size_limits_are_errors_not_aborts() {
|
||||
if !python_available() {
|
||||
assert!(
|
||||
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
|
||||
python()
|
||||
);
|
||||
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
|
||||
return;
|
||||
}
|
||||
let dir = fixture_dir();
|
||||
|
||||
// Read whole, /huge_u8 would widen 2^28 values to 64 bits (2 GiB): an
|
||||
// error naming readHyperslab, before its chunks are read. A window of
|
||||
// it reads.
|
||||
let data = std::fs::read(dir.path().join("limits.h5")).unwrap();
|
||||
let n = (1u64 << 28) + 1024;
|
||||
let window = Hyperslab {
|
||||
start: vec![n - 4],
|
||||
count: vec![4],
|
||||
stride: None,
|
||||
block: None,
|
||||
};
|
||||
let local = Reader::open(data.clone()).unwrap();
|
||||
let lazy = Lazy::open(data, LazyConfig::default()).unwrap();
|
||||
let before = lazy.storage.stats().requests;
|
||||
for e in [
|
||||
local.read("/huge_u8", None).unwrap_err(),
|
||||
lazy.call(|r| r.read("/huge_u8", None)).unwrap_err(),
|
||||
] {
|
||||
assert!(e.contains("readHyperslab"), "{e}");
|
||||
}
|
||||
assert_eq!(lazy.storage.stats().requests, before, "nothing fetched");
|
||||
for part in [
|
||||
local.read("/huge_u8", Some(&window)).unwrap(),
|
||||
lazy.call(|r| r.read("/huge_u8", Some(&window))).unwrap(),
|
||||
] {
|
||||
assert_eq!(format!("{:?}", part.data), "U8([0, 0, 0, 7])");
|
||||
}
|
||||
|
||||
// A server that claims 3 GiB and a heap collection of 2 GiB + 4 KiB:
|
||||
// reading the strings fails at once, fetching a few blocks.
|
||||
let data = std::fs::read(dir.path().join("hostile_vl.h5")).unwrap();
|
||||
let storage = Arc::new(LazyStorage::new(3 << 30, LazyConfig::default()));
|
||||
let s = storage.clone();
|
||||
let reader = storage
|
||||
.run_blocking(
|
||||
|| Reader::open_storage(s.clone()),
|
||||
|r| fetch_padded(&data, r),
|
||||
)
|
||||
.unwrap()
|
||||
.unwrap();
|
||||
let e = storage
|
||||
.run_blocking(|| reader.read("/a", None), |r| fetch_padded(&data, r))
|
||||
.unwrap()
|
||||
.unwrap_err();
|
||||
assert!(e.contains("maxFetch"), "{e}");
|
||||
let st = storage.stats();
|
||||
assert!(st.requests <= 4 && st.bytes_fetched <= 4 << 20, "{st:?}");
|
||||
}
|
||||
|
||||
/// make_fixture.py's files, written to a temporary directory.
|
||||
fn fixture_dir() -> tempfile::TempDir {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let generator = Path::new(env!("CARGO_MANIFEST_DIR"))
|
||||
.join("../../examples/wasm-viewer/test/make_fixture.py");
|
||||
let out = Command::new(python())
|
||||
.arg(&generator)
|
||||
.arg(dir.path())
|
||||
.output()
|
||||
.unwrap();
|
||||
assert!(
|
||||
out.status.success(),
|
||||
"{}",
|
||||
String::from_utf8_lossy(&out.stderr)
|
||||
);
|
||||
dir
|
||||
}
|
||||
|
||||
fn hdf5_files(dir: &Path, out: &mut Vec<PathBuf>) {
|
||||
let Ok(entries) = std::fs::read_dir(dir) else {
|
||||
return;
|
||||
};
|
||||
for e in entries.flatten() {
|
||||
let p = e.path();
|
||||
if p.is_dir() {
|
||||
hdf5_files(&p, out);
|
||||
} else if std::fs::read(&p)
|
||||
.ok()
|
||||
.is_some_and(|b| b.len() <= 64 << 20 && is_hdf5(&b))
|
||||
{
|
||||
out.push(p);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The HDF5 signature at 0 or a power-of-two user-block offset.
|
||||
fn is_hdf5(b: &[u8]) -> bool {
|
||||
const SIG: &[u8] = b"\x89HDF\r\n\x1a\n";
|
||||
let mut at = 0usize;
|
||||
loop {
|
||||
if b.get(at..at + 8) == Some(SIG) {
|
||||
return true;
|
||||
}
|
||||
at = if at == 0 { 512 } else { at * 2 };
|
||||
if at >= b.len() {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn corpus_files_read_the_same_lazily() {
|
||||
let Ok(dirs) = std::env::var("CLAWHDF5_WASM_CORPUS") else {
|
||||
eprintln!("CLAWHDF5_WASM_CORPUS not set; skipping the corpus");
|
||||
return;
|
||||
};
|
||||
let mut files = Vec::new();
|
||||
for d in std::env::split_paths(&dirs) {
|
||||
hdf5_files(&d, &mut files);
|
||||
}
|
||||
files.sort();
|
||||
assert!(!files.is_empty(), "no HDF5 files under {dirs}");
|
||||
let (mut requests, mut bytes, mut total) = (0u64, 0u64, 0u64);
|
||||
for f in &files {
|
||||
let data = std::fs::read(f).unwrap();
|
||||
total += data.len() as u64;
|
||||
let (r, b, _) = check_equal(&f.display().to_string(), &data, 64 * 1024);
|
||||
requests += r;
|
||||
bytes += b;
|
||||
}
|
||||
eprintln!(
|
||||
"{} files ({total} bytes): {requests} requests, {bytes} bytes fetched",
|
||||
files.len()
|
||||
);
|
||||
}
|
||||
|
||||
/// Passes and requests `list(path)` takes on a file opened lazily at
|
||||
/// `block`-byte blocks (the open not counted), checking the listing against
|
||||
/// the in-memory one.
|
||||
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
|
||||
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
|
||||
let lazy = Lazy::open(
|
||||
data.to_vec(),
|
||||
LazyConfig {
|
||||
block_size: block,
|
||||
..LazyConfig::default()
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
let before = lazy.storage.stats();
|
||||
assert_eq!(lazy.call(|r| r.list(path)).unwrap(), want);
|
||||
let after = lazy.storage.stats();
|
||||
(
|
||||
after.passes - before.passes,
|
||||
after.requests - before.requests,
|
||||
)
|
||||
}
|
||||
|
||||
/// Listing a group reads every child's object header, and its index (B-tree
|
||||
/// and symbol table nodes, or B-tree v2 and heap blocks) before that. Each
|
||||
/// pass asks for every node of a level it is missing, not the first one
|
||||
/// only, so the passes (network round trips) grow with the depth of the
|
||||
/// index, not with the number of children: 2000 children with headers
|
||||
/// scattered over 512-byte blocks list in a handful of passes, where each
|
||||
/// header block used to cost its own.
|
||||
#[test]
|
||||
fn listing_a_large_group_takes_a_few_passes() {
|
||||
let mut b = FileBuilder::new();
|
||||
let mut g = b.create_group("many");
|
||||
for i in 0..600 {
|
||||
g.create_dataset(&format!("d{i}")).with_i32_data(&[i; 64]);
|
||||
}
|
||||
b.add_group(g.finish());
|
||||
let data = b.finish().unwrap();
|
||||
let (passes, requests) = listing_cost(&data, "/many", 512);
|
||||
eprintln!("FileBuilder, 600 children: {passes} passes, {requests} requests");
|
||||
assert!(passes <= 6, "{passes} passes");
|
||||
|
||||
if !python_available() {
|
||||
return;
|
||||
}
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
for libver in ["earliest", "latest"] {
|
||||
let path = dir.path().join(format!("{libver}.h5"));
|
||||
let script = format!(
|
||||
"import h5py, numpy as np\n\
|
||||
with h5py.File({:?}, 'w', libver='{libver}') as f:\n\
|
||||
\x20 for i in range(2000):\n\
|
||||
\x20 f.create_dataset('d%d' % i, data=np.full(256, i, np.float32))\n",
|
||||
path.display().to_string()
|
||||
);
|
||||
let out = Command::new(python())
|
||||
.args(["-c", &script])
|
||||
.output()
|
||||
.unwrap();
|
||||
assert!(
|
||||
out.status.success(),
|
||||
"{}",
|
||||
String::from_utf8_lossy(&out.stderr)
|
||||
);
|
||||
let data = std::fs::read(&path).unwrap();
|
||||
let (passes, requests) = listing_cost(&data, "/", 512);
|
||||
eprintln!("h5py libver={libver}, 2000 children: {passes} passes, {requests} requests");
|
||||
assert!(passes <= 12, "{libver}: {passes} passes");
|
||||
}
|
||||
}
|
||||
|
||||
/// `CLAWHDF5_WASM_LIST_FILE=file.h5`: what listing the root group of that
|
||||
/// file costs lazily, at 1 MiB and 64 KiB blocks (a measurement, printed).
|
||||
#[test]
|
||||
fn listing_cost_of_a_given_file() {
|
||||
let Ok(path) = std::env::var("CLAWHDF5_WASM_LIST_FILE") else {
|
||||
return;
|
||||
};
|
||||
let data = std::fs::read(&path).unwrap();
|
||||
for block in [1 << 20, 64 << 10] {
|
||||
let lazy = Lazy::open(
|
||||
data.clone(),
|
||||
LazyConfig {
|
||||
block_size: block,
|
||||
..LazyConfig::default()
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
let open = lazy.storage.stats();
|
||||
let n = lazy.call(|r| r.list("/")).unwrap().len();
|
||||
let st = lazy.storage.stats();
|
||||
eprintln!(
|
||||
"{path} ({} bytes), {block}-byte blocks: open {} requests / {} passes; list('/') of {n}: {} passes, {} requests, {} bytes",
|
||||
data.len(),
|
||||
open.requests,
|
||||
open.passes,
|
||||
st.passes - open.passes,
|
||||
st.requests - open.requests,
|
||||
st.bytes_fetched - open.bytes_fetched
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -65,7 +65,8 @@ const MSG_FLAG_DONTSHARE: u8 = 0x04;
|
||||
/// that may be mid-update.
|
||||
///
|
||||
/// Every method is one self-contained edit: it re-reads the file's
|
||||
/// metadata, applies the change, and syncs the file before returning.
|
||||
/// metadata (from the file it holds open, never by path), applies the
|
||||
/// change, and syncs the file before returning.
|
||||
///
|
||||
/// # What it can change
|
||||
///
|
||||
@@ -750,9 +751,17 @@ impl FileEditor {
|
||||
/// consistent: a metadata cache image, paged or persistent free-space
|
||||
/// management, a multi-file driver, a file another writer has marked
|
||||
/// open (superblock version 3 consistency flags).
|
||||
///
|
||||
/// The path is only used to open the file: every edit is planned from
|
||||
/// and written to the file opened here, even if the path is renamed,
|
||||
/// replaced or (relative) resolved from another working directory
|
||||
/// later. [`path`](Self::path) is the absolute path it had at open.
|
||||
pub fn open<P: AsRef<Path>>(path: P) -> Result<Self, Error> {
|
||||
let path = path.as_ref().to_path_buf();
|
||||
let file = OpenOptions::new().read(true).write(true).open(&path)?;
|
||||
let file = OpenOptions::new()
|
||||
.read(true)
|
||||
.write(true)
|
||||
.open(path.as_ref())?;
|
||||
let path = std::fs::canonicalize(path.as_ref())?;
|
||||
match file.try_lock() {
|
||||
Ok(()) => {}
|
||||
Err(TryLockError::WouldBlock) => {
|
||||
@@ -768,16 +777,81 @@ impl FileEditor {
|
||||
file,
|
||||
free: FreeList::default(),
|
||||
};
|
||||
let f = File::open(&ed.path)?;
|
||||
let f = ed.plan_reader()?;
|
||||
check_editable(&f)?;
|
||||
Ok(ed)
|
||||
}
|
||||
|
||||
/// The file's path.
|
||||
/// The file's absolute path when it was opened (it may have been
|
||||
/// renamed since; the editor keeps editing the file it opened).
|
||||
pub fn path(&self) -> &Path {
|
||||
&self.path
|
||||
}
|
||||
|
||||
/// A reader over the file this editor holds, as last written: the file
|
||||
/// opened by [`open`](Self::open), not whatever its path names now.
|
||||
///
|
||||
/// It opens the file anew (read-only), so it does not share the
|
||||
/// editor's lock and stays usable after the editor is dropped: on Linux
|
||||
/// through `/proc/self/fd`, which reaches the held file even after its
|
||||
/// path was renamed or replaced; elsewhere by the path the file had at
|
||||
/// open, refused with [`Error::Io`] when that path no longer names the
|
||||
/// held file (on Unix, compared by device and inode; Windows cannot
|
||||
/// check). With the `mmap` feature the reader maps the file: edits
|
||||
/// through the editor change the bytes it sees, so take a new reader
|
||||
/// after each edit rather than reading through an old one while an edit
|
||||
/// runs.
|
||||
pub fn reader(&self) -> Result<File, Error> {
|
||||
let dir = self.path.parent().map(Path::to_path_buf);
|
||||
File::from_std_file(self.reopen()?, dir)
|
||||
}
|
||||
|
||||
/// A new read-only open file description of the held file (see
|
||||
/// [`reader`](Self::reader)).
|
||||
fn reopen(&self) -> Result<std::fs::File, Error> {
|
||||
#[cfg(target_os = "linux")]
|
||||
{
|
||||
use std::os::fd::AsRawFd;
|
||||
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
|
||||
if let Ok(f) = std::fs::File::open(proc) {
|
||||
return Ok(f);
|
||||
}
|
||||
}
|
||||
let f = std::fs::File::open(&self.path)?;
|
||||
#[cfg(unix)]
|
||||
{
|
||||
use std::os::unix::fs::MetadataExt;
|
||||
let (a, b) = (self.file.metadata()?, f.metadata()?);
|
||||
if (a.dev(), a.ino()) != (b.dev(), b.ino()) {
|
||||
return Err(Error::Io(std::io::Error::other(format!(
|
||||
"{} no longer names the file being edited (renamed or replaced)",
|
||||
self.path.display()
|
||||
))));
|
||||
}
|
||||
}
|
||||
Ok(f)
|
||||
}
|
||||
|
||||
/// A reader over the held file for planning an edit (dropped before the
|
||||
/// edit writes). On Linux a new open file description through
|
||||
/// `/proc/self/fd`: a mapping of a clone of the held descriptor would
|
||||
/// share its `flock`, and a process forked meanwhile (any
|
||||
/// `std::process::Command` on another thread) would briefly keep the
|
||||
/// lock alive after the editor is dropped. Elsewhere a clone of the
|
||||
/// held descriptor, which follows the file wherever its path goes.
|
||||
fn plan_reader(&self) -> Result<File, Error> {
|
||||
let dir = self.path.parent().map(Path::to_path_buf);
|
||||
#[cfg(target_os = "linux")]
|
||||
{
|
||||
use std::os::fd::AsRawFd;
|
||||
let proc = format!("/proc/self/fd/{}", self.file.as_raw_fd());
|
||||
if let Ok(f) = std::fs::File::open(proc) {
|
||||
return File::from_std_file(f, dir);
|
||||
}
|
||||
}
|
||||
File::from_std_file(self.file.try_clone()?, dir)
|
||||
}
|
||||
|
||||
/// Bytes earlier edits of this editor freed that later ones can still
|
||||
/// reuse.
|
||||
pub fn reusable_bytes(&self) -> u64 {
|
||||
@@ -794,7 +868,7 @@ impl FileEditor {
|
||||
&mut self,
|
||||
op: impl FnOnce(&File, &mut Image<'_>) -> Result<R, Error>,
|
||||
) -> Result<R, Error> {
|
||||
let f = File::open(&self.path)?;
|
||||
let f = self.plan_reader()?;
|
||||
check_editable(&f)?;
|
||||
let sb = f.superblock().clone();
|
||||
let user_block = f.user_block_size();
|
||||
@@ -877,7 +951,7 @@ impl FileEditor {
|
||||
/// the fill value (`H5D__chunk_prune_by_extent`).
|
||||
pub fn resize(&mut self, path: &str, shape: &[u64]) -> Result<(), Error> {
|
||||
self.edit(|f, img| {
|
||||
let t = Target::load(f, path)?;
|
||||
let mut t = Target::load(f, path)?;
|
||||
let dims = t.dims().to_vec();
|
||||
if shape.len() != dims.len() {
|
||||
return Err(Error::InvalidArgument(format!(
|
||||
@@ -889,7 +963,12 @@ impl FileEditor {
|
||||
if shape == dims.as_slice() {
|
||||
return Ok(());
|
||||
}
|
||||
let max = t.ds.max_dimensions.clone().unwrap_or_else(|| dims.clone());
|
||||
// No maximum recorded means the current dimensions (see below).
|
||||
let record_max = t.ds.max_dimensions.is_none();
|
||||
let max =
|
||||
t.ds.max_dimensions
|
||||
.get_or_insert_with(|| dims.clone())
|
||||
.clone();
|
||||
for d in 0..dims.len() {
|
||||
if shape[d] > max[d] {
|
||||
return Err(Error::InvalidArgument(format!(
|
||||
@@ -928,7 +1007,36 @@ impl FileEditor {
|
||||
}
|
||||
put_uint(&mut dims_bytes[d * ls..], n, img.ls);
|
||||
}
|
||||
hdr.patch(img, i, first, &dims_bytes)?;
|
||||
if !record_max {
|
||||
hdr.patch(img, i, first, &dims_bytes)?;
|
||||
} else {
|
||||
// No maximum recorded (clawhdf5's writer, for a dataset
|
||||
// created without a maxshape). libhdf5 never writes such a
|
||||
// dataspace: `H5S_set_extent_simple` records the maximum,
|
||||
// equal to the dimensions when none is given. Reading one,
|
||||
// libhdf5 takes the maximum to be the *current* dimensions
|
||||
// (`H5S_extent_get_dims`), so changing them would also
|
||||
// change the maximum the chunk index was built with — the
|
||||
// Fixed Array linearises chunks by it — and move every
|
||||
// existing chunk. Record the maximum libhdf5 would have
|
||||
// written, the dimensions before this resize, so the index
|
||||
// keeps its layout and the dataset can grow back to them.
|
||||
let body_len = first + dims.len() * ls;
|
||||
if body.len() < body_len || body[2] & !0x01 != 0 {
|
||||
return Err(Error::Unsupported("dataspace message layout".into()));
|
||||
}
|
||||
let mut new_body = body[..first].to_vec();
|
||||
new_body[2] |= 0x01;
|
||||
new_body.extend_from_slice(&dims_bytes);
|
||||
let at = new_body.len();
|
||||
new_body.resize(at + dims.len() * ls, 0);
|
||||
for (d, &n) in dims.iter().enumerate() {
|
||||
put_uint(&mut new_body[at + d * ls..], n, img.ls);
|
||||
}
|
||||
let (flags, corder) = (hdr.msgs[i].flags, hdr.msgs[i].corder);
|
||||
hdr.delete(img, i)?;
|
||||
hdr.insert(img, MSG_DATASPACE, flags, &new_body, corder)?;
|
||||
}
|
||||
let fill = fill_info(img, &hdr)?;
|
||||
hdr.finish(img)?;
|
||||
let expand = shape.iter().zip(&dims).any(|(n, o)| n > o);
|
||||
|
||||
@@ -47,6 +47,7 @@ pub mod lazy;
|
||||
#[cfg(feature = "mmap")]
|
||||
pub mod mmap_file;
|
||||
pub mod reader;
|
||||
mod swmr;
|
||||
pub mod types;
|
||||
pub mod vlen;
|
||||
pub mod writer;
|
||||
@@ -57,6 +58,7 @@ pub use lazy::{LazyDataset, LazyFile, LazyGroup};
|
||||
#[cfg(feature = "mmap")]
|
||||
pub use mmap_file::{MmapDataset, MmapFile, MmapGroup};
|
||||
pub use reader::{Dataset, File, Group, SharedStorage, VdsResolver};
|
||||
pub use swmr::{FileStorage, SWMR_READ_ATTEMPTS};
|
||||
pub use types::{AttrValue, DType};
|
||||
pub use vlen::VlenValue;
|
||||
pub use writer::FileBuilder;
|
||||
|
||||
+608
-218
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,362 @@
|
||||
//! Reading files a SWMR writer is still appending to (see
|
||||
//! `docs/design/swmr.md` and [`File::open_swmr`](crate::File::open_swmr)):
|
||||
//! a [`FileStorage`] whose length grows with the file, and the bounded
|
||||
//! retries of operations a concurrent write can make fail.
|
||||
|
||||
use std::borrow::Cow;
|
||||
use std::cell::Cell;
|
||||
use std::path::Path;
|
||||
use std::sync::atomic::{AtomicU64, Ordering};
|
||||
use std::time::Duration;
|
||||
|
||||
use clawhdf5_format::error::FormatError;
|
||||
use clawhdf5_format::storage::Storage;
|
||||
|
||||
use crate::error::Error;
|
||||
|
||||
/// How many times a live file ([`File::open_swmr`](crate::File::open_swmr))
|
||||
/// tries an operation that fails with an error a concurrent write can cause
|
||||
/// before returning the error: 100, libhdf5's default number of metadata
|
||||
/// read attempts for SWMR access (`H5Pset_metadata_read_attempts`).
|
||||
pub const SWMR_READ_ATTEMPTS: u32 = 100;
|
||||
|
||||
/// Longest pause between two attempts.
|
||||
const MAX_PAUSE: Duration = Duration::from_millis(10);
|
||||
|
||||
/// A local file read with positioned reads (`pread` on Unix, `seek_read` on
|
||||
/// Windows), never mapped, whose [`Storage::len`] is the file's length at
|
||||
/// the time of the call: a [`Storage`] for a file that another process is
|
||||
/// appending to. A read past the end is short, as the trait allows.
|
||||
#[derive(Debug)]
|
||||
pub struct FileStorage {
|
||||
file: std::fs::File,
|
||||
#[cfg(not(any(unix, windows)))]
|
||||
lock: std::sync::Mutex<()>,
|
||||
}
|
||||
|
||||
impl FileStorage {
|
||||
/// Open the file at `path` for reading.
|
||||
pub fn open<P: AsRef<Path>>(path: P) -> std::io::Result<Self> {
|
||||
Ok(Self::new(std::fs::File::open(path)?))
|
||||
}
|
||||
|
||||
/// A storage over an open file.
|
||||
pub fn new(file: std::fs::File) -> Self {
|
||||
Self {
|
||||
file,
|
||||
#[cfg(not(any(unix, windows)))]
|
||||
lock: std::sync::Mutex::new(()),
|
||||
}
|
||||
}
|
||||
|
||||
/// Up to `buf.len()` bytes at `offset`; fewer only at the end of file.
|
||||
fn read_into(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
|
||||
let mut got = 0;
|
||||
while got < buf.len() {
|
||||
match self.read_once(offset + got as u64, &mut buf[got..]) {
|
||||
Ok(0) => break,
|
||||
Ok(n) => got += n,
|
||||
Err(e) if e.kind() == std::io::ErrorKind::Interrupted => {}
|
||||
Err(e) => return Err(e),
|
||||
}
|
||||
}
|
||||
Ok(got)
|
||||
}
|
||||
|
||||
#[cfg(unix)]
|
||||
fn read_once(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
|
||||
std::os::unix::fs::FileExt::read_at(&self.file, buf, offset)
|
||||
}
|
||||
|
||||
#[cfg(windows)]
|
||||
fn read_once(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
|
||||
std::os::windows::fs::FileExt::seek_read(&self.file, buf, offset)
|
||||
}
|
||||
|
||||
#[cfg(not(any(unix, windows)))]
|
||||
fn read_once(&self, offset: u64, buf: &mut [u8]) -> std::io::Result<usize> {
|
||||
use std::io::{Read, Seek, SeekFrom};
|
||||
let _guard = self.lock.lock().unwrap_or_else(|p| p.into_inner());
|
||||
let mut f = &self.file;
|
||||
f.seek(SeekFrom::Start(offset))?;
|
||||
f.read(buf)
|
||||
}
|
||||
}
|
||||
|
||||
impl Storage for FileStorage {
|
||||
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
|
||||
// Never allocate more than the file holds, whatever a (possibly
|
||||
// hostile) size field asked for.
|
||||
let avail = self.len().saturating_sub(offset);
|
||||
let len = usize::try_from(avail).map_or(len, |a| a.min(len));
|
||||
let mut buf = vec![0u8; len];
|
||||
let got = self
|
||||
.read_into(offset, &mut buf)
|
||||
.map_err(|e| FormatError::Storage(format!("read of {len} bytes at {offset}: {e}")))?;
|
||||
buf.truncate(got);
|
||||
Ok(Cow::Owned(buf))
|
||||
}
|
||||
|
||||
fn len(&self) -> u64 {
|
||||
self.file.metadata().map_or(0, |m| m.len())
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether `e` can be caused by reading a structure while a SWMR writer
|
||||
/// rewrites it or has not yet written it, so that the operation is worth
|
||||
/// running again (see [`is_transient_format`]).
|
||||
pub(crate) fn is_transient(e: &Error) -> bool {
|
||||
match e {
|
||||
Error::Format(f) => is_transient_format(f),
|
||||
Error::Io(io) => io.kind() == std::io::ErrorKind::UnexpectedEof,
|
||||
_ => false,
|
||||
}
|
||||
}
|
||||
|
||||
/// The failures libhdf5's SWMR reader retries, and nothing else. libhdf5
|
||||
/// (`H5C__load_entry`) reads a metadata structure again when its checksum
|
||||
/// fails, and when the prefix it decodes before the checksum to learn the
|
||||
/// structure's size does not decode (for an object header, its signature
|
||||
/// and version: a header whose every byte is garbled fails there). A read
|
||||
/// past the file's current end is short here; libhdf5 reads zeros there,
|
||||
/// which then fail the checksum. So:
|
||||
///
|
||||
/// - [`FormatError::ChecksumMismatch`], of any checksummed structure;
|
||||
/// - [`FormatError::UnexpectedEof`], a read past the current end;
|
||||
/// - [`FormatError::InvalidObjectHeaderSignature`] and
|
||||
/// [`FormatError::InvalidObjectHeaderVersion`], the object header prefix.
|
||||
///
|
||||
/// Every other error — an unsupported version or message, a file that is not
|
||||
/// HDF5, a structure that is corrupt behind a valid checksum — is returned
|
||||
/// at once: a concurrent write does not cause it, and retrying it only
|
||||
/// costs up to a second of pauses.
|
||||
pub(crate) fn is_transient_format(e: &FormatError) -> bool {
|
||||
use FormatError as F;
|
||||
matches!(
|
||||
e,
|
||||
F::ChecksumMismatch { .. }
|
||||
| F::UnexpectedEof { .. }
|
||||
| F::InvalidObjectHeaderSignature
|
||||
| F::InvalidObjectHeaderVersion(_)
|
||||
)
|
||||
}
|
||||
|
||||
thread_local! {
|
||||
/// Set while this thread runs an operation under [`retry`], so an
|
||||
/// operation made of retried operations retries as a whole, not each
|
||||
/// part up to the limit.
|
||||
static RETRYING: Cell<bool> = const { Cell::new(false) };
|
||||
}
|
||||
|
||||
/// A live file's retry policy: attempts per operation, and a count of the
|
||||
/// retries made.
|
||||
#[derive(Debug)]
|
||||
pub(crate) struct Retries {
|
||||
pub(crate) attempts: u32,
|
||||
retried: AtomicU64,
|
||||
}
|
||||
|
||||
impl Default for Retries {
|
||||
fn default() -> Self {
|
||||
Self {
|
||||
attempts: SWMR_READ_ATTEMPTS,
|
||||
retried: AtomicU64::new(0),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl Retries {
|
||||
/// [`retry`] with this policy, counting the retries.
|
||||
pub(crate) fn retry<T>(&self, op: impl FnMut() -> Result<T, Error>) -> Result<T, Error> {
|
||||
retry(self.attempts, &self.retried, op)
|
||||
}
|
||||
|
||||
/// Retries made so far.
|
||||
pub(crate) fn retries(&self) -> u64 {
|
||||
self.retried.load(Ordering::Relaxed)
|
||||
}
|
||||
}
|
||||
|
||||
/// Run `op` up to `attempts` times while it fails with a transient error
|
||||
/// ([`is_transient`]), pausing 1 µs, 2 µs, 4 µs, … up to 10 ms between
|
||||
/// attempts, and adding each retry to `retried`; return its first success
|
||||
/// or last error. Inside another `retry` on the same thread, `op` runs
|
||||
/// once.
|
||||
fn retry<T>(
|
||||
attempts: u32,
|
||||
retried: &AtomicU64,
|
||||
mut op: impl FnMut() -> Result<T, Error>,
|
||||
) -> Result<T, Error> {
|
||||
if RETRYING.with(Cell::get) {
|
||||
return op();
|
||||
}
|
||||
struct Reset;
|
||||
impl Drop for Reset {
|
||||
fn drop(&mut self) {
|
||||
RETRYING.with(|r| r.set(false));
|
||||
}
|
||||
}
|
||||
RETRYING.with(|r| r.set(true));
|
||||
let _reset = Reset;
|
||||
let mut pause = Duration::from_micros(1);
|
||||
let mut attempt = 1;
|
||||
loop {
|
||||
match op() {
|
||||
Err(e) if attempt < attempts && is_transient(&e) => {
|
||||
retried.fetch_add(1, Ordering::Relaxed);
|
||||
std::thread::sleep(pause);
|
||||
pause = (pause * 2).min(MAX_PAUSE);
|
||||
attempt += 1;
|
||||
}
|
||||
result => return result,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn file_storage_reads_what_the_file_holds_now() {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let path = dir.path().join("grow.bin");
|
||||
std::fs::write(&path, b"hello").unwrap();
|
||||
let s = FileStorage::open(&path).unwrap();
|
||||
assert_eq!(s.len(), 5);
|
||||
assert_eq!(&*s.read_at(1, 3).unwrap(), b"ell");
|
||||
assert_eq!(&*s.read_at(3, 10).unwrap(), b"lo");
|
||||
assert!(s.read_at(9, 4).unwrap().is_empty());
|
||||
// The file grows after the storage was opened.
|
||||
use std::io::Write;
|
||||
std::fs::OpenOptions::new()
|
||||
.append(true)
|
||||
.open(&path)
|
||||
.unwrap()
|
||||
.write_all(b", world")
|
||||
.unwrap();
|
||||
assert_eq!(s.len(), 12);
|
||||
assert_eq!(&*s.read_at(3, 100).unwrap(), b"lo, world");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn permanent_errors_are_returned_at_once() {
|
||||
// Errors a concurrent write does not cause: none is retried, so
|
||||
// none costs the pauses (about 0.9 s for 100 attempts).
|
||||
let permanent = [
|
||||
FormatError::SignatureNotFound,
|
||||
FormatError::UnsupportedVersion(9),
|
||||
FormatError::UnsupportedMessage(0x99),
|
||||
FormatError::TruncatedFile {
|
||||
stored_eof: 10,
|
||||
actual_len: 5,
|
||||
},
|
||||
FormatError::InvalidDatatypeClass(15),
|
||||
FormatError::InvalidLayoutVersion(9),
|
||||
FormatError::InvalidBTreeSignature,
|
||||
FormatError::ChunkedReadError("x".into()),
|
||||
FormatError::DecompressionError("x".into()),
|
||||
FormatError::Fletcher32Mismatch {
|
||||
expected: 1,
|
||||
computed: 2,
|
||||
},
|
||||
];
|
||||
let n = AtomicU64::new(0);
|
||||
let started = std::time::Instant::now();
|
||||
for e in permanent {
|
||||
assert!(!is_transient_format(&e), "{e:?}");
|
||||
let mut calls = 0;
|
||||
let r: Result<(), Error> = retry(SWMR_READ_ATTEMPTS, &n, || {
|
||||
calls += 1;
|
||||
Err(Error::Format(e.clone()))
|
||||
});
|
||||
assert!(r.is_err());
|
||||
assert_eq!(calls, 1, "{e:?}");
|
||||
}
|
||||
assert_eq!(n.load(Ordering::Relaxed), 0);
|
||||
assert!(started.elapsed() < Duration::from_millis(50));
|
||||
assert!(!is_transient(&Error::Io(std::io::Error::other("x"))));
|
||||
|
||||
// The transient ones are retried to the limit.
|
||||
for e in [
|
||||
FormatError::ChecksumMismatch {
|
||||
expected: 1,
|
||||
computed: 2,
|
||||
},
|
||||
FormatError::UnexpectedEof {
|
||||
expected: 8,
|
||||
available: 0,
|
||||
},
|
||||
FormatError::InvalidObjectHeaderSignature,
|
||||
FormatError::InvalidObjectHeaderVersion(0x4f),
|
||||
] {
|
||||
let mut calls = 0;
|
||||
let _: Result<(), Error> = retry(3, &n, || {
|
||||
calls += 1;
|
||||
Err(Error::Format(e.clone()))
|
||||
});
|
||||
assert_eq!(calls, 3, "{e:?}");
|
||||
}
|
||||
assert!(is_transient(&Error::Io(
|
||||
std::io::ErrorKind::UnexpectedEof.into()
|
||||
)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn retry_runs_again_only_for_transient_errors() {
|
||||
let n = AtomicU64::new(0);
|
||||
let mut calls = 0;
|
||||
let r: Result<u32, Error> = retry(5, &n, || {
|
||||
calls += 1;
|
||||
if calls < 3 {
|
||||
Err(Error::Format(FormatError::ChecksumMismatch {
|
||||
expected: 1,
|
||||
computed: 2,
|
||||
}))
|
||||
} else {
|
||||
Ok(7)
|
||||
}
|
||||
});
|
||||
assert_eq!(r.unwrap(), 7);
|
||||
assert_eq!(calls, 3);
|
||||
assert_eq!(n.load(Ordering::Relaxed), 2);
|
||||
|
||||
// Gives up after `attempts`.
|
||||
let mut calls = 0;
|
||||
let r: Result<(), Error> = retry(4, &n, || {
|
||||
calls += 1;
|
||||
Err(Error::Format(FormatError::UnexpectedEof {
|
||||
expected: 8,
|
||||
available: 0,
|
||||
}))
|
||||
});
|
||||
assert!(r.is_err());
|
||||
assert_eq!(calls, 4);
|
||||
assert_eq!(n.load(Ordering::Relaxed), 5);
|
||||
|
||||
// A permanent error is returned at once.
|
||||
let mut calls = 0;
|
||||
let r: Result<(), Error> = retry(4, &n, || {
|
||||
calls += 1;
|
||||
Err(Error::Format(FormatError::UnsupportedFilter(999)))
|
||||
});
|
||||
assert!(r.is_err());
|
||||
assert_eq!(calls, 1);
|
||||
assert_eq!(n.load(Ordering::Relaxed), 5);
|
||||
|
||||
// Nested: the inner operation runs once per outer attempt.
|
||||
let mut inner = 0;
|
||||
let mut outer = 0;
|
||||
let r: Result<(), Error> = retry(3, &n, || {
|
||||
outer += 1;
|
||||
retry(3, &n, || {
|
||||
inner += 1;
|
||||
Err(Error::Format(FormatError::InvalidObjectHeaderVersion(0x76)))
|
||||
})
|
||||
});
|
||||
assert!(r.is_err());
|
||||
assert_eq!((outer, inner), (3, 3));
|
||||
// The flag is reset afterwards.
|
||||
assert!(!RETRYING.with(Cell::get));
|
||||
}
|
||||
}
|
||||
@@ -170,16 +170,38 @@ pub(crate) fn read_attrs<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
),
|
||||
crate::Error,
|
||||
> {
|
||||
read_attrs_reporting(file_data, header, offset_size, length_size)
|
||||
.map(|(attrs, errors, _)| (attrs, errors))
|
||||
}
|
||||
|
||||
/// What [`read_attrs_reporting`] returns: the attributes, the errors of
|
||||
/// those left out, and the errors behind values returned as
|
||||
/// [`AttrValue::Raw`].
|
||||
pub(crate) type AttrsReport = (
|
||||
HashMap<String, AttrValue>,
|
||||
Vec<clawhdf5_format::error::FormatError>,
|
||||
Vec<clawhdf5_format::error::FormatError>,
|
||||
);
|
||||
|
||||
/// [`read_attrs`], also returning the errors of reading the file for a
|
||||
/// value (a variable-length string's global heap) that was returned as
|
||||
/// [`AttrValue::Raw`] instead: a live file retries on them (see
|
||||
/// `File::open_swmr`).
|
||||
pub(crate) fn read_attrs_reporting<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
file_data: &S,
|
||||
header: &clawhdf5_format::object_header::ObjectHeader,
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
) -> Result<AttrsReport, crate::Error> {
|
||||
let (msgs, errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in(
|
||||
file_data,
|
||||
header,
|
||||
offset_size,
|
||||
length_size,
|
||||
)?;
|
||||
Ok((
|
||||
attrs_to_map(&msgs, file_data, offset_size, length_size),
|
||||
errors,
|
||||
))
|
||||
let mut read_errors = Vec::new();
|
||||
let map = attrs_to_map_reporting(&msgs, file_data, offset_size, length_size, &mut read_errors);
|
||||
Ok((map, errors, read_errors))
|
||||
}
|
||||
|
||||
/// The attribute called `name` on the object with header `header`, decoded
|
||||
@@ -192,30 +214,50 @@ pub(crate) fn read_attr<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
) -> Result<Option<AttrValue>, crate::Error> {
|
||||
let Some(msg) = clawhdf5_format::attribute::find_attribute_in(
|
||||
read_attr_reporting(file_data, header, name, offset_size, length_size).map(|(v, _)| v)
|
||||
}
|
||||
|
||||
/// [`read_attr`], also returning the errors of the attributes it could not
|
||||
/// read on the way (the one asked for may be among them) and of reading the
|
||||
/// file for its value, if one made it [`AttrValue::Raw`] (see
|
||||
/// [`read_attrs_reporting`]).
|
||||
pub(crate) fn read_attr_reporting<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
file_data: &S,
|
||||
header: &clawhdf5_format::object_header::ObjectHeader,
|
||||
name: &str,
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
) -> Result<(Option<AttrValue>, Vec<clawhdf5_format::error::FormatError>), crate::Error> {
|
||||
let (found, mut read_errors) = clawhdf5_format::attribute::find_attribute_reporting_in(
|
||||
file_data,
|
||||
header,
|
||||
name,
|
||||
offset_size,
|
||||
length_size,
|
||||
)?
|
||||
else {
|
||||
return Ok(None);
|
||||
)?;
|
||||
let Some(msg) = found else {
|
||||
return Ok((None, read_errors));
|
||||
};
|
||||
Ok(attrs_to_map(
|
||||
let value = attrs_to_map_reporting(
|
||||
std::slice::from_ref(&msg),
|
||||
file_data,
|
||||
offset_size,
|
||||
length_size,
|
||||
&mut read_errors,
|
||||
)
|
||||
.remove(name))
|
||||
.remove(name);
|
||||
Ok((value, read_errors))
|
||||
}
|
||||
|
||||
pub(crate) fn attrs_to_map<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
/// The attributes `attrs` by name, decoded; an error reading the file for
|
||||
/// a value that is returned as [`AttrValue::Raw`] because of it is added
|
||||
/// to `read_errors`.
|
||||
pub(crate) fn attrs_to_map_reporting<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
attrs: &[clawhdf5_format::attribute::AttributeMessage],
|
||||
file_data: &S,
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
read_errors: &mut Vec<clawhdf5_format::error::FormatError>,
|
||||
) -> HashMap<String, AttrValue> {
|
||||
let mut map = HashMap::new();
|
||||
for attr in attrs {
|
||||
@@ -224,13 +266,11 @@ pub(crate) fn attrs_to_map<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
// verbatim as `AttrValue::Raw` rather than dropped — a partial
|
||||
// attribute list with no indication anything is missing is worse than
|
||||
// an undecoded value.
|
||||
let val =
|
||||
decode_attr_value(attr, file_data, offset_size, length_size).unwrap_or_else(|| {
|
||||
AttrValue::Raw {
|
||||
datatype: attr.datatype.clone(),
|
||||
shape: attr.dataspace.dimensions.clone(),
|
||||
data: attr.raw_data.clone(),
|
||||
}
|
||||
let val = decode_attr_value(attr, file_data, offset_size, length_size, read_errors)
|
||||
.unwrap_or_else(|| AttrValue::Raw {
|
||||
datatype: attr.datatype.clone(),
|
||||
shape: attr.dataspace.dimensions.clone(),
|
||||
data: attr.raw_data.clone(),
|
||||
});
|
||||
map.insert(attr.name.clone(), val);
|
||||
}
|
||||
@@ -271,6 +311,7 @@ fn decode_attr_value<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
file_data: &S,
|
||||
offset_size: u8,
|
||||
length_size: u8,
|
||||
read_errors: &mut Vec<clawhdf5_format::error::FormatError>,
|
||||
) -> Option<AttrValue> {
|
||||
use clawhdf5_format::datatype::Datatype;
|
||||
|
||||
@@ -310,9 +351,13 @@ fn decode_attr_value<S: clawhdf5_format::storage::Storage + ?Sized>(
|
||||
Datatype::VariableLength {
|
||||
is_string: true, ..
|
||||
} => {
|
||||
let strings = attr
|
||||
.read_vl_strings_in(file_data, offset_size, length_size)
|
||||
.ok()?;
|
||||
let strings = match attr.read_vl_strings_in(file_data, offset_size, length_size) {
|
||||
Ok(strings) => strings,
|
||||
Err(e) => {
|
||||
read_errors.push(e);
|
||||
return None;
|
||||
}
|
||||
};
|
||||
if strings.len() == 1 {
|
||||
Some(AttrValue::String(strings[0].clone()))
|
||||
} else {
|
||||
|
||||
@@ -0,0 +1,287 @@
|
||||
//! `FileEditor::resize` on chunked datasets whose dataspace records no
|
||||
//! maximum dimensions, as clawhdf5's writer stored a dataset created without
|
||||
//! a `maxshape` up to 2.7.0 (`fixtures/chunked_no_maxshape_v2_7_0.h5`). libhdf5 never writes such a dataspace (`H5S_set_extent_simple`
|
||||
//! always records the maximum, equal to the dimensions when none is given),
|
||||
//! and its Fixed Array chunk index linearises chunks by the maximum
|
||||
//! dimensions. The editor therefore records the maximum libhdf5 would have
|
||||
//! written (the dimensions the index was built with) before it changes the
|
||||
//! current ones, so existing chunks stay where the index put them and the
|
||||
//! dataset can grow back to its original extent.
|
||||
//!
|
||||
//! The writer now records the maximum too, so h5py can resize what it writes.
|
||||
//!
|
||||
//! Checked against a model of the expected values, with our reader and with
|
||||
//! h5py (`CLAWHDF5_PYTHON`; skipped without it unless
|
||||
//! `CLAWHDF5_REQUIRE_INTEROP=1`), on files clawhdf5 (old and new) and h5py
|
||||
//! wrote.
|
||||
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::process::Command;
|
||||
|
||||
use clawhdf5::{Error, File, FileBuilder, FileEditor};
|
||||
|
||||
fn python() -> String {
|
||||
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||
}
|
||||
|
||||
fn h5py_ok() -> bool {
|
||||
let ok = Command::new(python())
|
||||
.args(["-c", "import h5py, numpy"])
|
||||
.output()
|
||||
.is_ok_and(|o| o.status.success());
|
||||
if !ok {
|
||||
assert!(
|
||||
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||
"CLAWHDF5_REQUIRE_INTEROP=1 but h5py/numpy is not available"
|
||||
);
|
||||
eprintln!("SKIP (h5py part): h5py/numpy not available");
|
||||
}
|
||||
ok
|
||||
}
|
||||
|
||||
fn py(script: &str) -> String {
|
||||
let o = Command::new(python())
|
||||
.args(["-c", script])
|
||||
.output()
|
||||
.expect("run python");
|
||||
assert!(
|
||||
o.status.success(),
|
||||
"python failed:\n{script}\nSTDERR: {}",
|
||||
String::from_utf8_lossy(&o.stderr)
|
||||
);
|
||||
String::from_utf8_lossy(&o.stdout).trim().to_string()
|
||||
}
|
||||
|
||||
/// Row-major values of a 2-D model after resizing `data` (shape `old`) to
|
||||
/// `new`: kept elements keep their values, new ones are 0 (the fill value).
|
||||
fn resized(data: &[f32], old: [u64; 2], new: [u64; 2]) -> Vec<f32> {
|
||||
let mut out = vec![0f32; (new[0] * new[1]) as usize];
|
||||
for r in 0..old[0].min(new[0]) {
|
||||
for c in 0..old[1].min(new[1]) {
|
||||
out[(r * new[1] + c) as usize] = data[(r * old[1] + c) as usize];
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Our reader and (when available) h5py read `expect` at `shape`.
|
||||
fn check(path: &Path, name: &str, shape: [u64; 2], expect: &[f32], with_h5py: bool) {
|
||||
let f = File::open(path).unwrap();
|
||||
let d = f.dataset(name).unwrap();
|
||||
assert_eq!(d.shape().unwrap(), shape);
|
||||
assert_eq!(
|
||||
d.read_f32().unwrap(),
|
||||
expect,
|
||||
"{name}: our reader at {shape:?}"
|
||||
);
|
||||
if with_h5py {
|
||||
let got = py(&format!(
|
||||
"import h5py, numpy as np\n\
|
||||
with h5py.File({p:?}, 'r') as f:\n\
|
||||
\x20 d = f[{name:?}][()]\n\
|
||||
print(d.shape, ','.join(repr(float(x)) for x in d.ravel()))",
|
||||
p = path.to_str().unwrap()
|
||||
));
|
||||
let want = format!(
|
||||
"({}, {}) {}",
|
||||
shape[0],
|
||||
shape[1],
|
||||
expect
|
||||
.iter()
|
||||
.map(|x| format!("{:?}", f64::from(*x)))
|
||||
.collect::<Vec<_>>()
|
||||
.join(",")
|
||||
);
|
||||
assert_eq!(got, want.trim(), "{name}: h5py at {shape:?}");
|
||||
}
|
||||
}
|
||||
|
||||
/// Resize `name` (20 x 20, values 0..400) through a sequence of shrinks,
|
||||
/// zero extents and growth back, checking every step.
|
||||
fn run(path: &Path, name: &str, with_h5py: bool) {
|
||||
let orig: Vec<f32> = (0..400).map(|i| i as f32).collect();
|
||||
let mut shape = [20u64, 20];
|
||||
let mut data = orig.clone();
|
||||
check(path, name, shape, &data, with_h5py);
|
||||
for next in [
|
||||
[15, 15],
|
||||
[3, 2],
|
||||
[20, 20],
|
||||
[1, 1],
|
||||
[1, 0],
|
||||
[0, 0],
|
||||
[7, 20],
|
||||
[20, 13],
|
||||
[20, 20],
|
||||
] {
|
||||
let mut ed = FileEditor::open(path).unwrap();
|
||||
ed.resize(name, &next).unwrap();
|
||||
drop(ed);
|
||||
data = resized(&data, shape, next);
|
||||
shape = next;
|
||||
check(path, name, shape, &data, with_h5py);
|
||||
}
|
||||
// The maximum is the extent the dataset was created with.
|
||||
let mut ed = FileEditor::open(path).unwrap();
|
||||
assert!(matches!(
|
||||
ed.resize(name, &[21, 20]),
|
||||
Err(Error::InvalidArgument(_))
|
||||
));
|
||||
drop(ed);
|
||||
let f = File::open(path).unwrap();
|
||||
assert_eq!(
|
||||
f.dataset(name).unwrap().max_dimensions().unwrap(),
|
||||
Some(vec![20, 20])
|
||||
);
|
||||
drop(f);
|
||||
// A shrink keeps the values it keeps.
|
||||
let mut ed = FileEditor::open(path).unwrap();
|
||||
let vals: Vec<f32> = orig.iter().map(|v| v + 0.5).collect();
|
||||
ed.write_values(name, &clawhdf5::Selection::All, &vals)
|
||||
.unwrap();
|
||||
ed.resize(name, &[15, 15]).unwrap();
|
||||
drop(ed);
|
||||
check(
|
||||
path,
|
||||
name,
|
||||
[15, 15],
|
||||
&resized(&vals, [20, 20], [15, 15]),
|
||||
with_h5py,
|
||||
);
|
||||
}
|
||||
|
||||
fn fixture(dir: &Path, name: &str) -> PathBuf {
|
||||
let path = dir.join(name);
|
||||
std::fs::copy(
|
||||
Path::new(env!("CARGO_MANIFEST_DIR"))
|
||||
.join("tests/fixtures")
|
||||
.join(name),
|
||||
&path,
|
||||
)
|
||||
.unwrap();
|
||||
path
|
||||
}
|
||||
|
||||
/// Files clawhdf5 2.7.0 wrote: a Fixed Array index (Single Chunk for `s`)
|
||||
/// and a dataspace with no maximum. Shrinking scrambled the values
|
||||
/// (released in 2.7.0's `FileEditor`, PR #18).
|
||||
#[test]
|
||||
fn resize_without_stored_maxshape_keeps_values() {
|
||||
let with_h5py = h5py_ok();
|
||||
for name in ["d", "z"] {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
|
||||
run(&path, name, with_h5py);
|
||||
}
|
||||
// A single-chunk dataset and an empty one keep their extents as maxima.
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let path = fixture(dir.path(), "chunked_no_maxshape_v2_7_0.h5");
|
||||
let mut ed = FileEditor::open(&path).unwrap();
|
||||
ed.resize("s", &[2, 4]).unwrap();
|
||||
ed.resize("s", &[3, 4]).unwrap();
|
||||
assert!(matches!(
|
||||
ed.resize("s", &[4, 4]),
|
||||
Err(Error::InvalidArgument(_))
|
||||
));
|
||||
assert!(matches!(
|
||||
ed.resize("e", &[1, 5]),
|
||||
Err(Error::InvalidArgument(_))
|
||||
));
|
||||
ed.resize("e", &[0, 3]).unwrap();
|
||||
drop(ed);
|
||||
let f = File::open(&path).unwrap();
|
||||
let s = f.dataset("s").unwrap();
|
||||
assert_eq!(s.max_dimensions().unwrap(), Some(vec![3, 4]));
|
||||
let mut want: Vec<i32> = (0..12).collect();
|
||||
want[8..].fill(0);
|
||||
assert_eq!(s.read_i32().unwrap(), want);
|
||||
assert_eq!(
|
||||
f.dataset("e").unwrap().max_dimensions().unwrap(),
|
||||
Some(vec![0, 5])
|
||||
);
|
||||
}
|
||||
|
||||
fn written(path: &Path, deflate: bool) {
|
||||
let data: Vec<f32> = (0..400).map(|i| i as f32).collect();
|
||||
let mut b = FileBuilder::new();
|
||||
let d = b
|
||||
.create_dataset("d")
|
||||
.with_f32_data(&data)
|
||||
.with_shape(&[20, 20])
|
||||
.with_chunks(&[6, 6]);
|
||||
if deflate {
|
||||
d.with_deflate(4);
|
||||
}
|
||||
b.write(path).unwrap();
|
||||
}
|
||||
|
||||
/// clawhdf5's writer now records the maximum, as libhdf5 does.
|
||||
#[test]
|
||||
fn resize_file_written_without_maxshape_keeps_values() {
|
||||
let with_h5py = h5py_ok();
|
||||
for deflate in [false, true] {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let path = dir.path().join("cw.h5");
|
||||
written(&path, deflate);
|
||||
let f = File::open(&path).unwrap();
|
||||
assert_eq!(
|
||||
f.dataset("d").unwrap().max_dimensions().unwrap(),
|
||||
Some(vec![20, 20])
|
||||
);
|
||||
drop(f);
|
||||
run(&path, "d", with_h5py);
|
||||
}
|
||||
}
|
||||
|
||||
/// h5py resizing a file clawhdf5 wrote without a maxshape keeps its values
|
||||
/// (it scrambled them while the writer recorded no maximum).
|
||||
#[test]
|
||||
fn h5py_resizes_what_clawhdf5_writes() {
|
||||
if !h5py_ok() {
|
||||
return;
|
||||
}
|
||||
for deflate in [false, true] {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let path = dir.path().join("cw.h5");
|
||||
written(&path, deflate);
|
||||
let out = py(&format!(
|
||||
"import h5py, numpy as np\n\
|
||||
exp = np.arange(400, dtype='f4').reshape(20, 20)\n\
|
||||
with h5py.File({p:?}, 'r+') as f:\n\
|
||||
\x20 f['d'].resize((15, 15))\n\
|
||||
\x20 ok = np.array_equal(f['d'][()], exp[:15, :15])\n\
|
||||
\x20 f['d'].resize((20, 20))\n\
|
||||
\x20 back = f['d'][()]\n\
|
||||
want = np.zeros((20, 20), 'f4'); want[:15, :15] = exp[:15, :15]\n\
|
||||
print(ok and np.array_equal(back, want))",
|
||||
p = path.to_str().unwrap()
|
||||
));
|
||||
assert_eq!(out, "True");
|
||||
let mut want = vec![0f32; 400];
|
||||
for r in 0..15 {
|
||||
for c in 0..15 {
|
||||
want[r * 20 + c] = (r * 20 + c) as f32;
|
||||
}
|
||||
}
|
||||
check(&path, "d", [20, 20], &want, false);
|
||||
}
|
||||
}
|
||||
|
||||
/// h5py's files record the maximum; the same sequence must hold.
|
||||
#[test]
|
||||
fn resize_h5py_file_without_maxshape_keeps_values() {
|
||||
if !h5py_ok() {
|
||||
return;
|
||||
}
|
||||
for libver in ["earliest", "v110", "latest"] {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let path = dir.path().join("hp.h5");
|
||||
py(&format!(
|
||||
"import h5py, numpy as np\n\
|
||||
with h5py.File({p:?}, 'w', libver=({libver:?}, 'latest')) as f:\n\
|
||||
\x20 f.create_dataset('d', data=np.arange(400, dtype='f4').reshape(20, 20), chunks=(6, 6))",
|
||||
p = path.to_str().unwrap()
|
||||
));
|
||||
run(&path, "d", true);
|
||||
}
|
||||
}
|
||||
@@ -162,3 +162,59 @@ fn shrink_then_grow_reads_fill() {
|
||||
raw[..4].fill(0.5);
|
||||
assert_eq!(f.dataset("raw").unwrap().read_f64().unwrap(), raw);
|
||||
}
|
||||
|
||||
/// The editor plans every edit from the file it holds open, never by
|
||||
/// re-opening its path: with the path renamed away and another file put in
|
||||
/// its place, edits go to the held file, planned from its own metadata,
|
||||
/// and the file now at the path is untouched (planning from it and writing
|
||||
/// into the held file corrupted the held one).
|
||||
#[test]
|
||||
fn edits_go_to_the_file_held_not_the_path() {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let a = dir.path().join("a.h5");
|
||||
let b = dir.path().join("b.h5");
|
||||
let mut fb = FileBuilder::new();
|
||||
fb.create_dataset("x")
|
||||
.with_i32_data(&[0; 10])
|
||||
.with_shape(&[10]);
|
||||
fb.create_dataset("big")
|
||||
.with_f64_data(&[1.5; 5000])
|
||||
.with_shape(&[5000]);
|
||||
fb.set_attr("title", AttrValue::String("a".into()));
|
||||
fb.write(&a).unwrap();
|
||||
let mut fb = FileBuilder::new();
|
||||
fb.create_dataset("pad")
|
||||
.with_f64_data(&[2.5; 3000])
|
||||
.with_shape(&[3000]);
|
||||
fb.create_dataset("x")
|
||||
.with_i32_data(&[500; 10])
|
||||
.with_shape(&[10]);
|
||||
fb.write(&b).unwrap();
|
||||
|
||||
let mut ed = FileEditor::open(&a).unwrap();
|
||||
assert!(ed.path().is_absolute());
|
||||
let moved = dir.path().join("moved.h5");
|
||||
std::fs::rename(&a, &moved).unwrap();
|
||||
std::fs::rename(&b, &a).unwrap();
|
||||
let other = std::fs::read(&a).unwrap();
|
||||
|
||||
ed.write_values("x", &Selection::All, &[7i32; 10]).unwrap();
|
||||
let vals: Vec<f64> = (0..50).map(f64::from).collect();
|
||||
ed.set_attr("/", "note", &AttrValue::F64Array(vals.clone()))
|
||||
.unwrap();
|
||||
// The editor's own reader sees the held file.
|
||||
let r = ed.reader().unwrap();
|
||||
assert_eq!(r.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
|
||||
assert!(r.dataset("pad").is_err());
|
||||
drop(r);
|
||||
drop(ed);
|
||||
|
||||
assert!(
|
||||
std::fs::read(&a).unwrap() == other,
|
||||
"the file at the path changed"
|
||||
);
|
||||
let f = File::open(&moved).unwrap();
|
||||
assert_eq!(f.dataset("x").unwrap().read_i32().unwrap(), [7; 10]);
|
||||
assert_eq!(f.dataset("big").unwrap().read_f64().unwrap(), [1.5; 5000]);
|
||||
assert!(matches!(f.root().attr("note").unwrap(), Some(AttrValue::F64Array(v)) if v == vals));
|
||||
}
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -455,3 +455,70 @@ with h5py.File("{p}", "w") as f:
|
||||
assert!(lazy.dataset("line").unwrap().read_i32().is_err());
|
||||
assert!(lazy.dataset("grid").unwrap().read_f64().is_err());
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Several damaged chunks: which one is reported
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// With more than one damaged chunk, the error names the first damaged chunk
|
||||
/// in the chunk index's order, every time and on every read path. The cached
|
||||
/// reader used to walk the chunks in hash-map order, so two opens of the same
|
||||
/// file could report different chunks (`cve-2025-2310.h5`).
|
||||
#[test]
|
||||
fn several_damaged_chunks_report_the_same_chunk_every_time() {
|
||||
skip_if_no_python!();
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let path = dir.path().join("damaged.h5");
|
||||
let p = path.display().to_string();
|
||||
run_python(&format!(
|
||||
r#"
|
||||
import h5py, numpy as np, zlib
|
||||
with h5py.File("{p}", "w") as f:
|
||||
ds = f.create_dataset("d", shape=(512,), chunks=(8,), dtype="<f8", compression="gzip")
|
||||
ds[...] = np.arange(512.0)
|
||||
# Each damaged chunk inflates to a different short length, so each
|
||||
# error names its own chunk.
|
||||
for k, off in enumerate((40, 136, 320, 488)):
|
||||
ds.id.write_direct_chunk((off,), zlib.compress(bytes(8 * (k + 1))))
|
||||
"#
|
||||
));
|
||||
let first = File::open(&path)
|
||||
.unwrap()
|
||||
.dataset("d")
|
||||
.unwrap()
|
||||
.read_f64()
|
||||
.expect_err("damaged chunks must fail")
|
||||
.to_string();
|
||||
// A whole-dataset selection goes through the file's chunk cache, whose index is built
|
||||
// afresh (a new hash map) for every open.
|
||||
let first_raw = File::open(&path)
|
||||
.unwrap()
|
||||
.dataset("d")
|
||||
.unwrap()
|
||||
.read_selection(&Selection::All)
|
||||
.expect_err("damaged chunks must fail")
|
||||
.to_string();
|
||||
for _ in 0..40 {
|
||||
let f = File::open(&path).unwrap();
|
||||
let ds = f.dataset("d").unwrap();
|
||||
assert_eq!(ds.read_f64().unwrap_err().to_string(), first);
|
||||
assert_eq!(
|
||||
ds.read_selection(&Selection::All).unwrap_err().to_string(),
|
||||
first_raw
|
||||
);
|
||||
// A second read goes through the cached chunk index.
|
||||
assert_eq!(
|
||||
ds.read_selection(&Selection::All).unwrap_err().to_string(),
|
||||
first_raw
|
||||
);
|
||||
}
|
||||
let lazy = clawhdf5::LazyFile::open_mmap(&path).unwrap();
|
||||
assert_eq!(
|
||||
lazy.dataset("d")
|
||||
.unwrap()
|
||||
.read_f64()
|
||||
.unwrap_err()
|
||||
.to_string(),
|
||||
first
|
||||
);
|
||||
}
|
||||
|
||||
@@ -966,7 +966,10 @@ fn skipped_optional_filters_are_masked_as_libhdf5_masks_them() {
|
||||
|
||||
/// Files whose chunks all compress are written exactly as before optional
|
||||
/// filters could be skipped: every mask is 0 and nothing else changed. The
|
||||
/// hashes are of the files the writer produced before that change.
|
||||
/// hashes are of the files the writer produced before that change, except
|
||||
/// that a chunked dataset without a maxshape now records its maximum
|
||||
/// dimensions (8 bytes per dimension; `lzf_fixed`, `lzf_single`, and
|
||||
/// `blosc_fixed`).
|
||||
#[cfg(feature = "lzf")]
|
||||
#[test]
|
||||
fn files_whose_chunks_all_compress_are_unchanged() {
|
||||
@@ -985,7 +988,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
|
||||
.with_chunks(&[500])
|
||||
.with_lzf();
|
||||
},
|
||||
(3965, 449169442),
|
||||
(3973, 452644487),
|
||||
),
|
||||
(
|
||||
"lzf_ea_noshuffle",
|
||||
@@ -1017,7 +1020,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
|
||||
.with_chunks(&[3000])
|
||||
.with_lzf();
|
||||
},
|
||||
(546, 690805477),
|
||||
(554, 1394027497),
|
||||
),
|
||||
];
|
||||
#[cfg(feature = "blosc")]
|
||||
@@ -1029,7 +1032,7 @@ fn files_whose_chunks_all_compress_are_unchanged() {
|
||||
.with_chunks(&[1024])
|
||||
.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
|
||||
},
|
||||
(2776, 4278611376),
|
||||
(2784, 1180133244),
|
||||
));
|
||||
for (name, build, want) in &cases {
|
||||
let mut fb = clawhdf5::FileBuilder::new();
|
||||
|
||||
@@ -0,0 +1,847 @@
|
||||
//! Files a libhdf5 SWMR writer (h5py `f.swmr_mode = True`) has open: a copy
|
||||
//! taken mid-write (fixture), and a live file appended to by an h5py writer
|
||||
//! process while clawhdf5 and h5py's own SWMR reader read it (see
|
||||
//! `docs/design/swmr.md`).
|
||||
//!
|
||||
//! The live tests need python3 with h5py; they are skipped without it,
|
||||
//! unless `CLAWHDF5_REQUIRE_INTEROP=1`.
|
||||
|
||||
use std::io::{BufRead, BufReader, Read};
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::process::{Command, Stdio};
|
||||
use std::sync::Arc;
|
||||
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
|
||||
|
||||
use clawhdf5::{File, Selection};
|
||||
|
||||
fn python() -> String {
|
||||
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||
}
|
||||
|
||||
fn interop_required() -> bool {
|
||||
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
|
||||
}
|
||||
|
||||
fn python_available() -> bool {
|
||||
Command::new(python())
|
||||
.args(["-c", "import h5py, numpy"])
|
||||
.output()
|
||||
.map(|o| o.status.success())
|
||||
.unwrap_or(false)
|
||||
}
|
||||
|
||||
macro_rules! skip_if_no_python {
|
||||
() => {
|
||||
if !python_available() {
|
||||
assert!(
|
||||
!interop_required(),
|
||||
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
|
||||
);
|
||||
eprintln!("SKIP: python3 with h5py not available");
|
||||
return;
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
fn fixture(name: &str) -> PathBuf {
|
||||
Path::new(env!("CARGO_MANIFEST_DIR"))
|
||||
.join("tests/fixtures")
|
||||
.join(name)
|
||||
}
|
||||
|
||||
/// `swmr_mid_write.h5`: a copy h5py 3.16 (HDF5 2.0) made of its own file
|
||||
/// while writing it in SWMR mode, after 4 appends of 37 rows to
|
||||
/// `/a` (int64, chunks of 100, no filter: `a[i] = i + 1`) and `/b` (float64
|
||||
/// `(n, 4)`, chunks of 16 x 4, gzip: `b[i, j] = 10 i + j + 1`), each
|
||||
/// followed by a flush. Its superblock (v3) still has the SWMR-write flag
|
||||
/// set and records an end of file of 715 in a 6 030-byte file.
|
||||
fn check_mid_write_copy(f: &File) {
|
||||
let sb = f.superblock();
|
||||
assert_eq!(sb.version, 3);
|
||||
assert!(sb.is_swmr_write());
|
||||
let a = f.dataset("a").unwrap();
|
||||
assert_eq!(a.shape().unwrap(), vec![148]);
|
||||
let want_a: Vec<i64> = (1..=148).collect();
|
||||
assert_eq!(a.read_i64().unwrap(), want_a);
|
||||
let b = f.dataset("b").unwrap();
|
||||
assert_eq!(b.shape().unwrap(), vec![148, 4]);
|
||||
let want_b: Vec<f64> = (0..148)
|
||||
.flat_map(|i| (0..4).map(move |j| (10 * i + j + 1) as f64))
|
||||
.collect();
|
||||
assert_eq!(b.read_f64().unwrap(), want_b);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_copy_taken_mid_write_reads_past_its_recorded_end_of_file() {
|
||||
let path = fixture("swmr_mid_write.h5");
|
||||
// The recorded end of file (715) is far below the file's length; the
|
||||
// chunk indexes and chunks lie past it. libhdf5's SWMR reader does not
|
||||
// bound reads by it, and neither does any open path here.
|
||||
check_mid_write_copy(&File::open(&path).unwrap());
|
||||
check_mid_write_copy(&File::open_buffered(&path).unwrap());
|
||||
check_mid_write_copy(&File::from_bytes(std::fs::read(&path).unwrap()).unwrap());
|
||||
let want_a: Vec<i64> = (1..=148).collect();
|
||||
let want_b: Vec<f64> = (0..148)
|
||||
.flat_map(|i| (0..4).map(move |j| (10 * i + j + 1) as f64))
|
||||
.collect();
|
||||
let mm = clawhdf5::MmapFile::open(&path).unwrap();
|
||||
assert_eq!(mm.dataset("a").unwrap().read_i64().unwrap(), want_a);
|
||||
assert_eq!(mm.dataset("b").unwrap().read_f64().unwrap(), want_b);
|
||||
let lazy = clawhdf5::LazyFile::open_mmap(&path).unwrap();
|
||||
assert_eq!(lazy.dataset("a").unwrap().read_i64().unwrap(), want_a);
|
||||
assert_eq!(lazy.dataset("b").unwrap().read_f64().unwrap(), want_b);
|
||||
let storage = File::open_storage(Arc::new(std::fs::read(&path).unwrap())).unwrap();
|
||||
check_mid_write_copy(&storage);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_copy_taken_mid_write_reads_as_h5py_swmr_reader_reads_it() {
|
||||
skip_if_no_python!();
|
||||
let path = fixture("swmr_mid_write.h5");
|
||||
// libhdf5 refuses a non-SWMR open of this file ("file is already open
|
||||
// for write"); its SWMR reader reads the values checked above.
|
||||
let script = format!(
|
||||
r#"
|
||||
import h5py, numpy as np
|
||||
with h5py.File("{p}", "r", swmr=True, locking=False) as f:
|
||||
a = f["a"][()]
|
||||
b = f["b"][()]
|
||||
assert np.array_equal(a, np.arange(148) + 1), a
|
||||
assert np.array_equal(b, (np.arange(148)[:, None] * 10 + np.arange(4) + 1).astype("f8")), b
|
||||
print("ok")
|
||||
"#,
|
||||
p = path.display()
|
||||
);
|
||||
let out = Command::new(python())
|
||||
.args(["-c", &script])
|
||||
.output()
|
||||
.unwrap();
|
||||
assert!(
|
||||
out.status.success(),
|
||||
"h5py failed:\n{}",
|
||||
String::from_utf8_lossy(&out.stderr)
|
||||
);
|
||||
check_mid_write_copy(&File::open(&path).unwrap());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn open_swmr_reads_the_mid_write_copy() {
|
||||
let f = File::open_swmr(fixture("swmr_mid_write.h5")).unwrap();
|
||||
assert!(f.is_swmr_read());
|
||||
assert_eq!(f.swmr_read_attempts(), clawhdf5::SWMR_READ_ATTEMPTS);
|
||||
// The copy still has the SWMR-write flag: as far as it says, its writer
|
||||
// is still writing.
|
||||
assert!(f.swmr_writer_active().unwrap());
|
||||
check_mid_write_copy(&f);
|
||||
let mut a = f.dataset("a").unwrap();
|
||||
a.refresh().unwrap();
|
||||
assert_eq!(a.shape().unwrap(), vec![148]);
|
||||
assert_eq!(f.swmr_retries(), 0);
|
||||
assert!(
|
||||
!File::open(fixture("swmr_mid_write.h5"))
|
||||
.unwrap()
|
||||
.is_swmr_read()
|
||||
);
|
||||
}
|
||||
|
||||
/// The mid-write copy with its superblock's flags set to `flags` (and the
|
||||
/// superblock checksum updated).
|
||||
fn mid_write_copy_with_flags(flags: u8) -> Vec<u8> {
|
||||
let mut bytes = std::fs::read(fixture("swmr_mid_write.h5")).unwrap();
|
||||
// Superblock v3, 8-byte offsets: 12 bytes, 4 addresses, checksum.
|
||||
bytes[11] = flags;
|
||||
let sum = clawhdf5_format::checksum::jenkins_lookup3(&bytes[..44]);
|
||||
bytes[44..48].copy_from_slice(&sum.to_le_bytes());
|
||||
bytes
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn open_swmr_refuses_a_file_open_for_writing_without_swmr() {
|
||||
// libhdf5 refuses it too: "file is already open for write".
|
||||
let bytes = mid_write_copy_with_flags(0x01);
|
||||
let err = File::open_storage_swmr(Arc::new(bytes)).unwrap_err();
|
||||
assert!(matches!(err, clawhdf5::Error::Locked(_)), "{err}");
|
||||
// Closed (flags clear): opens, and no writer is active.
|
||||
let f = File::open_storage_swmr(Arc::new(mid_write_copy_with_flags(0))).unwrap();
|
||||
assert!(!f.swmr_writer_active().unwrap());
|
||||
}
|
||||
|
||||
/// The outcome of reading every dataset of `f` in full, as text: the
|
||||
/// values, or the error.
|
||||
fn read_all(f: &File) -> Vec<String> {
|
||||
["a", "b"]
|
||||
.iter()
|
||||
.map(|n| match f.dataset(n).and_then(|d| d.read_f64()) {
|
||||
Ok(v) => format!("{n}: {} values", v.len()),
|
||||
Err(e) => format!("{n}: error {e}"),
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn open_swmr_reads_a_file_without_the_swmr_write_flag_as_file_open_does() {
|
||||
// The mid-write copy with its flags cleared: a closed file that
|
||||
// records an end of file of 715 in 6 030 bytes, its chunk indexes past
|
||||
// that end. File::open (and libhdf5's plain reader) refuse what lies
|
||||
// past the recorded end; open_swmr must too, and must not retry (the
|
||||
// file is not live). See the h5py test below for libhdf5's SWMR reader.
|
||||
let bytes = mid_write_copy_with_flags(0);
|
||||
let dir = tempfile::tempdir_in(env!("CARGO_TARGET_TMPDIR")).unwrap();
|
||||
let path = dir.path().join("closed_short_eof.h5");
|
||||
std::fs::write(&path, &bytes).unwrap();
|
||||
|
||||
let plain = File::open(&path).unwrap();
|
||||
let swmr = File::open_swmr(&path).unwrap();
|
||||
let swmr_storage = File::open_storage_swmr(Arc::new(bytes)).unwrap();
|
||||
let want = read_all(&plain);
|
||||
assert!(
|
||||
want.iter().all(|r| r.contains("error")),
|
||||
"File::open reads past the recorded end: {want:?}"
|
||||
);
|
||||
for f in [&swmr, &swmr_storage] {
|
||||
assert!(!f.is_swmr_read());
|
||||
assert!(!f.swmr_writer_active().unwrap());
|
||||
let started = std::time::Instant::now();
|
||||
assert_eq!(read_all(f), want);
|
||||
assert!(started.elapsed() < std::time::Duration::from_millis(100));
|
||||
assert_eq!(f.swmr_retries(), 0);
|
||||
}
|
||||
|
||||
// The same bytes with the SWMR-write flag (and write access) set are
|
||||
// read live, past the recorded end.
|
||||
let live = File::open_storage_swmr(Arc::new(mid_write_copy_with_flags(0x05))).unwrap();
|
||||
assert!(live.is_swmr_read());
|
||||
check_mid_write_copy(&live);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn how_h5py_reads_the_closed_copy_past_its_recorded_end_of_file() {
|
||||
skip_if_no_python!();
|
||||
let dir = tempfile::tempdir_in(env!("CARGO_TARGET_TMPDIR")).unwrap();
|
||||
let path = dir.path().join("closed_short_eof.h5");
|
||||
std::fs::write(&path, mid_write_copy_with_flags(0)).unwrap();
|
||||
// libhdf5's plain reader refuses `a` (its chunk index is past the
|
||||
// recorded end of file), as File::open and open_swmr do above.
|
||||
// libhdf5's SWMR reader skips its end-of-allocation check for every
|
||||
// file it opens (H5FD_read), and reads it; it still refuses an object
|
||||
// *header* past the end ("address of object past end of allocation",
|
||||
// H5O_protect). open_swmr does not copy that half-way rule for files
|
||||
// without the SWMR-write flag: they read as File::open reads them.
|
||||
let script = format!(
|
||||
r#"
|
||||
import h5py
|
||||
for swmr in (False, True):
|
||||
try:
|
||||
with h5py.File("{p}", "r", swmr=swmr, locking=False) as f:
|
||||
f["a"][()]
|
||||
except Exception as e:
|
||||
print("refused", swmr)
|
||||
else:
|
||||
print("read", swmr)
|
||||
"#,
|
||||
p = path.display()
|
||||
);
|
||||
let out = Command::new(python())
|
||||
.args(["-c", &script])
|
||||
.output()
|
||||
.unwrap();
|
||||
let stdout = String::from_utf8_lossy(&out.stdout);
|
||||
assert_eq!(
|
||||
stdout.split_whitespace().collect::<Vec<_>>(),
|
||||
["refused", "False", "read", "True"],
|
||||
"{}",
|
||||
String::from_utf8_lossy(&out.stderr)
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn open_swmr_returns_a_permanent_error_at_once() {
|
||||
// Not HDF5 at all: SignatureNotFound is not something a writer
|
||||
// causes, so it is not retried (100 attempts would pause about 0.9 s).
|
||||
let started = std::time::Instant::now();
|
||||
let err = File::open_storage_swmr(Arc::new(vec![7u8; 4096])).unwrap_err();
|
||||
assert!(
|
||||
matches!(
|
||||
err,
|
||||
clawhdf5::Error::Format(clawhdf5_format::error::FormatError::SignatureNotFound)
|
||||
),
|
||||
"{err}"
|
||||
);
|
||||
assert!(started.elapsed() < std::time::Duration::from_millis(100));
|
||||
|
||||
// A live file: a lookup of a name it does not have fails at once, and
|
||||
// a torn object header read (below) is retried.
|
||||
let f = File::open_storage_swmr(Arc::new(mid_write_copy_with_flags(0x05))).unwrap();
|
||||
assert!(f.is_swmr_read());
|
||||
let started = std::time::Instant::now();
|
||||
assert!(f.dataset("no_such").is_err());
|
||||
assert!(started.elapsed() < std::time::Duration::from_millis(100));
|
||||
assert_eq!(f.swmr_retries(), 0);
|
||||
}
|
||||
|
||||
/// A storage whose next `torn` reads come back garbled, as a read racing a
|
||||
/// rewrite of the structure can see them: the middle byte changed, or with
|
||||
/// `invert` every byte (signatures included).
|
||||
struct Torn {
|
||||
bytes: Vec<u8>,
|
||||
torn: AtomicU32,
|
||||
invert: AtomicBool,
|
||||
}
|
||||
|
||||
impl clawhdf5::Storage for Torn {
|
||||
fn read_at(
|
||||
&self,
|
||||
offset: u64,
|
||||
len: usize,
|
||||
) -> Result<std::borrow::Cow<'_, [u8]>, clawhdf5_format::error::FormatError> {
|
||||
let got = self.bytes.as_slice().read_at(offset, len)?;
|
||||
let tear = self
|
||||
.torn
|
||||
.fetch_update(Ordering::SeqCst, Ordering::SeqCst, |n| n.checked_sub(1))
|
||||
.is_ok();
|
||||
if tear && !got.is_empty() {
|
||||
let mut v = got.into_owned();
|
||||
if self.invert.load(Ordering::SeqCst) {
|
||||
v.iter_mut().for_each(|b| *b = !*b);
|
||||
} else {
|
||||
let mid = v.len() / 2;
|
||||
v[mid] ^= 0x5a;
|
||||
}
|
||||
return Ok(v.into());
|
||||
}
|
||||
Ok(got)
|
||||
}
|
||||
|
||||
fn len(&self) -> u64 {
|
||||
self.bytes.len() as u64
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_torn_metadata_read_is_retried_and_never_returned() {
|
||||
use clawhdf5::Storage as _;
|
||||
let torn = Arc::new(Torn {
|
||||
bytes: std::fs::read(fixture("swmr_mid_write.h5")).unwrap(),
|
||||
torn: AtomicU32::new(0),
|
||||
invert: AtomicBool::new(false),
|
||||
});
|
||||
let f = File::open_storage_swmr(torn.clone()).unwrap();
|
||||
assert_eq!(f.swmr_retries(), 0);
|
||||
assert!(torn.len() > 0);
|
||||
|
||||
// The next 3 reads (object headers) come back garbled: the lookup
|
||||
// fails its checks, is run again, and returns the right dataset. The
|
||||
// same with every byte garbled, signatures included.
|
||||
for invert in [false, true] {
|
||||
torn.invert.store(invert, Ordering::SeqCst);
|
||||
let before = f.swmr_retries();
|
||||
torn.torn.store(3, Ordering::SeqCst);
|
||||
let a = f.dataset("a").unwrap();
|
||||
assert!(f.swmr_retries() > before, "invert {invert}");
|
||||
assert_eq!(a.read_i64().unwrap(), (1..=148).collect::<Vec<i64>>());
|
||||
// Two garbled reads: the prefix of the root group's header (whose
|
||||
// garbled byte may go unused, as it is read again whole) and the
|
||||
// whole header, which fails its checksum.
|
||||
let before = f.swmr_retries();
|
||||
torn.torn.store(2, Ordering::SeqCst);
|
||||
let b = f.dataset("b").unwrap();
|
||||
assert_eq!(b.read_f64().unwrap().len(), 148 * 4);
|
||||
assert!(f.swmr_retries() > before, "invert {invert}");
|
||||
}
|
||||
torn.invert.store(false, Ordering::SeqCst);
|
||||
|
||||
// With one attempt the error is returned instead (both reads of the
|
||||
// header garbled, as above).
|
||||
let mut f = File::open_storage_swmr(torn.clone()).unwrap();
|
||||
f.set_swmr_read_attempts(1);
|
||||
torn.torn.store(2, Ordering::SeqCst);
|
||||
assert!(f.dataset("a").is_err());
|
||||
assert_eq!(f.swmr_retries(), 0);
|
||||
torn.torn.store(0, Ordering::SeqCst);
|
||||
|
||||
// A refresh that keeps failing gives up after the attempts, and the
|
||||
// handle keeps its extent.
|
||||
f.set_swmr_read_attempts(5);
|
||||
let mut b = f.dataset("b").unwrap();
|
||||
torn.torn.store(1000, Ordering::SeqCst);
|
||||
assert!(b.refresh().is_err());
|
||||
assert_eq!(f.swmr_retries(), 4);
|
||||
torn.torn.store(0, Ordering::SeqCst);
|
||||
assert_eq!(b.shape().unwrap(), vec![148, 4]);
|
||||
b.refresh().unwrap();
|
||||
assert_eq!(b.read_f64().unwrap().len(), 148 * 4);
|
||||
|
||||
// A file not opened for SWMR reading does not retry.
|
||||
let plain = File::open_storage(torn.clone()).unwrap();
|
||||
torn.torn.store(2, Ordering::SeqCst);
|
||||
assert!(plain.dataset("a").is_err());
|
||||
assert_eq!(plain.swmr_retries(), 0);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// A live file: an h5py SWMR writer appends while clawhdf5 and h5py read.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// The h5py SWMR writer. It creates three chunked datasets whose values are
|
||||
/// a function of their position, so a reader can check every value it
|
||||
/// reads without knowing when it was written:
|
||||
///
|
||||
/// - `a`: int64 `(n,)`, chunks of 100, no filter — `a[i] = i + 1`
|
||||
/// (Extensible Array index);
|
||||
/// - `b`: float64 `(n, 4)`, chunks of 16 x 4, gzip — `b[i, j] = 10 i + j + 1`
|
||||
/// (Extensible Array index);
|
||||
/// - `c`: int32 `(r, k)`, both unlimited, chunks of 8 x 8, no filter —
|
||||
/// `c[i, j] = 1000 i + j + 1` (version-2 B-tree index).
|
||||
///
|
||||
/// It switches to SWMR mode, prints `ready`, then for `steps` steps appends a
|
||||
/// random number of rows to each dataset (and every 7th step a column to
|
||||
/// `c`), resizing before writing as SWMR writers must, flushing each
|
||||
/// dataset and pausing 2 ms, and finally closes the file and writes the
|
||||
/// final contents to `<path>.<name>.bin`. `CLAWHDF5_SWMR_STEPS` sets the
|
||||
/// number of steps (default 2500).
|
||||
const WRITER: &str = r#"
|
||||
import sys, time, random, h5py, numpy as np
|
||||
path, steps = sys.argv[1], int(sys.argv[2])
|
||||
rng = random.Random(1234)
|
||||
f = h5py.File(path, "w", libver="latest")
|
||||
a = f.create_dataset("a", shape=(0,), maxshape=(None,), chunks=(100,), dtype="i8")
|
||||
b = f.create_dataset("b", shape=(0, 4), maxshape=(None, 4), chunks=(16, 4), dtype="f8",
|
||||
compression="gzip")
|
||||
c = f.create_dataset("c", shape=(0, 1), maxshape=(None, None), chunks=(8, 8), dtype="i4")
|
||||
f.swmr_mode = True
|
||||
print("ready", flush=True)
|
||||
na = nb = rc = 0
|
||||
kc = 1
|
||||
for step in range(steps):
|
||||
k = rng.randint(1, 60)
|
||||
a.resize((na + k,)); a[na:na + k] = np.arange(na, na + k) + 1; na += k
|
||||
a.flush()
|
||||
k = rng.randint(1, 20)
|
||||
rows = np.arange(nb, nb + k)[:, None] * 10 + np.arange(4) + 1
|
||||
b.resize((nb + k, 4)); b[nb:nb + k] = rows; nb += k
|
||||
b.flush()
|
||||
if step % 7 == 6:
|
||||
c.resize((rc, kc + 1))
|
||||
if rc:
|
||||
c[:, kc] = np.arange(rc) * 1000 + kc + 1
|
||||
kc += 1
|
||||
k = rng.randint(1, 3)
|
||||
c.resize((rc + k, kc))
|
||||
c[rc:rc + k] = np.arange(rc, rc + k)[:, None] * 1000 + np.arange(kc) + 1
|
||||
rc += k
|
||||
c.flush()
|
||||
time.sleep(0.002)
|
||||
f.close()
|
||||
with h5py.File(path, "r") as f:
|
||||
for name in "abc":
|
||||
open(f"{path}.{name}.bin", "wb").write(f[name][()].tobytes())
|
||||
print(name, *f[name].shape, flush=True)
|
||||
"#;
|
||||
|
||||
/// h5py's own SWMR reader, run beside ours as the reference: the same
|
||||
/// checks, until `<path>.stop` exists. Prints its iteration count.
|
||||
const H5PY_READER: &str = r#"
|
||||
import os, sys, h5py, numpy as np
|
||||
path = sys.argv[1]
|
||||
f = h5py.File(path, "r", swmr=True)
|
||||
ds = {n: f[n] for n in "abc"}
|
||||
last = {n: (0,) * ds[n].ndim for n in "abc"}
|
||||
its = 0
|
||||
while True:
|
||||
stop = os.path.exists(path + ".stop")
|
||||
for n, d in ds.items():
|
||||
d.refresh()
|
||||
shape = d.shape
|
||||
assert all(s >= l for s, l in zip(shape, last[n])), (n, shape, last[n])
|
||||
last[n] = shape
|
||||
v = d[()]
|
||||
if n == "a":
|
||||
want = np.arange(shape[0]) + 1
|
||||
elif n == "b":
|
||||
want = np.arange(shape[0])[:, None] * 10 + np.arange(4) + 1
|
||||
else:
|
||||
want = np.arange(shape[0])[:, None] * 1000 + np.arange(shape[1]) + 1
|
||||
bad = np.argwhere(v != want)
|
||||
assert len(bad) == 0, (n, shape, bad[:5], v[tuple(bad[0])], want[tuple(bad[0])])
|
||||
its += 1
|
||||
if stop:
|
||||
break
|
||||
print("iterations", its, *last["a"], *last["b"], *last["c"], flush=True)
|
||||
"#;
|
||||
|
||||
fn want_a(n: u64) -> Vec<i64> {
|
||||
(1..=n as i64).collect()
|
||||
}
|
||||
|
||||
fn want_b(rows: std::ops::Range<u64>) -> Vec<f64> {
|
||||
rows.flat_map(|i| (0..4).map(move |j| (10 * i + j + 1) as f64))
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn want_c(rows: std::ops::Range<u64>, cols: u64) -> Vec<i32> {
|
||||
rows.flat_map(|i| (0..cols).map(move |j| (1000 * i + j + 1) as i32))
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// One pass of the Rust reader: refresh every dataset, check its extent
|
||||
/// did not shrink, and check what it reads (the last rows of each, and all
|
||||
/// of them every `full_every` passes).
|
||||
fn check_pass(
|
||||
datasets: &mut [clawhdf5::Dataset<'_>; 3],
|
||||
last: &mut [Vec<u64>; 3],
|
||||
pass: u64,
|
||||
full_every: u64,
|
||||
) {
|
||||
for (k, d) in datasets.iter_mut().enumerate() {
|
||||
d.refresh().unwrap_or_else(|e| panic!("refresh {k}: {e}"));
|
||||
let shape = d.shape().unwrap();
|
||||
assert!(
|
||||
shape.iter().zip(&last[k]).all(|(s, l)| s >= l),
|
||||
"dataset {k} shrank: {shape:?} after {:?}",
|
||||
last[k]
|
||||
);
|
||||
last[k] = shape;
|
||||
}
|
||||
let full = pass.is_multiple_of(full_every);
|
||||
let na = last[0][0];
|
||||
let from = if full { 0 } else { na.saturating_sub(200) };
|
||||
let got = if full {
|
||||
datasets[0].read_i64()
|
||||
} else {
|
||||
datasets[0].read_i64_selection(&Selection::slice(std::slice::from_ref(&(from..na))))
|
||||
}
|
||||
.unwrap_or_else(|e| panic!("read a: {e}"));
|
||||
let want: Vec<i64> = want_a(na).split_off(from as usize);
|
||||
assert!(got == want, "a {from}..{na}: {:?}", first_diff(&got, &want));
|
||||
|
||||
let nb = last[1][0];
|
||||
let from = if full { 0 } else { nb.saturating_sub(50) };
|
||||
let sel = Selection::slice(&[from..nb, 0..4]);
|
||||
let got = datasets[1]
|
||||
.read_f64_selection(&sel)
|
||||
.unwrap_or_else(|e| panic!("read b: {e}"));
|
||||
let want = want_b(from..nb);
|
||||
assert!(got == want, "b {from}..{nb}: {:?}", first_diff(&got, &want));
|
||||
|
||||
let (rc, kc) = (last[2][0], last[2][1]);
|
||||
let from = if full { 0 } else { rc.saturating_sub(9) };
|
||||
let got = if full {
|
||||
datasets[2].read_i32()
|
||||
} else {
|
||||
datasets[2].read_i32_selection(&Selection::slice(&[from..rc, 0..kc]))
|
||||
}
|
||||
.unwrap_or_else(|e| panic!("read c: {e}"));
|
||||
let want = want_c(from..rc, kc);
|
||||
assert!(
|
||||
got == want,
|
||||
"c {from}..{rc} x {kc}: {:?}",
|
||||
first_diff(&got, &want)
|
||||
);
|
||||
}
|
||||
|
||||
fn first_diff<T: PartialEq + std::fmt::Debug>(got: &[T], want: &[T]) -> String {
|
||||
if got.len() != want.len() {
|
||||
return format!("{} values, want {}", got.len(), want.len());
|
||||
}
|
||||
let i = got.iter().zip(want).position(|(g, w)| g != w).unwrap();
|
||||
format!("first difference at {i}: {:?}, want {:?}", got[i], want[i])
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_live_file_reads_consistently_while_an_h5py_swmr_writer_appends() {
|
||||
skip_if_no_python!();
|
||||
let steps: u32 = std::env::var("CLAWHDF5_SWMR_STEPS")
|
||||
.ok()
|
||||
.and_then(|s| s.parse().ok())
|
||||
.unwrap_or(2500);
|
||||
let dir = tempfile::tempdir_in(env!("CARGO_TARGET_TMPDIR")).unwrap();
|
||||
let path = dir.path().join("live.h5");
|
||||
|
||||
let mut writer = Command::new(python())
|
||||
.args(["-c", WRITER, path.to_str().unwrap(), &steps.to_string()])
|
||||
.stdout(Stdio::piped())
|
||||
.spawn()
|
||||
.unwrap();
|
||||
let mut out = BufReader::new(writer.stdout.take().unwrap());
|
||||
let mut line = String::new();
|
||||
out.read_line(&mut line).unwrap();
|
||||
assert_eq!(line.trim(), "ready", "writer did not start");
|
||||
|
||||
// h5py's SWMR reader on the same file, as the reference.
|
||||
let h5py_reader = Command::new(python())
|
||||
.args(["-c", H5PY_READER, path.to_str().unwrap()])
|
||||
.stdout(Stdio::piped())
|
||||
.stderr(Stdio::piped())
|
||||
.spawn()
|
||||
.unwrap();
|
||||
|
||||
let file = File::open_swmr(&path).unwrap();
|
||||
assert!(file.is_swmr_read());
|
||||
assert!(file.swmr_writer_active().unwrap());
|
||||
// Two reader threads share the file, each with its own handles; each
|
||||
// makes passes until it has seen the writer exit, then one more.
|
||||
let writer_done = AtomicBool::new(false);
|
||||
let reader = || {
|
||||
let mut datasets = ["a", "b", "c"].map(|n| file.dataset(n).unwrap());
|
||||
let mut last = [vec![0], vec![0, 4], vec![0, 1]];
|
||||
let (mut passes, mut grew) = (0u64, 0u64);
|
||||
loop {
|
||||
let done = writer_done.load(Ordering::Acquire);
|
||||
let before = last[0][0];
|
||||
check_pass(&mut datasets, &mut last, passes, 50);
|
||||
passes += 1;
|
||||
grew += u64::from(last[0][0] > before);
|
||||
if done {
|
||||
break;
|
||||
}
|
||||
}
|
||||
// The writer has closed the file: one more pass reads the final
|
||||
// extents, all of them.
|
||||
check_pass(&mut datasets, &mut last, 0, 50);
|
||||
(passes, grew, last)
|
||||
};
|
||||
let (status, results) = std::thread::scope(|s| {
|
||||
let readers = [s.spawn(reader), s.spawn(reader)];
|
||||
let status = writer.wait().unwrap();
|
||||
writer_done.store(true, Ordering::Release);
|
||||
(status, readers.map(|r| r.join().unwrap()))
|
||||
});
|
||||
assert!(status.success(), "writer failed");
|
||||
assert!(!file.swmr_writer_active().unwrap());
|
||||
let last = results[0].2.clone();
|
||||
assert_eq!(results[1].2, last);
|
||||
let passes: Vec<u64> = results.iter().map(|r| r.0).collect();
|
||||
let grew: u64 = results.iter().map(|r| r.1).min().unwrap();
|
||||
|
||||
std::fs::write(path.with_extension("h5.stop"), b"").unwrap();
|
||||
let reference = h5py_reader.wait_with_output().unwrap();
|
||||
assert!(
|
||||
reference.status.success(),
|
||||
"h5py's SWMR reader failed:\n{}",
|
||||
String::from_utf8_lossy(&reference.stderr)
|
||||
);
|
||||
let reference = String::from_utf8_lossy(&reference.stdout).into_owned();
|
||||
|
||||
// The writer's final contents, as h5py reads the closed file.
|
||||
let mut finals = String::new();
|
||||
out.read_to_string(&mut finals).unwrap();
|
||||
let shapes: std::collections::HashMap<&str, Vec<u64>> = finals
|
||||
.lines()
|
||||
.map(|l| {
|
||||
let mut w = l.split_whitespace();
|
||||
let name = w.next().unwrap();
|
||||
(name, w.map(|v| v.parse().unwrap()).collect())
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(shapes["a"], last[0]);
|
||||
assert_eq!(shapes["b"], last[1]);
|
||||
assert_eq!(shapes["c"], last[2]);
|
||||
let bin = |name: &str| std::fs::read(format!("{}.{name}.bin", path.display())).unwrap();
|
||||
let le = |b: &[u8], n: usize| -> Vec<[u8; 8]> {
|
||||
b.chunks(n)
|
||||
.map(|c| {
|
||||
let mut v = [0u8; 8];
|
||||
v[..n].copy_from_slice(c);
|
||||
v
|
||||
})
|
||||
.collect()
|
||||
};
|
||||
// The live handle and a fresh non-SWMR open of the closed file both
|
||||
// read exactly what h5py reads.
|
||||
for f in [&file, &File::open(&path).unwrap()] {
|
||||
let a: Vec<[u8; 8]> = f
|
||||
.dataset("a")
|
||||
.unwrap()
|
||||
.read_i64()
|
||||
.unwrap()
|
||||
.iter()
|
||||
.map(|v| v.to_le_bytes())
|
||||
.collect();
|
||||
assert_eq!(a, le(&bin("a"), 8));
|
||||
let b: Vec<[u8; 8]> = f
|
||||
.dataset("b")
|
||||
.unwrap()
|
||||
.read_f64()
|
||||
.unwrap()
|
||||
.iter()
|
||||
.map(|v| v.to_le_bytes())
|
||||
.collect();
|
||||
assert_eq!(b, le(&bin("b"), 8));
|
||||
let c: Vec<[u8; 8]> = f
|
||||
.dataset("c")
|
||||
.unwrap()
|
||||
.read_i32()
|
||||
.unwrap()
|
||||
.iter()
|
||||
.map(|v| {
|
||||
let mut x = [0u8; 8];
|
||||
x[..4].copy_from_slice(&v.to_le_bytes());
|
||||
x
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(c, le(&bin("c"), 4));
|
||||
}
|
||||
eprintln!(
|
||||
"SWMR: {steps} writer steps; clawhdf5 readers made {passes:?} passes (each saw `a` \
|
||||
grow at least {grew} times), {} retries; h5py reader: {}",
|
||||
file.swmr_retries(),
|
||||
reference.trim()
|
||||
);
|
||||
assert!(
|
||||
grew >= 2,
|
||||
"a reader never saw the file grow ({passes:?} passes)"
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Every read path retries: strings, variable-length data, attributes.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// A storage over `bytes` whose read number `fail_at` (counted from the
|
||||
/// last [`Flaky::arm`]) fails with a checksum mismatch, as a read racing a
|
||||
/// SWMR writer's flush can; it counts the reads.
|
||||
struct Flaky {
|
||||
bytes: Vec<u8>,
|
||||
reads: AtomicU32,
|
||||
fail_at: AtomicU32,
|
||||
}
|
||||
|
||||
impl Flaky {
|
||||
fn arm(&self, fail_at: u32) {
|
||||
self.reads.store(0, Ordering::SeqCst);
|
||||
self.fail_at.store(fail_at, Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
|
||||
impl clawhdf5::Storage for Flaky {
|
||||
fn read_at(
|
||||
&self,
|
||||
offset: u64,
|
||||
len: usize,
|
||||
) -> Result<std::borrow::Cow<'_, [u8]>, clawhdf5_format::error::FormatError> {
|
||||
let n = self.reads.fetch_add(1, Ordering::SeqCst);
|
||||
if n == self.fail_at.load(Ordering::SeqCst) {
|
||||
return Err(clawhdf5_format::error::FormatError::ChecksumMismatch {
|
||||
expected: 0,
|
||||
computed: 1,
|
||||
});
|
||||
}
|
||||
self.bytes.as_slice().read_at(offset, len)
|
||||
}
|
||||
|
||||
fn len(&self) -> u64 {
|
||||
self.bytes.len() as u64
|
||||
}
|
||||
}
|
||||
|
||||
/// `swmr_strings_attrs.h5`: a copy h5py 3.16 (HDF5 2.0) made of its own
|
||||
/// file right after switching it to SWMR mode (the SWMR-write flag is set),
|
||||
/// holding a variable-length string dataset `s` (`alpha beta gamma
|
||||
/// delta`), a variable-length int32 dataset `v` (`[1] [2 3] [4 5 6]`), an
|
||||
/// int64 dataset `d` (`0..10`) with 12 string attributes `k00`..`k11`
|
||||
/// (`value 0`..; dense storage: a fractal heap and a v2 B-tree), a group
|
||||
/// `g` with attribute `note`, and root attributes `title` (a
|
||||
/// variable-length string) and `n`.
|
||||
#[test]
|
||||
fn every_read_path_of_a_live_file_retries_a_failed_read() {
|
||||
let flaky = Arc::new(Flaky {
|
||||
bytes: std::fs::read(fixture("swmr_strings_attrs.h5")).unwrap(),
|
||||
reads: AtomicU32::new(0),
|
||||
fail_at: AtomicU32::new(u32::MAX),
|
||||
});
|
||||
let f = File::open_storage_swmr(flaky.clone()).unwrap();
|
||||
assert!(f.is_swmr_read());
|
||||
let root = f.root();
|
||||
let g = f.group("g").unwrap();
|
||||
let s = f.dataset("s").unwrap();
|
||||
let v = f.dataset("v").unwrap();
|
||||
let d = f.dataset("d").unwrap();
|
||||
let sorted = |m: std::collections::HashMap<String, clawhdf5::AttrValue>| {
|
||||
format!(
|
||||
"{:?}",
|
||||
m.into_iter().collect::<std::collections::BTreeMap<_, _>>()
|
||||
)
|
||||
};
|
||||
type Op<'a> = Box<dyn Fn() -> Result<String, clawhdf5::Error> + 'a>;
|
||||
let ops: Vec<(&str, Op<'_>)> = vec![
|
||||
("root.attrs", Box::new(|| root.attrs().map(sorted))),
|
||||
(
|
||||
"root.attr",
|
||||
Box::new(|| root.attr("title").map(|a| format!("{a:?}"))),
|
||||
),
|
||||
("g.attrs", Box::new(|| g.attrs().map(sorted))),
|
||||
("d.attrs", Box::new(|| d.attrs().map(sorted))),
|
||||
(
|
||||
"d.attr",
|
||||
Box::new(|| d.attr("k05").map(|a| format!("{a:?}"))),
|
||||
),
|
||||
(
|
||||
"root.datasets",
|
||||
Box::new(|| root.datasets().map(|n| format!("{n:?}"))),
|
||||
),
|
||||
(
|
||||
"root.groups",
|
||||
Box::new(|| root.groups().map(|n| format!("{n:?}"))),
|
||||
),
|
||||
("d.shape", Box::new(|| d.shape().map(|n| format!("{n:?}")))),
|
||||
(
|
||||
"d.read_i64",
|
||||
Box::new(|| d.read_i64().map(|n| format!("{n:?}"))),
|
||||
),
|
||||
(
|
||||
"s.read_string",
|
||||
Box::new(|| s.read_string().map(|n| format!("{n:?}"))),
|
||||
),
|
||||
(
|
||||
"s.read_string_bytes",
|
||||
Box::new(|| s.read_string_bytes().map(|n| format!("{n:?}"))),
|
||||
),
|
||||
(
|
||||
"s.read_string_selection",
|
||||
Box::new(|| {
|
||||
s.read_string_selection(&Selection::slice(std::slice::from_ref(&(1..3))))
|
||||
.map(|n| format!("{n:?}"))
|
||||
}),
|
||||
),
|
||||
(
|
||||
"v.read_vlen",
|
||||
Box::new(|| v.read_vlen::<i32>().map(|n| format!("{n:?}"))),
|
||||
),
|
||||
(
|
||||
"v.read_vlen_selection",
|
||||
Box::new(|| {
|
||||
v.read_vlen_selection::<i32>(&Selection::slice(std::slice::from_ref(&(1..3))))
|
||||
.map(|n| format!("{n:?}"))
|
||||
}),
|
||||
),
|
||||
];
|
||||
for (name, op) in &ops {
|
||||
flaky.arm(u32::MAX);
|
||||
let want = op().unwrap_or_else(|e| panic!("{name}: {e}"));
|
||||
let reads = flaky.reads.load(Ordering::SeqCst);
|
||||
// Each read the operation makes fails once in turn: the operation
|
||||
// is run again and returns the same result.
|
||||
for k in 0..reads {
|
||||
flaky.arm(k);
|
||||
let before = f.swmr_retries();
|
||||
let got = op().unwrap_or_else(|e| panic!("{name}, read {k} of {reads} failing: {e}"));
|
||||
assert_eq!(got, want, "{name}, read {k} of {reads} failing");
|
||||
assert_eq!(f.swmr_retries(), before + 1, "{name}, read {k} of {reads}");
|
||||
}
|
||||
}
|
||||
flaky.arm(u32::MAX);
|
||||
assert!(ops[0].1().unwrap().contains("live strings"));
|
||||
assert!(ops[3].1().unwrap().contains("value 11"));
|
||||
assert_eq!(
|
||||
ops[9].1().unwrap(),
|
||||
r#"["alpha", "beta", "gamma", "delta"]"#
|
||||
);
|
||||
assert_eq!(ops[12].1().unwrap(), "[[1], [2, 3], [4, 5, 6]]");
|
||||
|
||||
// With one attempt the failure reaches the caller.
|
||||
let mut f = File::open_storage_swmr(flaky.clone()).unwrap();
|
||||
f.set_swmr_read_attempts(1);
|
||||
let s = f.dataset("s").unwrap();
|
||||
flaky.arm(0);
|
||||
assert!(s.read_string().is_err());
|
||||
}
|
||||
+106
-5
@@ -7,20 +7,33 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
|
||||
`File::open_storage` gives the facade's read API over any `Storage` (see
|
||||
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
|
||||
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
|
||||
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is
|
||||
object stores) and URLs in `h5rs` (see the M3 status below). M5 (SWMR) is
|
||||
done on branch `feat/p3-m5-swmr-reader`, with its own design in
|
||||
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
|
||||
next. Every count in §1–§2 was
|
||||
|
||||
object stores) and URLs in `h5rs` (see the M3 status below); the Python
|
||||
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
|
||||
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
|
||||
|
||||
change. Progress: M1, first part (the `Storage` trait and the metadata
|
||||
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
||||
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
||||
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||
|
||||
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
|
||||
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
|
||||
reader, through the restartable `NeedBytes` mode (see the M4 status
|
||||
below). M5 (SWMR) is not started.
|
||||
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||
converted code without changing the plan: object-header continuation chunks
|
||||
are followed without recursion (still one bounded `read_at` per chunk), and
|
||||
implicit chunk indexes are addressed over the maximum chunk grid (in
|
||||
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
|
||||
working on the whole file in memory; it is not part of this design. Every count below was
|
||||
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
|
||||
next to it. No timing numbers appear here on purpose: the machine was shared
|
||||
working on the whole file in memory; it is not part of this design. Every
|
||||
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
|
||||
every count in a milestone's status on the date it gives, with the
|
||||
commands given next to it. No timing numbers appear here on purpose: the machine was shared
|
||||
with other build jobs when this was written.
|
||||
|
||||
## The problem
|
||||
@@ -475,7 +488,7 @@ fast path within benchmark noise.
|
||||
a request counter exposed for tests and users.
|
||||
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
|
||||
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
|
||||
Python bindings, with these choices:
|
||||
Python bindings (done 2026-09-27, below), with these choices:
|
||||
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
|
||||
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
|
||||
sits below the facade.
|
||||
@@ -512,6 +525,17 @@ fast path within benchmark noise.
|
||||
same work without a cache: 141 936 requests. Per file: A lists in 2
|
||||
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
|
||||
6.4 MB: 35 001 object headers spread over the file).
|
||||
- *Status 2026-09-27, Python bindings:* done on branch
|
||||
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
|
||||
`File.open_url(url, **options)` (cache and HTTP options) go through
|
||||
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
|
||||
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
|
||||
parsing (path lookups, headers, attributes, listings, the global heap)
|
||||
moved from `File::as_bytes` to `File::storage()` and the `*_in`
|
||||
functions, and every file access, metadata included, runs with the GIL
|
||||
released. Checked by running the whole read-vs-h5py suite over an
|
||||
in-process range server (1 MiB and 1 KiB blocks) and by request counts
|
||||
in `crates/clawhdf5-py/tests/test_remote.py`.
|
||||
|
||||
**M4 — wasm lazy loading (1–2 weeks).**
|
||||
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
|
||||
@@ -519,10 +543,87 @@ fast path within benchmark noise.
|
||||
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
|
||||
to a whole download when the server does not answer 206.
|
||||
- `examples/wasm-viewer`: open by URL.
|
||||
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
|
||||
with these choices:
|
||||
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
|
||||
`Storage` over the blocks fetched so far. A call (open, list, read)
|
||||
runs as a pass; a read that misses records its blocks and fails with
|
||||
a storage error. The error itself is not the signal: parsers catch
|
||||
errors and carry on (a listing leaves out a link it cannot resolve),
|
||||
so *any* pass that recorded a miss is thrown away, whatever it
|
||||
returned, and re-run once the missing ranges have been fetched
|
||||
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
|
||||
synchronous XHR would have kept the parsers' single pass, but it
|
||||
needs the page to run the reader off the main thread (a second
|
||||
module, messages for every call and every typed array copied back),
|
||||
since synchronous XHR can return binary data only there; the
|
||||
restartable loop costs a
|
||||
re-parse per wave of misses instead, which is CPU, not network.
|
||||
- **Progress:** no block is evicted while a call is in flight
|
||||
(`LazyStorage::operation`), so each pass that does not finish asks for
|
||||
at least one new block, and a call ends after at most one pass per
|
||||
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
|
||||
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
|
||||
evicted before metadata. The price: a call holds everything it reads
|
||||
until it finishes.
|
||||
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
|
||||
backend by blocking, evicts during a read, and does not keep reads
|
||||
that miss more than half its budget; a restartable pass needs every
|
||||
block it has read to be there when it is re-run. The block
|
||||
arithmetic and coalescing (runs of consecutive blocks, a one-block
|
||||
hole filled, at most 8 MiB per request) follow it; the cache is
|
||||
about 450 lines (with its documentation) in the wasm crate, with no
|
||||
in-flight tracking or HTTP. (A hole already cached is not filled:
|
||||
it would be fetched again.)
|
||||
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
|
||||
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
|
||||
requests at a time, every answer checked (206, `Content-Range` when
|
||||
visible, body length, the ETag/Last-Modified and length of the
|
||||
first answer). A `200` to the first request is kept as the whole
|
||||
file up to `maxDownload` (512 MiB), or refused with
|
||||
`fallback: "error"`. The first block comes with the request that
|
||||
learns the length, as in M3.
|
||||
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
|
||||
--test lazy` compares what the viewer shows of each file (kinds,
|
||||
listings, attributes, info, whole reads, hyperslabs) read lazily at
|
||||
512 B to 1 MiB blocks with the in-memory reader and the facade's
|
||||
range-storage path, for built files, the h5py/netCDF4 fixture and,
|
||||
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
|
||||
files. The built package, under Node and headless Chromium
|
||||
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
|
||||
`test/serve.py`, which counts requests): every fixture check over
|
||||
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
|
||||
attributes and a 10-value window of the 25-million-value dataset
|
||||
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
|
||||
with `open(bytes)`.
|
||||
- After review (2026-09-27): a call may fetch at most `maxFetch`
|
||||
(512 MiB, at most 1 GiB) and a longer single read is refused before
|
||||
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
|
||||
more is refused at open on wasm32 (offsets become `usize`); every
|
||||
response body is cut off at the length asked for. Parsers that walk
|
||||
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
|
||||
the wasm listing's child headers) read the siblings after the first
|
||||
failure before returning it, so a pass asks for a whole level's
|
||||
missing blocks: listing 3000 datasets went from 185 passes to 6.
|
||||
The 32-bit risk below is covered by a Node test that reads data at
|
||||
3 GiB from a mock server and is refused a 4 GiB file.
|
||||
|
||||
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
||||
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
||||
blocks past the old end. Needs libhdf5 SWMR semantics research first.
|
||||
- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader`; design and
|
||||
libhdf5 research in [`swmr.md`](swmr.md). Differences from the sketch
|
||||
above: the refresh is per dataset (`Dataset::refresh`, as libhdf5's
|
||||
`H5Drefresh`), not per file — a SWMR writer only grows datasets, and the
|
||||
superblock's EOF is not kept up to date by it, so there is nothing to
|
||||
re-read there (`File::swmr_writer_active` re-reads its flags). A live
|
||||
file (`File::open_swmr`) reads through `FileStorage` (positioned reads,
|
||||
`len()` the current length) with no block or chunk cache, rather than
|
||||
invalidating cached blocks: `BlockCache`/`HttpStorage` stay snapshot
|
||||
readers. Operations that fail with an error a racing write can cause are
|
||||
retried up to 100 times (libhdf5's default for SWMR readers).
|
||||
Tested against a live h5py writer and h5py's SWMR reader
|
||||
(`crates/clawhdf5/tests/swmr_interop.rs`).
|
||||
|
||||
Total: roughly 6–10 engineer-weeks for M0–M4 (estimate, not measured).
|
||||
|
||||
|
||||
@@ -0,0 +1,178 @@
|
||||
# Design: reading files a SWMR writer is still appending to (range-read M5)
|
||||
|
||||
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
|
||||
(see "Status" at the end). This is milestone M5 of
|
||||
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
|
||||
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
|
||||
|
||||
## What libhdf5 does
|
||||
|
||||
A SWMR ("single writer, multiple readers") writer is a libhdf5 process that
|
||||
opened a file with `libver='latest'` and switched to SWMR mode (h5py
|
||||
`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file
|
||||
with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the
|
||||
writer keeps appending. What the format and the library guarantee:
|
||||
|
||||
- **Superblock v3, flags set.** The writer sets the superblock's
|
||||
file-consistency flags to write access + SWMR write (`0x05`) and clears
|
||||
them on close. The superblock's end-of-file address is *not* kept up to
|
||||
date while writing: a copy of a file taken mid-write records an EOF of a
|
||||
few hundred bytes while the file is tens of kilobytes (checked on tank,
|
||||
2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR
|
||||
reader therefore skips libhdf5's end-of-allocation check for every read
|
||||
(`H5FD_read`: "allow access to data past the end of the allocated space
|
||||
… for SWMR read access"), and bounds reads by the file's real length.
|
||||
- **A non-SWMR open of such a file fails** in libhdf5: "file is already
|
||||
open for write (may use <h5clear file> to clear file consistency flags)".
|
||||
- **The writer only appends.** New objects and attributes cannot be created
|
||||
in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked
|
||||
datasets with one unlimited dimension use an Extensible Array index, with
|
||||
more than one a version-2 B-tree; both are updated in a SWMR-safe way.
|
||||
Fixed Array and single-chunk indexes are for datasets that cannot grow.
|
||||
- **Flush ordering.** Every metadata structure the writer uses is
|
||||
checksummed, and flush dependencies order the writes: a chunk's data is
|
||||
written before the index entry that points to it, and index blocks before
|
||||
the object header whose dataspace announces the new extent. A reader
|
||||
that reads the object header first and the index after sees an index at
|
||||
least as new as the extent, so every chunk inside the extent it read is
|
||||
either in the index or never written (then it reads as the fill value, as
|
||||
it does for libhdf5's reader).
|
||||
- **Refresh.** A reader sees a dataset's new extent only when it refreshes
|
||||
it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's
|
||||
cached metadata and reads the object header again.
|
||||
- **Retries.** Reads are not atomic against writes on every system, so a
|
||||
checksum can fail when a structure is read while the writer rewrites it.
|
||||
A SWMR reader reads checksummed metadata up to 100 times before failing
|
||||
(`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1
|
||||
otherwise — `H5Ppublic.h` of HDF5 1.14.6).
|
||||
|
||||
## What clawhdf5 did before
|
||||
|
||||
- `File::open` of a SWMR-flagged file bounded every read by the recorded
|
||||
EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the
|
||||
file). A file copied or read mid-write therefore listed, but every
|
||||
chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and
|
||||
`h5rs check` reported chunk indexes "past the end of the file".
|
||||
- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a
|
||||
per-file chunk cache that keeps each dataset's chunk index and decoded
|
||||
chunks for the life of the `File`. A reader could not see growth at all,
|
||||
and a mapping of a file that is being rewritten can change under a read.
|
||||
|
||||
## Design
|
||||
|
||||
1. **Bound SWMR files by their length.** `Superblock::data_end` returns the
|
||||
file's length for a version-3 superblock with the SWMR-write flag,
|
||||
whatever EOF it records, as libhdf5's SWMR reader does. Every existing
|
||||
open path (`File::open`, `open_storage`, `h5rs`) then reads a finished
|
||||
copy of a live file. We keep opening such files without a SWMR flag
|
||||
(libhdf5 refuses): it is read-only and the alternative is an error.
|
||||
2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.**
|
||||
- The file is read with positioned reads (`FileStorage`, `pread` on
|
||||
Unix, `seek_read` on Windows), never mapped, and `len()` is the file's
|
||||
current length, so reads past the length seen at open work.
|
||||
- The facade's view of the file (`FileData`) is *live*: reads are not
|
||||
clamped to an end fixed at open, only by the storage's current length.
|
||||
- The chunk cache is not used: every read reads the chunk index and the
|
||||
chunks it needs again. A cached index would hide new chunks, and a
|
||||
cached partial edge chunk would read as fill where the writer has since
|
||||
written data (unfiltered edge chunks are rewritten in place).
|
||||
- Only a file whose superblock has the SWMR-write flag when it is opened
|
||||
is read live. Any other file is read exactly as `File::open` reads it
|
||||
(bounded by its recorded EOF, through the chunk cache, no retries), and
|
||||
`is_swmr_read()` is `false`. libhdf5's SWMR reader is looser: it skips
|
||||
the end-of-allocation check in `H5FD_read` for every file it opens,
|
||||
flagged or not, yet still refuses an object header past the EOF
|
||||
(`H5O_protect`, "address of object past end of allocation"). On a
|
||||
closed file whose EOF is below its length (checked 2026-09-27, h5py
|
||||
3.16 / HDF5 2.0, the mid-write fixture with its flags cleared) it
|
||||
therefore reads a dataset whose chunk index lies past the EOF, which
|
||||
its plain reader refuses. We follow the plain reader there: the EOF of
|
||||
a file no SWMR writer has open is what the file says it is.
|
||||
- A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would
|
||||
serve stale bytes; `HttpStorage` also pins a file by ETag and length and
|
||||
refuses a changed file. Remote SWMR is out of scope.
|
||||
3. **`Dataset::refresh()`** reads the dataset's object header again (same
|
||||
address) and replaces the handle's copy, so `shape()` and every later
|
||||
read use the new extent. Like h5py, a handle that is not refreshed keeps
|
||||
its extent; its reads still read the index as it is now, and only return
|
||||
elements inside that extent.
|
||||
4. **Bounded retries.** In a live file, an operation (open, lookups and
|
||||
listings, refresh, every dataset read — typed, raw, selections, strings
|
||||
and variable-length data with their global-heap decoding — attributes,
|
||||
`File::decode_*`, `verify_provenance`) that fails with an error a
|
||||
concurrent write can
|
||||
cause is run again from the start, up to `File::swmr_read_attempts()`
|
||||
times (default 100, libhdf5's default; `set_swmr_read_attempts` changes
|
||||
it), sleeping 1 µs, 2 µs, … up to 10 ms between attempts (under a second
|
||||
in all). libhdf5 retries the one structure whose checksum failed; we
|
||||
retry the whole operation, because the parsers are pure functions of the
|
||||
bytes they read. Which errors: those libhdf5 retries (`H5C__load_entry`)
|
||||
— a checksum mismatch, and a failure to decode the prefix it reads
|
||||
before the checksum to size the structure (for an object header its
|
||||
signature and version: a test that garbles every byte of a header gets
|
||||
`InvalidObjectHeaderVersion`) — and a read past the file's current end
|
||||
(short here; libhdf5 reads zeros there, which fail the checksum). Every
|
||||
other error (an unsupported version or message, a file that is not HDF5,
|
||||
a structure corrupt behind a valid checksum) is returned at once: an
|
||||
earlier version retried nearly every format error, so a permanent one
|
||||
cost about 0.9 s of pauses. Non-live files never retry.
|
||||
Results are only returned from a run where every structure verified, so
|
||||
a torn metadata read is an error, never data. `File::swmr_retries()`
|
||||
counts the retries (libhdf5: `H5Fget_metadata_read_retry_info`).
|
||||
Attribute reads leave out an attribute they cannot read (or return a
|
||||
variable-length string one as `AttrValue::Raw`) instead of failing;
|
||||
on a live file such an error of a retried kind runs the read again too,
|
||||
and after the last attempt the last result is returned as before. The
|
||||
zero-copy reads (`read_raw_ref`, `read_as_slice`, `read_*_zerocopy`)
|
||||
need the file in memory, which a live file never is: they report
|
||||
`None` / `ContiguousStorageRequired` without reading data.
|
||||
Global heap collections (variable-length data) have no checksum in
|
||||
HDF5, like raw data, so a torn read of one is only caught when it fails
|
||||
a check (its signature, a bound).
|
||||
Raw data has no checksum in HDF5 (unless Fletcher-32 is on), in libhdf5
|
||||
as here: correctness rests on the writer's ordering (chunk data before
|
||||
the index entry, only appends), as for libhdf5's reader.
|
||||
5. **Writer state.** `File::swmr_writer_active()` reads the superblock
|
||||
flags again, so a reader can tell when the writer has closed the file.
|
||||
A writer that crashed or was killed never clears the flag, so a reader
|
||||
that follows a file needs a second stop condition (the README example
|
||||
stops after a minute without growth).
|
||||
|
||||
What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol,
|
||||
not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR
|
||||
writer cannot add them), and `MmapFile`/`LazyFile`.
|
||||
|
||||
## Tests
|
||||
|
||||
- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose
|
||||
EOF is below the file length.
|
||||
- `crates/clawhdf5/tests/swmr_interop.rs`:
|
||||
- a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader;
|
||||
- the same bytes with the flags cleared read as `File::open` reads them
|
||||
(and how h5py's two readers read them);
|
||||
- a permanent error returns at once; every read path (listings,
|
||||
attributes, strings, variable-length data) with each of its reads
|
||||
failing once in turn returns the same result;
|
||||
- a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a
|
||||
2-D dataset with one unlimited dimension (Extensible Array, one of them
|
||||
gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step,
|
||||
while a Rust reader refreshes and reads them in a loop, and an h5py SWMR
|
||||
reader does the same as the reference. Every value read must be the value
|
||||
the writer wrote (a deterministic function of its position), extents
|
||||
never shrink, and after the writer closes both readers must read the
|
||||
same data as h5py.
|
||||
|
||||
## Status
|
||||
|
||||
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
|
||||
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
|
||||
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
|
||||
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
|
||||
build): no read returned a value the writer had not written at that
|
||||
position, h5py's reader agreed, and retries were needed but rare (the
|
||||
failures seen were checksum mismatches, each cured by one retry). A
|
||||
variant of the test with the chunk cache left on in live mode fails it
|
||||
(stale chunk index / edge chunk), which is why live files do not use it.
|
||||
|
||||
Also found: `File::open` of such a file had been failing since the
|
||||
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).
|
||||
+183
-10
@@ -7,6 +7,56 @@ deleting it.
|
||||
|
||||
---
|
||||
|
||||
## Files a SWMR writer had open could not be read past a stale end of file
|
||||
|
||||
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any
|
||||
release: reads have been bounded by the recorded end of file since
|
||||
`7d7a7e7` (2026-09-26), which no release contains.
|
||||
|
||||
A libhdf5 writer in SWMR mode (h5py `f.swmr_mode = True`) sets the
|
||||
superblock's SWMR-write flag and does not keep its end-of-file address up to
|
||||
date: a copy h5py made of its own file mid-write records 715 in a
|
||||
6 030-byte file (h5py 3.16 / HDF5 2.0, tank). `Superblock::data_end` only
|
||||
ignored the recorded end when it lay *past* the end of the file, so every
|
||||
reader bounded such a file at 715 bytes: it listed, but every chunked read
|
||||
failed ("unexpected EOF: need 787 bytes, have 715") and `h5rs check`
|
||||
reported the chunk indexes past the end of the file. Never wrong data.
|
||||
|
||||
**Fix:** for a v3 superblock with the SWMR-write flag, the data ends at the
|
||||
end of the file, as libhdf5's SWMR reader reads it. **Test:**
|
||||
`crates/clawhdf5/tests/swmr_interop.rs` (the mid-write copy,
|
||||
`tests/fixtures/swmr_mid_write.h5`, through every open path and against
|
||||
h5py's SWMR reader). A file still being written is read with
|
||||
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
|
||||
at its length at open and is not meant for files that change while open.
|
||||
|
||||
## Shrinking a chunked dataset with no recorded maximum scrambled it
|
||||
|
||||
**Status:** fixed 2026-09-27, before any release (`FileEditor::resize`
|
||||
shipped on main in PR #18, a4c2ace). Files the writer produced before the
|
||||
fix still lack the maximum; see *Existing files*.
|
||||
|
||||
clawhdf5's writer stored no maximum dimensions for a chunked dataset
|
||||
created without a `maxshape` (libhdf5 always stores one, equal to the
|
||||
dimensions when none is given). With none recorded, the maximum is the
|
||||
current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index
|
||||
places chunks by the maximum. `FileEditor::resize` changed only the
|
||||
current dimensions, so a shrink moved every existing chunk and every
|
||||
reader returned wrong values; after a shrink the dataset could not grow
|
||||
back. Found by the review of the Python editing work.
|
||||
|
||||
**Fix:** before the first resize of such a dataset the editor records the
|
||||
maximum libhdf5 would have written (the dimensions the index was built
|
||||
with); the writer now records it for every chunked dataset. **Test:**
|
||||
`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:**
|
||||
datasets written before the fix have no recorded maximum. The fixed
|
||||
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
|
||||
does not** — it scrambles them the same way, and lets them grow past
|
||||
their Fixed Array. Resize them once with a fixed `FileEditor` (a resize
|
||||
to the same shape changes nothing; shrink and grow back) before letting
|
||||
libhdf5 resize them. A dataset already shrunk by the unfixed
|
||||
editor holds misplaced chunks; rewrite it from a good copy.
|
||||
|
||||
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
|
||||
|
||||
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to
|
||||
@@ -116,6 +166,57 @@ space it leaves is too small for its next, larger version.
|
||||
**No journal.** A crash while an edit patches existing structures can leave
|
||||
the file inconsistent; see the `FileEditor` documentation.
|
||||
|
||||
**Renamed files outside Linux.** Edits always go to the file the editor
|
||||
opened. `FileEditor::reader()` (which the Python `'r+'` handle reads
|
||||
through) reopens it through `/proc/self/fd` on Linux; elsewhere it reopens
|
||||
the path, and after the path was renamed or replaced it fails (on Unix;
|
||||
Windows cannot tell and would read whatever the path names).
|
||||
|
||||
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
|
||||
|
||||
**Status:** open (added 2026-09-27). The Python bindings edit through
|
||||
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
|
||||
before anything is written. On top of them:
|
||||
|
||||
- **No new or deleted objects:** `create_dataset`/`create_group` in an
|
||||
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
|
||||
`NotImplementedError` (the editor changes values, shapes and
|
||||
attributes only). Mode `'a'` works on an existing file only.
|
||||
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
|
||||
whole elements of the same structured dtype are), HDF5 array-type
|
||||
elements, variable-length data, strings padded with spaces or
|
||||
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
|
||||
ones, h5py's, are writable), compounds containing such strings, null
|
||||
dataspaces, index-list writes of more than 2²² elements (write them
|
||||
in slices), and boolean-mask keys (`ds[mask] = v`, and mask reads;
|
||||
h5py supports both).
|
||||
- **`str` attributes are fixed-length UTF-8**, where h5py writes
|
||||
variable-length strings: h5py reads them back as `bytes`
|
||||
(`numpy.bytes_`), not `str`.
|
||||
- **Numeric conversion follows libhdf5's native-order results, not its
|
||||
bugs.** Arrays are converted as libhdf5 converts them (integers
|
||||
saturate, floats are truncated toward zero and clipped), checked value by
|
||||
value against h5py 3.16 / HDF5 2.0 in
|
||||
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
|
||||
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
|
||||
differs from h5py on purpose:
|
||||
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
|
||||
minimum or 2⁶³ depending on the type);
|
||||
- when the dataset or the array is not in native byte order, libhdf5's
|
||||
"soft" conversions store a float in (-1, 0) as the integer minimum and
|
||||
wrap an unsigned value too large for the signed type of the same size
|
||||
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
|
||||
native order;
|
||||
- libhdf5's native casts that are undefined in C: half floats into
|
||||
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
|
||||
integers (→ minimum), a float equal to the integer maximum rounded up
|
||||
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
|
||||
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
|
||||
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
|
||||
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
|
||||
- **Each edit reopens the file** (a new memory map) so that reads see it;
|
||||
reads from other threads wait while an edit is written.
|
||||
|
||||
## Selection reads that decode more than the selection
|
||||
|
||||
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
|
||||
@@ -844,9 +945,14 @@ cache, but:
|
||||
`read_*_zerocopy`) need the file in memory and answer
|
||||
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
|
||||
panics for such a file (`File::contiguous_bytes()` is the fallible form).
|
||||
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
|
||||
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
|
||||
`remote` feature).
|
||||
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
|
||||
(`h5rs` and the Python bindings read through `File::storage`, and take
|
||||
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
|
||||
|
||||
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
|
||||
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
|
||||
feature; the wasm reader's `openUrl` reads by range requests since
|
||||
2026-09-27, its `open(bytes)` takes a whole file).
|
||||
- The file's length is read once, at open: a growing file (SWMR) is not
|
||||
followed (milestone M5). A remote file is pinned at open, so one that
|
||||
grows is `RemoteError::FileChanged`.
|
||||
@@ -855,16 +961,25 @@ cache, but:
|
||||
and so `Dataset::read_*`) lists a damaged dataset's chunks in hash-map
|
||||
order, so which failing chunk it reports can differ from one `File` to
|
||||
the next (`cve-2025-2310.h5`); the values of a dataset that reads are
|
||||
not affected.
|
||||
not affected. **Fixed 2026-09-27:** the chunk cache keeps the chunks in
|
||||
the order the index lists them, as the uncached readers do
|
||||
(`several_damaged_chunks_report_the_same_chunk_every_time`).
|
||||
|
||||
## Remote files (`clawhdf5-remote`) limits
|
||||
|
||||
**Status:** open (added 2026-09-26, milestone M3 of
|
||||
`docs/design/range-reads.md`).
|
||||
|
||||
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
|
||||
parses through `File::as_bytes`, which a remote file does not have; the
|
||||
wasm reader's `openUrl` is milestone M4.
|
||||
- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
|
||||
milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
|
||||
the default wheel reads plain `http://` only: `https://` needs a wheel
|
||||
built with `--features https` (rustls with ring, which compiles C), and
|
||||
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
|
||||
The Python tests run against an in-process `http.server` only.
|
||||
|
||||
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
|
||||
`File::as_bytes`, which a remote file does not have. (The browser can
|
||||
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
|
||||
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
||||
The design's policy of using a paged file's page size as the block size
|
||||
is not implemented, and only the first block is read ahead.
|
||||
@@ -903,10 +1018,68 @@ cache, but:
|
||||
|
||||
## `clawhdf5-wasm` (browser) limits
|
||||
|
||||
**Status:** open (by design for now; added 2026-09-26).
|
||||
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
|
||||
|
||||
- The whole file is held in memory: `open()` takes its bytes. There are no
|
||||
HTTP range reads, so a multi-GB file does not fit a browser tab.
|
||||
- `open()` holds the whole file in memory (it takes its bytes), so a
|
||||
multi-GB local file does not fit a browser tab. A file on a web server
|
||||
can be opened with `openUrl()` instead, which fetches only the byte
|
||||
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
|
||||
with these limits:
|
||||
- **Round trips:** a call runs as passes over the blocks fetched so far
|
||||
and is re-run after each wave of misses, so a call costs one round
|
||||
trip per wave, not one for everything: a chunk index is walked a
|
||||
level (or a node) per round trip, while the chunks of a read are
|
||||
fetched together. Listing a group asks for every child's object
|
||||
header, and every node of a level of the group's index, in one pass
|
||||
(since 2026-09-27; it was one round trip per header block): 3000
|
||||
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
|
||||
`libver="latest"` file (dense links). Each pass re-parses what the
|
||||
call reads (CPU, not network). With headers spread through the file
|
||||
(h5py writes each next to its data) a listing still fetches most of
|
||||
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
|
||||
- **Memory:** a call keeps every block it reads until it finishes (the
|
||||
cache budget applies between calls). It may fetch at most `maxFetch`
|
||||
bytes (512 MiB by default, at most 1 GiB), and a single read longer
|
||||
than that fails before anything is fetched: a hostile server cannot
|
||||
make the page fetch or allocate what a file's lengths claim. `read()`
|
||||
of a dataset that would take more than 1 GiB while it is decoded
|
||||
(stored bytes, the values widened to 64 bits, the result) fails,
|
||||
naming `readHyperslab`; read windows of large datasets. On wasm32 a
|
||||
buffer past 2 GiB cannot exist and a failed allocation aborts the
|
||||
whole module (every open file on the page), which these limits keep
|
||||
from happening; before 2026-09-27 both did abort it.
|
||||
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
|
||||
open. The format code turns file offsets into `usize` to use them
|
||||
(with a clean error past it), which is 32 bits on wasm32, so nothing
|
||||
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
|
||||
tested (with a mock server); files above 200 MB have not been served
|
||||
for real.
|
||||
- **Cross-origin servers** must allow CORS for the page's origin and
|
||||
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
|
||||
answer `HEAD` with `Content-Length`. The file is pinned at open by its
|
||||
`ETag` or `Last-Modified` (and its length): when the page cannot see
|
||||
either header, only a change of length is detected.
|
||||
- **A server without range support** (it answers `200`) costs a whole
|
||||
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
|
||||
with `fallback: "error"`. Every body, this one and each `206`, is read
|
||||
as it arrives and cut off past its limit (the range asked for, or
|
||||
`maxDownload`): a server cannot make the page buffer more.
|
||||
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
|
||||
size is not used. No retries: a failed request fails the call (calling
|
||||
again retries it; what was fetched stays cached).
|
||||
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
||||
a local server, cross-origin included (a page on 127.0.0.1 reading a
|
||||
file from localhost, with and without exposed headers); not in Firefox
|
||||
or Safari.
|
||||
- The native corpus comparison (`tests/lazy.rs` with
|
||||
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
|
||||
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
|
||||
and which chunk's error is reported depends on the iteration order of
|
||||
the chunk index (a `HashMap`, seeded per process), so the lazy and
|
||||
the range-storage reads can name different errors. Both are errors;
|
||||
not specific to `openUrl` (it predates it). **Fixed 2026-09-27:** the
|
||||
chunk cache keeps the index's chunk order, so every read path names the
|
||||
same (first) damaged chunk.
|
||||
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
||||
refused with an error naming the type; attributes of those types come back
|
||||
as `value: null` with their `dtype`.
|
||||
|
||||
+100
-20
@@ -2,10 +2,12 @@
|
||||
|
||||
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
|
||||
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
|
||||
file, browse its groups, and look at a dataset's type, shape, attributes
|
||||
and values (a 50 x 12 window at a time, read as a hyperslab, with the
|
||||
leading dimensions of a 3-D+ dataset held at chosen indices). The file never
|
||||
leaves the page.
|
||||
file or give a URL, browse its groups, and look at a dataset's type, shape,
|
||||
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
|
||||
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
|
||||
local file never leaves the page. A file given by URL is not downloaded:
|
||||
only the byte ranges each view needs are fetched (HTTP range requests),
|
||||
and the header shows the requests and bytes that has cost so far.
|
||||
|
||||
## Build and open
|
||||
|
||||
@@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe
|
||||
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
|
||||
```
|
||||
|
||||
Then open <http://localhost:8000/>. `?file=<url>&path=<object>` opens a
|
||||
file from a URL (same origin, or one serving CORS headers) and selects an
|
||||
object in it, e.g. `?file=data/run1.h5&path=/results/energy`.
|
||||
Then open <http://localhost:8000/>. The URL box, or
|
||||
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
|
||||
sending CORS headers, see Limits) and selects an object in it, e.g.
|
||||
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
|
||||
answer range requests, so a file served by it is downloaded whole (the
|
||||
header says so); `test/serve.py` does:
|
||||
|
||||
```bash
|
||||
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
|
||||
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
|
||||
```
|
||||
|
||||
## JavaScript API
|
||||
|
||||
```js
|
||||
import init, { open } from "./pkg/clawhdf5_wasm.js";
|
||||
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
|
||||
await init();
|
||||
const f = open(new Uint8Array(await blob.arrayBuffer()));
|
||||
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
|
||||
@@ -32,8 +42,41 @@ f.attrs("/grid"); // [{ name, value, dtype }]
|
||||
f.read("/grid"); // { shape, dtype, data }
|
||||
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
|
||||
f.free();
|
||||
|
||||
// A file on a web server, read by HTTP range requests as needed. The same
|
||||
// methods, each returning a promise.
|
||||
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
|
||||
await r.list("/");
|
||||
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
|
||||
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
|
||||
r.free();
|
||||
```
|
||||
|
||||
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
|
||||
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
|
||||
between calls, default 64 MiB), `maxFetch` (bytes one call may fetch, and
|
||||
so the longest single read, up to 1 GiB, default 512 MiB), `fallback`
|
||||
(`"download"`, the default, reads the whole file when the server ignores
|
||||
`Range`, up to `maxDownload` bytes, default 512 MiB, at most 1 GiB;
|
||||
`"error"` refuses such a server), `headers` (a `Headers`, `[name, value]`
|
||||
pairs or an object) and `credentials` (passed to every request),
|
||||
`parallel` (requests in flight, 1 to 1024, default 6; when one fails the
|
||||
others are aborted), `fetch` (a `fetch`-compatible function to use).
|
||||
|
||||
How it works: the reader is synchronous and a page cannot block on the
|
||||
network, so each call runs as a *pass* over the blocks fetched so far. A
|
||||
pass that needs a block not yet fetched is abandoned, the missing blocks
|
||||
are fetched (in parallel, adjacent blocks in one request), and the pass is
|
||||
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
||||
costs one request (the first block, which also gives the file's size);
|
||||
listing a group whose metadata is in blocks already fetched costs none,
|
||||
and otherwise a round trip per level of the group's index plus one for
|
||||
its children's headers, all fetched together;
|
||||
reading a chunked dataset costs a round trip for its chunk index (a few
|
||||
for a deep one) and one batch of requests for its chunks. Every answer is
|
||||
checked — a `206` with exactly the bytes asked for, from the same file
|
||||
(ETag or Last-Modified, and length) — or the call fails.
|
||||
|
||||
`data` is the typed array of the stored width (`Float64Array`,
|
||||
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
|
||||
`BigUint64Array`), or an array of strings for fixed- and variable-length
|
||||
@@ -43,7 +86,17 @@ else throws an `Error` naming the type.
|
||||
|
||||
## Limits
|
||||
|
||||
- Read-only, and the whole file is held in memory (no range requests).
|
||||
- Read-only. `open(bytes)` holds the whole file in memory.
|
||||
- `openUrl`: a cross-origin server must allow CORS and expose
|
||||
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
|
||||
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
|
||||
`Last-Modified` either, a file replaced on the server is detected only by
|
||||
a change of length. A call keeps what it reads until it finishes, and
|
||||
may fetch at most `maxFetch`; `read()` of a dataset that would take more
|
||||
than 1 GiB to decode is refused (read windows of large datasets with
|
||||
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
|
||||
Every response body is cut off past the length asked for. More in
|
||||
`docs/known-issues.md`.
|
||||
- Compound, reference, opaque and variable-length-sequence datasets are
|
||||
refused with an error. Attributes of those types are listed with
|
||||
`value: null` and their `dtype`.
|
||||
@@ -56,28 +109,55 @@ else throws an `Error` naming the type.
|
||||
## Tests
|
||||
|
||||
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
|
||||
`fixture.nc` (netCDF4) with `test/make_fixture.py`, then:
|
||||
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
|
||||
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
|
||||
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
|
||||
counter, `/norange/...` for a server without range support,
|
||||
`/noexpose/...` and `/unexposed/...` for one that does not expose its
|
||||
headers to CORS), then:
|
||||
|
||||
- runs `test/test.mjs` under Node: every dataset (whole and a strided
|
||||
hyperslab), listing and attribute is compared with what libhdf5 reads
|
||||
back, error paths are checked, and so are the page's DOM-free helpers
|
||||
(`viewer-lib.js`);
|
||||
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
|
||||
and 512 B blocks), calls in flight at once, the request budget of
|
||||
`big.h5` (listing it and reading small things of it must take at most 8
|
||||
requests and under 5% of the file), the download fallback and its
|
||||
limit, and the errors: HTTP status, a file that changes, a server that
|
||||
sends the wrong bytes or stops honouring `Range`. Also the cross-origin
|
||||
path with no exposed headers (`/noexpose/`: length from `HEAD`), the
|
||||
size limits on `make_fixture.py`'s `limits.h5`, `hostile_vl.h5` (a heap
|
||||
collection claiming 2 GiB) and `far.h5` (data at 3 GiB, served by a
|
||||
mock), bodies longer than asked for, `headers` forms, `parallel`, and
|
||||
sibling requests aborted after a failure.
|
||||
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
|
||||
to 16 MiB) read by URL with the same file read from bytes;
|
||||
- runs `test/browser.sh`: loads the page in headless Chromium with
|
||||
`?file=fixture.h5&path=...` for eight objects and checks the rendered tree,
|
||||
types, shapes, attribute and value cells, and the error shown for an
|
||||
unsupported type. Skipped when no Chromium is found (`CHROME` names one;
|
||||
a Playwright download under `~/.cache/ms-playwright` is picked up).
|
||||
Drag-and-drop and the file picker are not driven by it; they share
|
||||
`load()` with the `?file=` path.
|
||||
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
|
||||
tree, types, shapes, attribute and value cells, the request counter, and
|
||||
the error shown for an unsupported type; then the file from another
|
||||
origin (localhost), with CORS exposing `Content-Range` and exposing
|
||||
nothing (`/unexposed/`, where the server must see a `HEAD`), a server
|
||||
without range support, and `big.h5` (a small dataset and a window of the large one,
|
||||
with a single-digit percentage of the file fetched). Skipped when no
|
||||
Chromium is found (`CHROME` names one; a Playwright download under
|
||||
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
|
||||
and the URL box are not driven by it; they share `setFile()` with the
|
||||
`?file=` path.
|
||||
|
||||
The same expectations are checked natively, without Node, by
|
||||
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI
|
||||
container has no Node or browser).
|
||||
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
|
||||
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
|
||||
corpus too), which is what CI runs (the CI container has no Node or
|
||||
browser).
|
||||
|
||||
## Size
|
||||
|
||||
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
|
||||
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`:
|
||||
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
|
||||
larger now and the table has not been re-measured: the reader has grown
|
||||
since, and `openUrl` (2026-09-27) made the facade's range-read path
|
||||
reachable from JavaScript and added promise glue and `remote.js`.
|
||||
|
||||
| | raw | gzip -9 |
|
||||
|---|---:|---:|
|
||||
|
||||
@@ -56,27 +56,39 @@
|
||||
border: 1px solid var(--line); border-radius: 4px; }
|
||||
.error { color: var(--bad); font-family: var(--mono); white-space: pre-wrap; }
|
||||
.muted { color: var(--muted); }
|
||||
form.url { display: flex; gap: 6px; margin-left: auto; flex: 1 1 320px; max-width: 560px; }
|
||||
form.url input { flex: 1; min-width: 0; font: 13px var(--mono); padding: 4px 8px; background: var(--panel);
|
||||
color: var(--ink); border: 1px solid var(--line); border-radius: 6px; }
|
||||
.netstats { flex-basis: 100%; color: var(--muted); font-size: 12.5px; font-family: var(--mono); }
|
||||
.netstats:empty { display: none; }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<header>
|
||||
<h1>HDF5 Viewer</h1>
|
||||
<span class="file" id="filename">no file</span>
|
||||
<form class="url" id="urlform">
|
||||
<input type="url" id="url" placeholder="https://…/file.h5 (read by HTTP range requests)" aria-label="File URL">
|
||||
<button class="button" type="submit">Open URL</button>
|
||||
</form>
|
||||
<label class="button">Open file…<input type="file" id="picker" accept=".h5,.hdf5,.he5,.nc,.nc4,.cdf" hidden></label>
|
||||
<span class="netstats" id="netstats"></span>
|
||||
</header>
|
||||
<main>
|
||||
<nav id="tree"></nav>
|
||||
<section id="detail">
|
||||
<div class="drop">
|
||||
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, or use “Open file…”.</p>
|
||||
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; it is not uploaded anywhere.</p>
|
||||
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, use “Open file…”, or give a URL.</p>
|
||||
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; a local file is not uploaded anywhere.
|
||||
A file given by URL is not downloaded: only the byte ranges each view needs are fetched (HTTP range requests),
|
||||
and the requests and bytes it has cost are shown at the top.</p>
|
||||
<p class="muted" id="version"></p>
|
||||
</div>
|
||||
</section>
|
||||
</main>
|
||||
<script type="module">
|
||||
import init, { open, version } from "./pkg/clawhdf5_wasm.js";
|
||||
import { joinPath, formatValue, viewWindow, toRows, perElement } from "./viewer-lib.js";
|
||||
import init, { open, openUrl, version } from "./pkg/clawhdf5_wasm.js";
|
||||
import { joinPath, formatValue, viewWindow, toRows, perElement, formatStats } from "./viewer-lib.js";
|
||||
|
||||
const $ = (id) => document.getElementById(id);
|
||||
const el = (tag, props = {}, ...kids) => {
|
||||
@@ -85,30 +97,58 @@ const el = (tag, props = {}, ...kids) => {
|
||||
return e;
|
||||
};
|
||||
|
||||
// The open file: an H5File (a local file, read in memory; synchronous
|
||||
// methods) or a RemoteFile (openUrl: bytes fetched by range requests as
|
||||
// needed; methods return promises). Every call below is awaited, so both
|
||||
// work the same.
|
||||
let file = null;
|
||||
let selectedNode = null;
|
||||
// Bumped by every selection, so a slow read for an object no longer shown
|
||||
// does not overwrite the current one.
|
||||
let showSeq = 0;
|
||||
|
||||
await init();
|
||||
$("version").textContent = `clawhdf5 ${version()}`;
|
||||
|
||||
async function load(blob) {
|
||||
const bytes = new Uint8Array(await blob.arrayBuffer());
|
||||
// Requests and bytes a remote file has cost so far.
|
||||
function updateStats() {
|
||||
$("netstats").textContent = file && file.stats ? formatStats(file.stats()) : "";
|
||||
}
|
||||
|
||||
async function setFile(name, opener) {
|
||||
if (file) file.free();
|
||||
file = null;
|
||||
$("filename").textContent = blob.name;
|
||||
selectedNode = null;
|
||||
showSeq++;
|
||||
$("filename").textContent = name;
|
||||
$("tree").replaceChildren();
|
||||
$("netstats").textContent = "";
|
||||
$("detail").replaceChildren(el("p", { className: "muted", textContent: `Opening ${name}…` }));
|
||||
try {
|
||||
file = open(bytes);
|
||||
file = await opener();
|
||||
} catch (e) {
|
||||
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${blob.name}: ${e.message}` }));
|
||||
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${name}: ${e.message}` }));
|
||||
return;
|
||||
}
|
||||
updateStats();
|
||||
const root = el("ul", { className: "tree" });
|
||||
root.append(treeNode("/", "/", "group"));
|
||||
$("tree").append(root);
|
||||
root.querySelector(".node").click();
|
||||
await activate(root.querySelector(".node"));
|
||||
}
|
||||
|
||||
async function load(blob) {
|
||||
const bytes = new Uint8Array(await blob.arrayBuffer());
|
||||
await setFile(blob.name, async () => open(bytes));
|
||||
}
|
||||
|
||||
async function loadUrl(url) {
|
||||
$("url").value = url;
|
||||
await setFile(url, () => openUrl(new URL(url, location.href).href));
|
||||
}
|
||||
|
||||
// A tree row: clicking it selects it (and expands a group); `activate`
|
||||
// does the same and resolves once the listing and the detail are shown.
|
||||
function treeNode(path, name, kind) {
|
||||
const li = el("li");
|
||||
const icon = el("span", { className: "icon", textContent: kind === "group" ? "▸" : "·" });
|
||||
@@ -116,34 +156,39 @@ function treeNode(path, name, kind) {
|
||||
row.dataset.path = path;
|
||||
li.append(row);
|
||||
let children = null;
|
||||
row.addEventListener("click", () => {
|
||||
row.activate = async () => {
|
||||
if (selectedNode) selectedNode.classList.remove("selected");
|
||||
row.classList.add("selected");
|
||||
selectedNode = row;
|
||||
const shown = show(path, kind);
|
||||
if (kind === "group") {
|
||||
if (children) {
|
||||
children.hidden = !children.hidden;
|
||||
} else {
|
||||
children = el("ul", { className: "tree" });
|
||||
li.append(children);
|
||||
try {
|
||||
for (const c of file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
|
||||
for (const c of await file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
|
||||
} catch (e) {
|
||||
children.append(el("li", { className: "error", textContent: e.message }));
|
||||
}
|
||||
li.append(children);
|
||||
updateStats();
|
||||
}
|
||||
icon.textContent = children.hidden ? "▸" : "▾";
|
||||
}
|
||||
show(path, kind);
|
||||
});
|
||||
await shown;
|
||||
};
|
||||
row.addEventListener("click", () => row.activate());
|
||||
return li;
|
||||
}
|
||||
|
||||
function attrsTable(path) {
|
||||
const activate = (row) => row.activate();
|
||||
|
||||
async function attrsTable(path) {
|
||||
let attrs, errors;
|
||||
try {
|
||||
attrs = file.attrs(path);
|
||||
errors = file.attrErrors(path);
|
||||
attrs = await file.attrs(path);
|
||||
errors = await file.attrErrors(path);
|
||||
} catch (e) {
|
||||
return el("p", { className: "error", textContent: e.message });
|
||||
}
|
||||
@@ -157,14 +202,16 @@ function attrsTable(path) {
|
||||
return el("div", { className: "scroll" }, t);
|
||||
}
|
||||
|
||||
function show(path, kind) {
|
||||
async function show(path, kind) {
|
||||
const seq = ++showSeq;
|
||||
const current = () => seq === showSeq;
|
||||
const out = [el("h2", { textContent: path })];
|
||||
if (kind === "dataset") {
|
||||
let info;
|
||||
try {
|
||||
info = file.info(path);
|
||||
info = await file.info(path);
|
||||
} catch (e) {
|
||||
$("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
|
||||
if (current()) $("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
|
||||
return;
|
||||
}
|
||||
const max = info.maxshape === null ? "—" : `(${info.maxshape.map((d) => d ?? "∞").join(", ")})`;
|
||||
@@ -172,15 +219,16 @@ function show(path, kind) {
|
||||
el("dt", { textContent: "type" }), el("dd", { textContent: info.dtype }),
|
||||
el("dt", { textContent: "shape" }), el("dd", { textContent: `(${info.shape.join(", ")})` }),
|
||||
el("dt", { textContent: "max shape" }), el("dd", { textContent: max })));
|
||||
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
|
||||
out.push(el("h3", { textContent: "Values" }), valuesView(path, info));
|
||||
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
|
||||
out.push(el("h3", { textContent: "Values" }), await valuesView(path, info, current));
|
||||
} else {
|
||||
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
|
||||
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
|
||||
}
|
||||
$("detail").replaceChildren(...out);
|
||||
updateStats();
|
||||
if (current()) $("detail").replaceChildren(...out);
|
||||
}
|
||||
|
||||
function valuesView(path, info) {
|
||||
async function valuesView(path, info, current) {
|
||||
const shape = info.shape;
|
||||
const per = perElement(info.elementShape);
|
||||
const state = { row: 0, col: 0, rows: 50, cols: 12, fixed: shape.slice(0, Math.max(0, shape.length - 2)).map(() => 0) };
|
||||
@@ -202,15 +250,19 @@ function valuesView(path, info) {
|
||||
if (shape.length >= 2) controls.append(num("column", "col"));
|
||||
state.fixed.forEach((_, i) => controls.append(num(`dim ${i}`, "fixed", i)));
|
||||
|
||||
function render() {
|
||||
let renderSeq = 0;
|
||||
async function render() {
|
||||
const seq = ++renderSeq;
|
||||
const w = viewWindow(shape, state);
|
||||
let res;
|
||||
try {
|
||||
res = w === null ? file.read(path) : file.readHyperslab(path, w.start, w.count);
|
||||
res = w === null ? await file.read(path) : await file.readHyperslab(path, w.start, w.count);
|
||||
} catch (e) {
|
||||
body.replaceChildren(el("p", { className: "error", textContent: e.message }));
|
||||
if (seq === renderSeq) body.replaceChildren(el("p", { className: "error", textContent: e.message }));
|
||||
return;
|
||||
}
|
||||
updateStats();
|
||||
if (seq !== renderSeq || !current()) return;
|
||||
if (w === null) {
|
||||
body.replaceChildren(el("pre", { textContent: toRows(res.data, 1, 1, per)[0][0] }));
|
||||
return;
|
||||
@@ -229,41 +281,39 @@ function valuesView(path, info) {
|
||||
(shape.length >= 2 ? `, columns ${w.col}–${w.col + w.cols - 1}` : "") + ` of ${total} values` });
|
||||
body.replaceChildren(t, note);
|
||||
}
|
||||
render();
|
||||
await render();
|
||||
return box;
|
||||
}
|
||||
|
||||
// Expand the tree down to `path` and select it.
|
||||
function reveal(path) {
|
||||
async function reveal(path) {
|
||||
const rowFor = (p) => [...document.querySelectorAll(".node")].find((n) => n.dataset.path === p);
|
||||
let cur = "/";
|
||||
for (const part of path.split("/").filter(Boolean)) {
|
||||
const row = rowFor(cur);
|
||||
if (!row) return;
|
||||
const kids = row.parentElement.querySelector(":scope > ul");
|
||||
if (!kids || kids.hidden) row.click();
|
||||
if (!kids || kids.hidden) await activate(row);
|
||||
cur = joinPath(cur, part);
|
||||
}
|
||||
const target = rowFor(cur);
|
||||
if (target && target !== selectedNode) target.click();
|
||||
if (target && target !== selectedNode) await activate(target);
|
||||
}
|
||||
|
||||
// ?file=<url>&path=<object> opens a file from a URL (same origin, or one
|
||||
// that allows CORS) and selects an object in it.
|
||||
// that allows CORS) by range requests and selects an object in it.
|
||||
const params = new URLSearchParams(location.search);
|
||||
if (params.get("file")) {
|
||||
const url = params.get("file");
|
||||
try {
|
||||
const resp = await fetch(url);
|
||||
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
|
||||
const blob = await resp.blob();
|
||||
await load(new File([blob], url.split("/").pop()));
|
||||
if (file && params.get("path")) reveal(params.get("path"));
|
||||
} catch (e) {
|
||||
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot fetch ${url}: ${e.message}` }));
|
||||
}
|
||||
await loadUrl(params.get("file"));
|
||||
if (file && params.get("path")) await reveal(params.get("path"));
|
||||
document.body.dataset.ready = "1";
|
||||
}
|
||||
|
||||
$("urlform").addEventListener("submit", (e) => {
|
||||
e.preventDefault();
|
||||
const url = $("url").value.trim();
|
||||
if (url) loadUrl(url);
|
||||
});
|
||||
$("picker").addEventListener("change", (e) => e.target.files[0] && load(e.target.files[0]));
|
||||
document.addEventListener("dragover", (e) => { e.preventDefault(); document.body.classList.add("dragging"); });
|
||||
document.addEventListener("dragleave", () => document.body.classList.remove("dragging"));
|
||||
|
||||
@@ -3,10 +3,13 @@
|
||||
#
|
||||
# browser.sh FIXTURE_DIR
|
||||
#
|
||||
# FIXTURE_DIR holds fixture.h5 from make_fixture.py; ../pkg must be built.
|
||||
# The page is opened with ?file=fixture.h5&path=<object>, which fetches the
|
||||
# file, builds the tree down to <object> and shows it; the rendered DOM is
|
||||
# dumped and checked for the values libhdf5 reads.
|
||||
# FIXTURE_DIR holds fixture.h5 from make_fixture.py (and big.h5 when it was
|
||||
# run with WASM_BIG_MB); ../pkg must be built. test/serve.py serves the
|
||||
# page and, under /fix, the fixtures, with HTTP range requests (no copies,
|
||||
# no links). The page is opened with ?file=fix/<name>&path=<object>, which
|
||||
# opens the file with openUrl (range requests), builds the tree down to
|
||||
# <object> and shows it; the rendered DOM is dumped and checked for the
|
||||
# values libhdf5 reads and for the request counter.
|
||||
#
|
||||
# Browser: $CHROME, else chromium/google-chrome on PATH, else a Playwright
|
||||
# download under ~/.cache/ms-playwright. Exit 3 when none is found.
|
||||
@@ -31,63 +34,90 @@ if [ -z "$chrome" ] || [ ! -x "$chrome" ]; then
|
||||
exit 3
|
||||
fi
|
||||
|
||||
root="$(mktemp -d)"
|
||||
profiles="$(mktemp -d)"
|
||||
server=""
|
||||
cleanup() {
|
||||
[ -n "$server" ] && kill "$server" 2>/dev/null || true
|
||||
rm -rf "$root"
|
||||
rm -rf "$profiles"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
ln -s "$HERE/../index.html" "$HERE/../viewer-lib.js" "$HERE/../pkg" "$FIX/fixture.h5" "$root/"
|
||||
|
||||
port=$("$PY" -c 'import socket; s = socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1])')
|
||||
"$PY" -m http.server --bind 127.0.0.1 --directory "$root" "$port" >/dev/null 2>&1 &
|
||||
"$PY" "$HERE/serve.py" --root "fix=$FIX" --root "=$HERE/.." > "$profiles/port" &
|
||||
server=$!
|
||||
for _ in $(seq 50); do
|
||||
"$PY" -c "import urllib.request; urllib.request.urlopen('http://127.0.0.1:$port/index.html')" 2>/dev/null && break
|
||||
[ -s "$profiles/port" ] && break
|
||||
sleep 0.1
|
||||
done
|
||||
port="$(head -1 "$profiles/port")"
|
||||
|
||||
fails=0
|
||||
# A fresh profile per page: a second instance on the same profile fails.
|
||||
render() {
|
||||
local profile
|
||||
profile="$(mktemp -d "$root/profile.XXXXXX")"
|
||||
profile="$(mktemp -d "$profiles/profile.XXXXXX")"
|
||||
"$chrome" --headless --no-sandbox --disable-gpu --user-data-dir="$profile" \
|
||||
--virtual-time-budget=20000 \
|
||||
--dump-dom "http://127.0.0.1:$port/index.html?file=fixture.h5&path=$1" 2>/dev/null
|
||||
--dump-dom "http://127.0.0.1:$port/index.html?file=$1&path=$2" 2>/dev/null
|
||||
}
|
||||
# expect PATH TEXT...: every TEXT appears in the page rendered for PATH.
|
||||
# expect FILE PATH TEXT...: every TEXT appears in the page rendered for PATH
|
||||
# of FILE (a URL relative to the page); "re:TEXT" is an extended regex.
|
||||
expect() {
|
||||
local path="$1" dom
|
||||
shift
|
||||
dom="$(render "$path")"
|
||||
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$path" | tr / _).html"
|
||||
local file="$1" path="$2" dom
|
||||
shift 2
|
||||
dom="$(render "$file" "$path")"
|
||||
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$file$path" | tr / _).html"
|
||||
for text in "$@"; do
|
||||
if ! grep -qF -- "$text" <<<"$dom"; then
|
||||
echo "FAIL: page for $path lacks: $text" >&2
|
||||
local flag=-qF
|
||||
if [ "${text#re:}" != "$text" ]; then flag=-qE; text="${text#re:}"; fi
|
||||
if ! grep $flag -- "$text" <<<"$dom"; then
|
||||
echo "FAIL: page for $file $path lacks: $text" >&2
|
||||
fails=$((fails + 1))
|
||||
fi
|
||||
done
|
||||
echo "rendered $path"
|
||||
echo "rendered $file $path"
|
||||
}
|
||||
|
||||
H5=fix/fixture.h5
|
||||
# Tree (root expanded; the group row carries its path) and root attributes.
|
||||
expect "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
|
||||
expect $H5 "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
|
||||
'<td>big</td><td>9223372036854775813</td>' '(compound{x: f64, n: i32})'
|
||||
# A chunked, deflated 2-D dataset: type, shape, the first window of values.
|
||||
expect "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
|
||||
'showing rows 0–5, columns 0–9 of 60 values'
|
||||
# A chunked, deflated 2-D dataset: type, shape, the first window of values;
|
||||
# the request counter.
|
||||
expect $H5 "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
|
||||
'showing rows 0–5, columns 0–9 of 60 values' 'request' 'fetched (100.0%)'
|
||||
# Nested path revealed through the tree; big-endian float32.
|
||||
expect "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
|
||||
expect $H5 "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
|
||||
# 64-bit integers stay exact; strings; array datatype cells.
|
||||
expect "/u64" '<td>18446744073709551615</td>'
|
||||
expect "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
|
||||
expect "/pairs" '<td>[2, 3]</td>' '<dd>array[2]<i32></dd>'
|
||||
expect $H5 "/u64" '<td>18446744073709551615</td>'
|
||||
expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
|
||||
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]<i32></dd>'
|
||||
# 3-D: leading dimension held at 0, window over the last two.
|
||||
expect "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
|
||||
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
|
||||
# Unsupported type: an error, not values.
|
||||
expect "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
|
||||
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
|
||||
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
|
||||
# exposes Content-Range, and CORS that exposes nothing (/unexposed/: the
|
||||
# browser hides Content-Range and ETag, so the page learns the length with
|
||||
# a HEAD request, which the server's log must show).
|
||||
expect "http://localhost:$port/$H5" "/grid" '<td>14.75</td>' 'request'
|
||||
get() { "$PY" -c 'import sys, urllib.request; print(urllib.request.urlopen(sys.argv[1]).read().decode())' "$1"; }
|
||||
get "http://127.0.0.1:$port/__reset" >/dev/null
|
||||
expect "http://localhost:$port/unexposed/$H5" "/grid" '<td>0.25</td>' '<td>14.75</td>' 'request'
|
||||
heads="$(get "http://127.0.0.1:$port/__stats" | "$PY" -c \
|
||||
'import json, sys; print(sum(1 for l in json.load(sys.stdin)["log"] if l[4] == "HEAD" and l[0].endswith("fixture.h5")))')"
|
||||
if [ "$heads" -lt 1 ]; then
|
||||
echo "FAIL: no HEAD request for the length with unexposed headers" >&2
|
||||
fails=$((fails + 1))
|
||||
fi
|
||||
# A server without range support: the file is downloaded whole, and says so.
|
||||
expect norange/$H5 "/grid" '<td>14.75</td>' 'downloaded whole'
|
||||
# A large file: a small dataset, and a window of the large one, fetch a few
|
||||
# blocks (the counter shows a small percentage: "(1.5%)", not "(100.0%)").
|
||||
if [ -f "$FIX/big.h5" ]; then
|
||||
expect fix/big.h5 "/small" '<td>1.5</td>' '<td>3.25</td>' 're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
|
||||
expect fix/big.h5 "/big" '<dd>(25000000)</dd>' '<td>0</td>' '<td>24.5</td>' \
|
||||
're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
|
||||
fi
|
||||
|
||||
if [ "$fails" -gt 0 ]; then
|
||||
echo "browser: $fails checks failed" >&2
|
||||
|
||||
@@ -3,7 +3,9 @@ reads back from them, for the clawhdf5-wasm tests.
|
||||
|
||||
python make_fixture.py OUT_DIR
|
||||
|
||||
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json.
|
||||
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json,
|
||||
and the limit-test files OUT_DIR/limits.h5 and OUT_DIR/hostile_vl.h5 (see
|
||||
write_limits).
|
||||
Both the Rust test (crates/clawhdf5-wasm/tests/h5py_interop.rs, native) and
|
||||
the Node test (test.mjs, the built wasm package) compare against the same
|
||||
expected.json, so the two check the same values.
|
||||
@@ -14,6 +16,7 @@ encoded as strings so JSON.parse keeps 64-bit values exact.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import warnings
|
||||
from pathlib import Path
|
||||
@@ -201,3 +204,105 @@ def slab_for(obj):
|
||||
|
||||
json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)},
|
||||
open(out / "expected.json", "w"), indent=1, ensure_ascii=False)
|
||||
|
||||
|
||||
HUGE_U8 = 2**28 + 1024
|
||||
# The collection size hostile_vl.h5 claims: past 2 GiB, which a wasm32
|
||||
# buffer cannot hold.
|
||||
HOSTILE_GCOL_SIZE = 2**31 + 4096
|
||||
# The file length a server claims for hostile_vl.h5 (the tests' mock fetch
|
||||
# answers every range with zeros past the real bytes): 3 GiB, within what
|
||||
# wasm32 opens, and room for the collection.
|
||||
HOSTILE_LENGTH = 3 << 30
|
||||
|
||||
|
||||
def write_limits(out):
|
||||
"""Files for the size limits (the tests must get errors, not aborts):
|
||||
|
||||
- limits.h5: /huge_u8, 2^28 + 1024 bytes of u8 in compressed chunks
|
||||
(a small file): read whole it would take over 2 GiB while decoding;
|
||||
its last value is 7.
|
||||
- hostile_vl.h5: a variable-length string dataset /a whose global heap
|
||||
collection claims HOSTILE_GCOL_SIZE bytes, with the superblock's end
|
||||
of file set to HOSTILE_LENGTH (libhdf5 cannot read it; it is only
|
||||
served by a mock that claims that length).
|
||||
- far.h5 and far.json: /x, 16 float64 values, whose contiguous data
|
||||
address is moved FAR_SHIFT bytes on (past 2 GiB, the sign bit of a
|
||||
wasm32 isize) in a file whose end of file is moved as far; the tests'
|
||||
mock serves the data there, to show offsets up to 4 GiB work on
|
||||
wasm32.
|
||||
"""
|
||||
with h5py.File(out / "limits.h5", "w") as f:
|
||||
d = f.create_dataset("huge_u8", shape=(HUGE_U8,), dtype="u1",
|
||||
chunks=(1 << 20,), compression="gzip")
|
||||
d[-1] = 7
|
||||
path = out / "hostile_vl.h5"
|
||||
with h5py.File(path, "w", libver="earliest") as f:
|
||||
f.create_dataset("a", data=["x", "yy"], dtype=h5py.string_dtype())
|
||||
b = bytearray(path.read_bytes())
|
||||
assert b[8] == 0, "a version 0 superblock"
|
||||
b[40:48] = HOSTILE_LENGTH.to_bytes(8, "little") # end of file address
|
||||
at = b.index(b"GCOL")
|
||||
b[at + 8:at + 16] = HOSTILE_GCOL_SIZE.to_bytes(8, "little")
|
||||
path.write_bytes(bytes(b))
|
||||
|
||||
path = out / "far.h5"
|
||||
values = np.arange(16, dtype="<f8") * 1.5
|
||||
with h5py.File(path, "w", libver="earliest") as f:
|
||||
f.create_dataset("x", data=values)
|
||||
data_at = f["x"].id.get_offset()
|
||||
b = bytearray(path.read_bytes())
|
||||
# The layout message: the data's address, then its size.
|
||||
old = data_at.to_bytes(8, "little") + (values.nbytes).to_bytes(8, "little")
|
||||
assert b.count(old) == 1
|
||||
at = b.index(old)
|
||||
b[at:at + 8] = (data_at + FAR_SHIFT).to_bytes(8, "little")
|
||||
length = len(b) + FAR_SHIFT
|
||||
b[40:48] = length.to_bytes(8, "little")
|
||||
path.write_bytes(bytes(b))
|
||||
json.dump({"data_at": data_at, "far_at": data_at + FAR_SHIFT,
|
||||
"nbytes": values.nbytes, "length": length,
|
||||
"values": [float(x) for x in values]},
|
||||
open(out / "far.json", "w"))
|
||||
|
||||
|
||||
# far.h5's data moves this far: past 2 GiB, below 4 GiB.
|
||||
FAR_SHIFT = 3 << 30
|
||||
|
||||
write_limits(out)
|
||||
|
||||
|
||||
def write_big(path, megabytes):
|
||||
"""A large file for the range-request tests (`openUrl`): `/big`, about
|
||||
`megabytes` MB of float64 in 1 MiB chunks, written after a small
|
||||
dataset and a group, so listing and reading `/small` touch a few blocks
|
||||
of the file and a window of `/big` one chunk. Returns what h5py reads
|
||||
back."""
|
||||
n = megabytes * 1_000_000 // 8
|
||||
chunk = 1 << 17
|
||||
with h5py.File(path, "w") as f:
|
||||
f.attrs["note"] = "large file for range reads"
|
||||
f.create_dataset("small", data=np.array([1.5, -2.0, 3.25]))
|
||||
g = f.create_group("meta")
|
||||
g.attrs["units"] = "m"
|
||||
g.create_dataset("ids", data=np.arange(10, dtype="<i4"))
|
||||
big = f.create_dataset("big", shape=(n,), dtype="<f8", chunks=(chunk,))
|
||||
for s in range(0, n, 1 << 22):
|
||||
e = min(n, s + (1 << 22))
|
||||
big[s:e] = np.arange(s, e, dtype="<f8") * 0.5
|
||||
start = n // 2 + 12_345
|
||||
with h5py.File(path, "r") as f:
|
||||
return {
|
||||
"size": path.stat().st_size,
|
||||
"list": {"groups": ["meta"], "datasets": ["big", "small"]},
|
||||
"small": [float(x) for x in f["small"][()]],
|
||||
"ids": [str(int(x)) for x in f["meta/ids"][()]],
|
||||
"big_shape": list(f["big"].shape),
|
||||
"window": {"start": start, "count": 10,
|
||||
"values": [float(x) for x in f["big"][start:start + 10]]},
|
||||
}
|
||||
|
||||
|
||||
big_mb = int(os.environ.get("WASM_BIG_MB", "0"))
|
||||
if big_mb > 0:
|
||||
json.dump(write_big(out / "big.h5", big_mb), open(out / "big.json", "w"), indent=1)
|
||||
|
||||
@@ -1,11 +1,17 @@
|
||||
#!/usr/bin/env bash
|
||||
# Build the wasm package (../build.sh), test it under Node against files
|
||||
# written by h5py and netCDF4 (make_fixture.py), then load the viewer page
|
||||
# in headless Chromium if one is found (browser.sh).
|
||||
# written by h5py and netCDF4 (make_fixture.py) — opened from bytes, and
|
||||
# opened by URL from a local range-capable HTTP server (serve.py) — then
|
||||
# load the viewer page in headless Chromium if one is found (browser.sh).
|
||||
#
|
||||
# Needs node, the wasm-bindgen CLI (see ../build.sh) and a Python with h5py,
|
||||
# netCDF4 and numpy: CLAWHDF5_PYTHON names it (default python3). Without that
|
||||
# Python the test is skipped, unless CLAWHDF5_REQUIRE_INTEROP=1.
|
||||
#
|
||||
# WASM_BIG_MB (default 200) sizes the large file of the range-request
|
||||
# budget test (0 leaves it out); it is written under TMPDIR.
|
||||
# CLAWHDF5_WASM_CORPUS=DIR also compares every HDF5 file under DIR (up to
|
||||
# 16 MiB) read over HTTP with the same file read from bytes.
|
||||
set -euo pipefail
|
||||
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
@@ -23,9 +29,24 @@ fi
|
||||
|
||||
bash "$HERE/../build.sh"
|
||||
fix="$(mktemp -d)"
|
||||
trap 'rm -rf "$fix"' EXIT
|
||||
"$PY" "$HERE/make_fixture.py" "$fix"
|
||||
node "$HERE/test.mjs" "$HERE/../pkg" "$fix"
|
||||
server=""
|
||||
cleanup() {
|
||||
[ -n "$server" ] && kill "$server" 2>/dev/null || true
|
||||
rm -rf "$fix"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
WASM_BIG_MB="${WASM_BIG_MB:-200}" "$PY" "$HERE/make_fixture.py" "$fix"
|
||||
|
||||
# The fixtures (and the corpus) over HTTP with range support.
|
||||
roots=(--root "fix=$fix")
|
||||
[ -n "${CLAWHDF5_WASM_CORPUS:-}" ] && roots+=(--root "corpus=$CLAWHDF5_WASM_CORPUS")
|
||||
"$PY" "$HERE/serve.py" "${roots[@]}" > "$fix/port" &
|
||||
server=$!
|
||||
for _ in $(seq 50); do
|
||||
[ -s "$fix/port" ] && break
|
||||
sleep 0.1
|
||||
done
|
||||
node "$HERE/test.mjs" "$HERE/../pkg" "$fix" "http://127.0.0.1:$(head -1 "$fix/port")"
|
||||
|
||||
# The page itself, in headless Chromium when one is available.
|
||||
status=0
|
||||
|
||||
@@ -0,0 +1,223 @@
|
||||
"""A static HTTP server for the wasm tests, with HTTP Range support and
|
||||
request counting.
|
||||
|
||||
python serve.py [--root PREFIX=DIR ...]
|
||||
|
||||
Serves each DIR under URL PREFIX (the first match wins; PREFIX "" is the
|
||||
site root), prints the port on its first line of stdout, and runs until
|
||||
killed. No symlinks or copies are made: files are read where they are.
|
||||
|
||||
- `Range: bytes=a-b`, `bytes=a-` and `bytes=-n` get 206 with Content-Range,
|
||||
an unsatisfiable range 416; every file answer carries an ETag, and CORS
|
||||
headers exposing Content-Range, so a page on another origin can use it.
|
||||
- Under `/norange/...` the same files are served but Range is ignored
|
||||
(200 with the whole file), as by a server without range support.
|
||||
- Under `/noexpose/...` ranges are served without Content-Range, ETag,
|
||||
Last-Modified or Accept-Ranges, and CORS exposes none of them: what a
|
||||
page sees of a cross-origin server that does not list them in
|
||||
Access-Control-Expose-Headers. The length comes from Content-Length of
|
||||
a HEAD request (always readable).
|
||||
- Under `/unexposed/...` ranges are served with all those headers, but
|
||||
CORS exposes none of them: a browser page on another origin cannot read
|
||||
them (test/browser.sh loads the page from 127.0.0.1 and the file from
|
||||
localhost), so it has to take the same path.
|
||||
- `GET /__stats` returns `{"requests": n, "bytes": n, "log": [...]}` for
|
||||
file requests since the last `GET /__reset`, which zeroes them.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import posixpath
|
||||
import sys
|
||||
import threading
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
from urllib.parse import unquote, urlsplit
|
||||
|
||||
TYPES = {
|
||||
".html": "text/html; charset=utf-8",
|
||||
".js": "text/javascript; charset=utf-8",
|
||||
".mjs": "text/javascript; charset=utf-8",
|
||||
".wasm": "application/wasm",
|
||||
".json": "application/json",
|
||||
".ts": "text/plain; charset=utf-8",
|
||||
}
|
||||
|
||||
lock = threading.Lock()
|
||||
stats = {"requests": 0, "bytes": 0, "log": []}
|
||||
|
||||
|
||||
def resolve(roots, path):
|
||||
"""The file for URL `path`, or None. `..` never leaves a root."""
|
||||
parts = [p for p in posixpath.normpath(unquote(path)).split("/") if p]
|
||||
if any(p in (".", "..") for p in parts):
|
||||
return None
|
||||
for prefix, root in roots:
|
||||
pre = [p for p in prefix.split("/") if p]
|
||||
if parts[: len(pre)] == pre:
|
||||
rest = parts[len(pre):] or ["index.html"]
|
||||
f = os.path.join(root, *rest)
|
||||
if os.path.isfile(f):
|
||||
return f
|
||||
return None
|
||||
|
||||
|
||||
def parse_range(header, size):
|
||||
"""(start, end exclusive) for a single `bytes=` range, "bad" when
|
||||
unsatisfiable, None when absent or unparsable (served whole)."""
|
||||
if not header or not header.startswith("bytes=") or "," in header:
|
||||
return None
|
||||
a, _, b = header[len("bytes="):].strip().partition("-")
|
||||
try:
|
||||
if a == "":
|
||||
n = int(b)
|
||||
return (max(0, size - n), size) if n > 0 and size > 0 else "bad"
|
||||
start = int(a)
|
||||
end = int(b) + 1 if b else size
|
||||
except ValueError:
|
||||
return None
|
||||
if start >= size or end <= start:
|
||||
return "bad"
|
||||
return start, min(end, size)
|
||||
|
||||
|
||||
def make_handler(roots):
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
protocol_version = "HTTP/1.1"
|
||||
|
||||
def log_message(self, *args):
|
||||
pass
|
||||
|
||||
def cors(self, expose=True):
|
||||
self.send_header("Access-Control-Allow-Origin", "*")
|
||||
if expose:
|
||||
self.send_header("Access-Control-Expose-Headers",
|
||||
"Content-Range, Content-Length, ETag, Accept-Ranges")
|
||||
|
||||
def do_OPTIONS(self):
|
||||
self.send_response(204)
|
||||
self.cors()
|
||||
self.send_header("Access-Control-Allow-Headers", "Range")
|
||||
self.send_header("Content-Length", "0")
|
||||
self.end_headers()
|
||||
|
||||
def do_HEAD(self):
|
||||
self.serve(head=True)
|
||||
|
||||
def do_GET(self):
|
||||
self.serve(head=False)
|
||||
|
||||
def json(self, obj):
|
||||
body = json.dumps(obj).encode()
|
||||
self.send_response(200)
|
||||
self.cors()
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(body)))
|
||||
self.send_header("Cache-Control", "no-store")
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
|
||||
def serve(self, head):
|
||||
path = urlsplit(self.path).path
|
||||
if path == "/__stats":
|
||||
with lock:
|
||||
return self.json(stats)
|
||||
if path == "/__reset":
|
||||
with lock:
|
||||
stats.update(requests=0, bytes=0, log=[])
|
||||
return self.json({})
|
||||
# ranges: honour Range; send_all: send Content-Range, ETag and
|
||||
# Accept-Ranges; expose: list them for CORS.
|
||||
ranges = send_all = expose = True
|
||||
if path.startswith("/norange/"):
|
||||
ranges = False
|
||||
path = path[len("/norange"):]
|
||||
elif path.startswith("/noexpose/"):
|
||||
expose = send_all = False
|
||||
path = path[len("/noexpose"):]
|
||||
elif path.startswith("/unexposed/"):
|
||||
expose = False
|
||||
path = path[len("/unexposed"):]
|
||||
f = resolve(roots, path)
|
||||
if f is None:
|
||||
self.send_response(404)
|
||||
self.cors()
|
||||
self.send_header("Content-Length", "0")
|
||||
self.end_headers()
|
||||
return
|
||||
size = os.path.getsize(f)
|
||||
st = os.stat(f)
|
||||
etag = '"%s"' % hashlib.sha1(
|
||||
f"{f}:{size}:{st.st_mtime_ns}".encode()).hexdigest()[:16]
|
||||
r = parse_range(self.headers.get("Range"), size) if ranges else None
|
||||
if r == "bad":
|
||||
self.send_response(416)
|
||||
self.cors()
|
||||
self.send_header("Content-Range", f"bytes */{size}")
|
||||
self.send_header("Content-Length", "0")
|
||||
self.end_headers()
|
||||
return
|
||||
start, end = r if r else (0, size)
|
||||
self.send_response(206 if r else 200)
|
||||
self.cors(expose)
|
||||
ext = os.path.splitext(f)[1]
|
||||
self.send_header("Content-Type", TYPES.get(ext, "application/octet-stream"))
|
||||
self.send_header("Content-Length", str(end - start))
|
||||
self.send_header("Cache-Control", "no-store")
|
||||
if send_all:
|
||||
self.send_header("ETag", etag)
|
||||
if ranges:
|
||||
self.send_header("Accept-Ranges", "bytes")
|
||||
if r:
|
||||
self.send_header("Content-Range", f"bytes {start}-{end - 1}/{size}")
|
||||
self.end_headers()
|
||||
with lock:
|
||||
# A HEAD is a request too (openUrl makes one when it cannot
|
||||
# see Content-Range); it sends no bytes.
|
||||
stats["requests"] += 1
|
||||
sent = 0 if head else end - start
|
||||
stats["bytes"] += sent
|
||||
stats["log"].append([path, start, end, 206 if r else 200, "HEAD" if head else "GET"])
|
||||
if not head:
|
||||
with open(f, "rb") as fh:
|
||||
fh.seek(start)
|
||||
left = end - start
|
||||
try:
|
||||
while left:
|
||||
buf = fh.read(min(left, 1 << 20))
|
||||
if not buf:
|
||||
break
|
||||
self.wfile.write(buf)
|
||||
left -= len(buf)
|
||||
except (BrokenPipeError, ConnectionResetError):
|
||||
pass
|
||||
|
||||
return Handler
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--root", action="append", default=[],
|
||||
help="PREFIX=DIR: serve DIR under URL PREFIX")
|
||||
args = ap.parse_args()
|
||||
roots = []
|
||||
for spec in args.root:
|
||||
prefix, _, d = spec.partition("=")
|
||||
roots.append((prefix, os.path.abspath(d)))
|
||||
class Server(ThreadingHTTPServer):
|
||||
def handle_error(self, request, client_address):
|
||||
# A client that drops a connection (a cancelled download) is
|
||||
# not an error of the server.
|
||||
if not isinstance(sys.exc_info()[1], (ConnectionError, TimeoutError)):
|
||||
super().handle_error(request, client_address)
|
||||
|
||||
httpd = Server(("127.0.0.1", 0), make_handler(roots))
|
||||
httpd.daemon_threads = True
|
||||
print(httpd.server_address[1], flush=True)
|
||||
sys.stdout.close()
|
||||
httpd.serve_forever()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,20 +1,31 @@
|
||||
// Node test of the built wasm package (the exact pkg/ the viewer page loads)
|
||||
// and the viewer's DOM-free helpers. Run by test/run.sh:
|
||||
// node test.mjs PKG_DIR FIXTURE_DIR
|
||||
// node test.mjs PKG_DIR FIXTURE_DIR [SERVER_URL]
|
||||
// FIXTURE_DIR holds fixture.h5, fixture.nc and expected.json from
|
||||
// make_fixture.py (values as libhdf5 reads them back).
|
||||
// make_fixture.py (values as libhdf5 reads them back), and big.h5/big.json
|
||||
// when it was run with WASM_BIG_MB. SERVER_URL is test/serve.py serving
|
||||
// FIXTURE_DIR under /fix (and $CLAWHDF5_WASM_CORPUS under /corpus): with
|
||||
// it, every check is repeated on files opened with openUrl (HTTP range
|
||||
// requests), and the request budget, the full-download fallback and the
|
||||
// error paths of openUrl are tested.
|
||||
import assert from "node:assert/strict";
|
||||
import { readFileSync } from "node:fs";
|
||||
import { join } from "node:path";
|
||||
import { existsSync, readFileSync, readdirSync, statSync } from "node:fs";
|
||||
import { join, relative } from "node:path";
|
||||
import { pathToFileURL } from "node:url";
|
||||
|
||||
const [pkgDir, fixDir] = process.argv.slice(2);
|
||||
const [pkgDir, fixDir, base] = process.argv.slice(2);
|
||||
const pkg = await import(pathToFileURL(join(pkgDir, "clawhdf5_wasm.js")));
|
||||
pkg.initSync({ module: readFileSync(join(pkgDir, "clawhdf5_wasm_bg.wasm")) });
|
||||
const lib = await import(pathToFileURL(join(import.meta.dirname, "..", "viewer-lib.js")));
|
||||
|
||||
let checks = 0;
|
||||
const eq = (a, b, msg) => { assert.deepEqual(a, b, msg); checks++; };
|
||||
// A call that must fail: a thrown Error (open) or a rejected promise
|
||||
// (openUrl), whose message matches `re`.
|
||||
const fails = async (fn, re, msg) => {
|
||||
await assert.rejects(async () => fn(), (e) => e instanceof Error && re.test(e.message), msg);
|
||||
checks++;
|
||||
};
|
||||
|
||||
const ARRAY_TYPES = {
|
||||
f32: Float32Array, f64: Float64Array, i8: Int8Array, i16: Int16Array, i32: Int32Array,
|
||||
@@ -54,13 +65,13 @@ function checkAttr(ctx, a, want) {
|
||||
assert.fail(`${ctx}: unknown expectation ${JSON.stringify(want)}`);
|
||||
}
|
||||
|
||||
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
|
||||
for (const [name, exp] of Object.entries(expected)) {
|
||||
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
|
||||
|
||||
// Every listing, dataset, error and attribute of `file` against what
|
||||
// libhdf5 reads (expected.json). `file` is an H5File (synchronous methods)
|
||||
// or a RemoteFile (promises): every call is awaited.
|
||||
async function checkFile(name, exp, file) {
|
||||
for (const [path, want] of Object.entries(exp.lists)) {
|
||||
eq(file.kind(path), "group", `${name}:${path} kind`);
|
||||
const list = file.list(path);
|
||||
eq(await file.kind(path), "group", `${name}:${path} kind`);
|
||||
const list = await file.list(path);
|
||||
for (const [kind, key] of [["group", "groups"], ["dataset", "datasets"]]) {
|
||||
eq(list.filter((c) => c.kind === kind).map((c) => c.name).sort(), want[key], `${name}:${path} ${key}`);
|
||||
}
|
||||
@@ -68,58 +79,67 @@ for (const [name, exp] of Object.entries(expected)) {
|
||||
|
||||
for (const [path, want] of Object.entries(exp.datasets)) {
|
||||
const ctx = `${name}:${path}`;
|
||||
eq(file.kind(path), "dataset", `${ctx} kind`);
|
||||
eq(await file.kind(path), "dataset", `${ctx} kind`);
|
||||
if (want.unavailable) {
|
||||
// The wasm build has no zstd (it links C): a clear error, no data.
|
||||
assert.throws(() => file.read(path), (e) => e.message.includes(want.unavailable), ctx);
|
||||
checks++;
|
||||
await fails(() => file.read(path), new RegExp(want.unavailable), ctx);
|
||||
continue;
|
||||
}
|
||||
const info = file.info(path);
|
||||
const info = await file.info(path);
|
||||
eq([...info.shape, ...info.elementShape], want.shape, `${ctx} info shape`);
|
||||
const r = file.read(path);
|
||||
const r = await file.read(path);
|
||||
eq(r.shape, want.shape, `${ctx} shape`);
|
||||
eq(r.dtype, info.dtype, `${ctx} dtype`);
|
||||
assert.ok(r.data instanceof ARRAY_TYPES[want.kind], `${ctx}: ${r.data.constructor.name} for ${want.kind}`);
|
||||
eq(values(want.kind, r.data), want.values, ctx);
|
||||
if (want.slab) {
|
||||
const s = want.slab;
|
||||
const part = file.readHyperslab(path, s.start, s.count, s.stride);
|
||||
const part = await file.readHyperslab(path, s.start, s.count, s.stride);
|
||||
eq(part.shape, s.shape, `${ctx} slab shape`);
|
||||
eq(values(want.kind, part.data), s.values, `${ctx} slab`);
|
||||
}
|
||||
}
|
||||
|
||||
for (const [path, what] of Object.entries(exp.errors)) {
|
||||
assert.throws(() => file.read(path), (e) => e instanceof Error && e.message.includes(what), `${name}:${path}`);
|
||||
checks++;
|
||||
await fails(() => file.read(path), new RegExp(what), `${name}:${path}`);
|
||||
}
|
||||
|
||||
for (const [path, want] of Object.entries(exp.attrs)) {
|
||||
const attrs = file.attrs(path);
|
||||
eq(file.attrErrors(path), [], `${name}:${path} attr errors`);
|
||||
const attrs = await file.attrs(path);
|
||||
eq(await file.attrErrors(path), [], `${name}:${path} attr errors`);
|
||||
const seen = attrs.filter((a) => !a.name.startsWith("_") && !exp.skip_attrs.includes(a.name));
|
||||
eq(seen.map((a) => a.name).sort(), Object.keys(want).sort(), `${name}:${path} attr names`);
|
||||
for (const a of seen) checkAttr(`${name}:${path}@${a.name}`, a, want[a.name]);
|
||||
}
|
||||
}
|
||||
|
||||
// The error paths of the reader, for an H5File or a RemoteFile of
|
||||
// fixture.h5.
|
||||
async function checkErrors(h5) {
|
||||
await fails(() => h5.read("/nope"), /./);
|
||||
await fails(() => h5.list("/grid"), /not a group/);
|
||||
await fails(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
|
||||
await fails(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
|
||||
await fails(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
|
||||
await fails(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
|
||||
// Big integers stay exact.
|
||||
eq((await h5.read("/u64")).data[0], 18446744073709551615n, "u64 max");
|
||||
}
|
||||
|
||||
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
|
||||
for (const [name, exp] of Object.entries(expected)) {
|
||||
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
|
||||
await checkFile(name, exp, file);
|
||||
file.free();
|
||||
}
|
||||
|
||||
// Errors reach JavaScript as thrown Errors, never as data.
|
||||
const h5 = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.h5"))));
|
||||
const throwsMsg = (fn, re) => { assert.throws(fn, (e) => e instanceof Error && re.test(e.message)); checks++; };
|
||||
throwsMsg(() => pkg.open(new Uint8Array(64)), /./);
|
||||
throwsMsg(() => h5.read("/nope"), /./);
|
||||
throwsMsg(() => h5.list("/grid"), /not a group/);
|
||||
throwsMsg(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
|
||||
throwsMsg(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
|
||||
throwsMsg(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
|
||||
throwsMsg(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
|
||||
await fails(() => pkg.open(new Uint8Array(64)), /./);
|
||||
await checkErrors(h5);
|
||||
// Info for a dataset with an unlimited dimension (netCDF "time").
|
||||
const nc = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.nc"))));
|
||||
eq(nc.info("/time").maxshape, [null], "unlimited dimension is null");
|
||||
// Big integers stay exact.
|
||||
eq(h5.read("/u64").data[0], 18446744073709551615n, "u64 max");
|
||||
eq(typeof pkg.version(), "string", "version");
|
||||
|
||||
// Viewer helpers.
|
||||
@@ -137,7 +157,486 @@ eq(lib.toRows(pairs.data, 2, 1, lib.perElement([2])), [["[2, 3]"], ["[4, 5]"]],
|
||||
eq(lib.formatValue(0.1 + 0.2), "0.3", "float formatting");
|
||||
eq(lib.formatValue(2n ** 64n - 1n), "18446744073709551615", "bigint formatting");
|
||||
eq(lib.formatValue("x"), '"x"', "string formatting");
|
||||
eq(lib.formatBytes(0), "0 B", "bytes");
|
||||
eq(lib.formatBytes(1536), "1.5 KiB", "KiB");
|
||||
eq(lib.formatBytes(200 * 1024 * 1024), "200 MiB", "MiB");
|
||||
eq(lib.formatStats({ lazy: true, requests: 3, bytesFetched: 2 << 20, size: 200 << 20 }),
|
||||
"3 requests, 2 MiB of 200 MiB fetched (1.0%)", "stats line");
|
||||
eq(lib.formatStats({ lazy: false, requests: 1, bytesFetched: 1024, size: 1024 }),
|
||||
"downloaded whole (1 KiB): the server does not support range requests", "stats line, no ranges");
|
||||
h5.free();
|
||||
nc.free();
|
||||
|
||||
console.log(`wasm package: ${checks} checks passed`);
|
||||
|
||||
if (base) await remoteTests();
|
||||
|
||||
async function serverStats() {
|
||||
return (await fetch(`${base}/__stats`)).json();
|
||||
}
|
||||
|
||||
async function remoteTests() {
|
||||
const before = checks;
|
||||
await fetch(`${base}/__reset`);
|
||||
|
||||
// The same checks over HTTP range requests, at the default block size
|
||||
// and at 512-byte blocks with a 4 KiB cache (almost every structure read
|
||||
// a miss, evictions between calls).
|
||||
for (const opts of [undefined, { blockSize: 512, cacheSize: 4096 }]) {
|
||||
for (const [name, exp] of Object.entries(expected)) {
|
||||
const f = await pkg.openUrl(`${base}/fix/${name}`, opts);
|
||||
await checkFile(`${name} (openUrl ${JSON.stringify(opts ?? {})})`, exp, f);
|
||||
const st = f.stats();
|
||||
eq(st.lazy, true, "read by ranges");
|
||||
eq(st.size, statSync(join(fixDir, name)).size, "size");
|
||||
f.free();
|
||||
}
|
||||
}
|
||||
const remote = await pkg.openUrl(`${base}/fix/fixture.h5`);
|
||||
await checkErrors(remote);
|
||||
eq((await (await pkg.openUrl(`${base}/fix/fixture.nc`)).info("/time")).maxshape, [null], "remote unlimited");
|
||||
|
||||
// What the page counts is what the server served.
|
||||
await fetch(`${base}/__reset`);
|
||||
const counted = await pkg.openUrl(`${base}/fix/fixture.nc`, { blockSize: 1024 });
|
||||
await counted.read("/temp");
|
||||
const server = await serverStats();
|
||||
eq(counted.stats().requests, server.requests, "requests counted");
|
||||
eq(counted.stats().bytesFetched, server.bytes, "bytes counted");
|
||||
|
||||
// A custom fetch is used for every request, with the caller's headers,
|
||||
// given in any form fetch takes (a caller's Range is not sent).
|
||||
for (const headers of [{ "X-Test": "1" }, new Headers({ "X-Test": "1", Range: "bytes=0-0" }), [["X-Test", "1"]]]) {
|
||||
let calls = 0;
|
||||
const seen = new Set();
|
||||
const viaCustom = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||
blockSize: 4096,
|
||||
headers,
|
||||
fetch: (url, init) => {
|
||||
calls++;
|
||||
const h = new Headers(init.headers);
|
||||
seen.add(`${h.get("X-Test")} ${h.get("Range").startsWith("bytes=0-0") ? "caller's range" : "ours"}`);
|
||||
return fetch(url, init);
|
||||
},
|
||||
});
|
||||
const kind = headers.constructor.name;
|
||||
eq((await viaCustom.read("/sensors/temp")).data[0], 21.5, `custom fetch values (${kind})`);
|
||||
eq(calls, viaCustom.stats().requests, `custom fetch calls (${kind})`);
|
||||
eq([...seen], ["1 ours"], `headers passed (${kind})`);
|
||||
}
|
||||
|
||||
// parallel must be a positive integer.
|
||||
for (const parallel of [0, -1, 1.5, "4", NaN, 5000]) {
|
||||
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { parallel }), /parallel/, `parallel ${String(parallel)}`);
|
||||
}
|
||||
|
||||
// When one range request fails, the others in flight are aborted and no
|
||||
// more start (fetchRanges, called directly: remote.js is in the package).
|
||||
{
|
||||
const snippets = join(pkgDir, "snippets");
|
||||
const dir = readdirSync(snippets).find((d) => existsSync(join(snippets, d, "js", "remote.js")));
|
||||
const remote = await import(pathToFileURL(join(snippets, dir, "js", "remote.js")));
|
||||
let started = 0;
|
||||
const aborted = [];
|
||||
const slowFetch = (url, init) => {
|
||||
const k = started++;
|
||||
if (k === 2) return Promise.resolve(new Response(null, { status: 500 }));
|
||||
return new Promise((resolve, reject) => {
|
||||
const t = setTimeout(() => resolve(new Response(new Uint8Array(10), { status: 206 })), 200);
|
||||
init.signal?.addEventListener("abort", () => {
|
||||
clearTimeout(t);
|
||||
aborted.push(k);
|
||||
reject(new DOMException("aborted", "AbortError"));
|
||||
});
|
||||
});
|
||||
};
|
||||
const ranges = Array.from({ length: 20 }, (_, i) => [i * 10, i * 10 + 10]).flat();
|
||||
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: 3 }, null, 200),
|
||||
/HTTP 500/, "a failed range request is the error");
|
||||
await new Promise((r) => setTimeout(r, 300));
|
||||
eq(started, 3, "no request starts after a failure");
|
||||
eq(aborted.sort(), [0, 1], "requests in flight are aborted");
|
||||
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: "x" }, null, 200),
|
||||
/parallel must be a positive integer/, "fetchRanges checks parallel");
|
||||
}
|
||||
|
||||
// Calls in flight at once share the cache (a 1 KiB budget: nothing is
|
||||
// evicted while any of them runs) and each gets its own answer.
|
||||
const both = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, cacheSize: 1024 });
|
||||
const [g, t, l, a] = await Promise.all([
|
||||
both.read("/grid"), both.read("/sensors/temp"), both.list("/sensors"), both.attrs("/"),
|
||||
]);
|
||||
const exp = expected["fixture.h5"];
|
||||
eq(Array.from(g.data), exp.datasets["/grid"].values, "concurrent /grid");
|
||||
eq(Array.from(t.data), exp.datasets["/sensors/temp"].values, "concurrent /sensors/temp");
|
||||
eq(l.map((c) => c.name).sort(), [...exp.lists["/sensors"].groups, ...exp.lists["/sensors"].datasets].sort(), "concurrent list");
|
||||
eq(a.length > 0, true, "concurrent attrs");
|
||||
assert.ok(both.stats().cachedBytes <= 1024, "trimmed to the budget when idle");
|
||||
|
||||
// Listing and reading small things of a large file fetches a few blocks,
|
||||
// not the file.
|
||||
if (existsSync(join(fixDir, "big.json"))) {
|
||||
const big = JSON.parse(readFileSync(join(fixDir, "big.json"), "utf8"));
|
||||
await fetch(`${base}/__reset`);
|
||||
const f = await pkg.openUrl(`${base}/fix/big.h5`);
|
||||
const list = await f.list("/");
|
||||
eq(list.filter((c) => c.kind === "group").map((c) => c.name), big.list.groups, "big: groups");
|
||||
eq(list.filter((c) => c.kind === "dataset").map((c) => c.name).sort(), big.list.datasets, "big: datasets");
|
||||
eq(Array.from((await f.read("/small")).data), big.small, "big: /small");
|
||||
eq(Array.from((await f.read("/meta/ids")).data, String), big.ids, "big: /meta/ids");
|
||||
eq((await f.info("/big")).shape, big.big_shape, "big: shape");
|
||||
eq((await f.attrs("/meta"))[0].value, "m", "big: group attribute");
|
||||
const win = await f.readHyperslab("/big", [big.window.start], [big.window.count]);
|
||||
eq(Array.from(win.data), big.window.values, "big: window");
|
||||
const st = f.stats();
|
||||
const server = await serverStats();
|
||||
eq(st.requests, server.requests, "big: requests counted");
|
||||
eq(st.bytesFetched, server.bytes, "big: bytes counted");
|
||||
console.log(`big.h5 (${big.size} bytes): listed, 2 small reads, attributes, info and a window in ` +
|
||||
`${server.requests} requests, ${server.bytes} bytes (${(100 * server.bytes / big.size).toFixed(2)}%)`);
|
||||
assert.ok(server.requests <= 8, `big: ${server.requests} requests`);
|
||||
assert.ok(server.bytes * 20 < big.size, `big: ${server.bytes} bytes fetched`);
|
||||
checks += 2;
|
||||
}
|
||||
|
||||
// A cross-origin server that does not expose Content-Range, ETag or
|
||||
// Last-Modified (serve.py's /noexpose/): the length comes from a HEAD
|
||||
// request and answers are checked by their length alone. Every fixture
|
||||
// check, calls in flight at once, and what the page counts.
|
||||
for (const opts of [undefined, { blockSize: 512, cacheSize: 1024 }]) {
|
||||
for (const [name, exp] of Object.entries(expected)) {
|
||||
await fetch(`${base}/__reset`);
|
||||
const f = await pkg.openUrl(`${base}/noexpose/fix/${name}`, opts);
|
||||
await checkFile(`${name} (no exposed headers, ${JSON.stringify(opts ?? {})})`, exp, f);
|
||||
const st = f.stats();
|
||||
eq(st.lazy, true, "no exposed headers: read by ranges");
|
||||
eq(st.size, statSync(join(fixDir, name)).size, "no exposed headers: size from HEAD");
|
||||
const server = await serverStats();
|
||||
eq(server.log.filter((l) => l[4] === "HEAD").length, 1, "no exposed headers: one HEAD");
|
||||
eq(st.requests, server.requests, "no exposed headers: requests counted");
|
||||
eq(st.bytesFetched, server.bytes, "no exposed headers: bytes counted");
|
||||
f.free();
|
||||
}
|
||||
}
|
||||
{
|
||||
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, { blockSize: 512, cacheSize: 0, parallel: 3 });
|
||||
const exp = expected["fixture.h5"];
|
||||
const paths = ["/grid", "/sensors/temp", "/vlen_str", "/cube"];
|
||||
const got = await Promise.all([...paths, ...paths].map((p) => f.read(p)));
|
||||
got.forEach((r, i) => eq(values(exp.datasets[paths[i % 4]].kind, r.data), exp.datasets[paths[i % 4]].values,
|
||||
`no exposed headers: concurrent ${paths[i % 4]}`));
|
||||
await checkErrors(f);
|
||||
}
|
||||
// Without a validator a changed file cannot be told apart; a short
|
||||
// answer still can.
|
||||
await fails(async () => {
|
||||
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
|
||||
blockSize: 512,
|
||||
fetch: async (url, init) => {
|
||||
const r = await fetch(url, init);
|
||||
if (init.method === "HEAD" || init.headers.Range === "bytes=0-511") return r;
|
||||
return new Response((await r.arrayBuffer()).slice(1), { status: 206 });
|
||||
},
|
||||
});
|
||||
await f.read("/grid");
|
||||
}, /got \d+/, "no exposed headers: short answer");
|
||||
// No HEAD length either: a clear error.
|
||||
await fails(() => pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
|
||||
fetch: async (url, init) => (init.method === "HEAD" ? new Response(null, { status: 405 }) : fetch(url, init)),
|
||||
}), /cannot learn the file's size/, "no exposed headers, no HEAD");
|
||||
|
||||
// A server without range support: downloaded whole (the default), or
|
||||
// refused.
|
||||
const whole = await pkg.openUrl(`${base}/norange/fix/fixture.h5`);
|
||||
await checkFile("fixture.h5 (no range support)", expected["fixture.h5"], whole);
|
||||
eq(whole.stats().lazy, false, "downloaded whole");
|
||||
eq(whole.stats().requests, 1, "one request");
|
||||
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fallback: "error" }),
|
||||
/does not support HTTP range requests/, "fallback: error");
|
||||
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000 }),
|
||||
/more than maxDownload/, "maxDownload");
|
||||
// Without a Content-Length the body is streamed, and stopped at the limit.
|
||||
const undeclared = async (url, init) => {
|
||||
const r = await fetch(url, init);
|
||||
return new Response(r.body, { status: r.status });
|
||||
};
|
||||
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000, fetch: undeclared }),
|
||||
/more than maxDownload/, "maxDownload, streamed");
|
||||
const streamed = await pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fetch: undeclared });
|
||||
eq(Array.from((await streamed.read("/sensors/temp")).data), [21.5, 22, 22.25], "streamed download");
|
||||
|
||||
// Errors: HTTP status, not HDF5, bad options, a file that changes, a
|
||||
// server that answers with the wrong bytes.
|
||||
await fails(() => pkg.openUrl(`${base}/fix/missing.h5`), /HTTP 404/, "404");
|
||||
await fails(() => pkg.openUrl(`${base}/fix/expected.json`), /./, "not HDF5");
|
||||
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 100 }), /blockSize/, "blockSize");
|
||||
const tamper = (edit) => async (url, init) => {
|
||||
const r = await fetch(url, init);
|
||||
return init.headers.Range === "bytes=0-511" ? r : edit(r);
|
||||
};
|
||||
const withHeaders = async (r, headers) => {
|
||||
const h = new Headers(r.headers);
|
||||
for (const [k, v] of Object.entries(headers)) h.set(k, v);
|
||||
return new Response(await r.arrayBuffer(), { status: r.status, headers: h });
|
||||
};
|
||||
await fails(async () => {
|
||||
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, fetch: tamper((r) => withHeaders(r, { ETag: '"other"' })) });
|
||||
await f.read("/grid");
|
||||
}, /changed on the server/, "changed file");
|
||||
await fails(async () => {
|
||||
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||
blockSize: 512,
|
||||
fetch: tamper(async (r) => new Response((await r.arrayBuffer()).slice(1), { status: 206 })),
|
||||
});
|
||||
await f.read("/grid");
|
||||
}, /got \d+/, "short answer");
|
||||
await fails(async () => {
|
||||
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||
blockSize: 512,
|
||||
fetch: tamper(async (r) => new Response(await r.arrayBuffer(), { status: 200 })),
|
||||
});
|
||||
await f.read("/grid");
|
||||
}, /stopped honouring range requests/, "200 mid-file");
|
||||
await fails(async () => {
|
||||
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||
blockSize: 512,
|
||||
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
|
||||
});
|
||||
await f.read("/grid");
|
||||
}, /the server sent 0-511/, "wrong range");
|
||||
await fails(async () => {
|
||||
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||
blockSize: 512,
|
||||
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes */25752" })),
|
||||
});
|
||||
await f.read("/grid");
|
||||
}, /unusable Content-Range/, "unusable Content-Range");
|
||||
|
||||
await limitTests();
|
||||
await floodTests();
|
||||
|
||||
// Corpus files: what the viewer can show of each is the same read by
|
||||
// ranges as in memory (an error wherever it gives one).
|
||||
const corpus = process.env.CLAWHDF5_WASM_CORPUS;
|
||||
if (corpus) await corpusTests(corpus);
|
||||
console.log(`openUrl: ${checks - before} checks passed`);
|
||||
}
|
||||
|
||||
// A fetch that serves `buf` as a file of `total` bytes (zeros past the end
|
||||
// of `buf`, and `extra` = [[offset, bytes], ...] laid over them), counting
|
||||
// its calls. `hide` leaves out Content-Range and the validators, as a
|
||||
// cross-origin server that exposes neither does; HEAD then gives the length.
|
||||
function mockFetch(buf, { total = buf.length, extra = [], hide = false } = {}) {
|
||||
const f = async (url, init) => {
|
||||
f.calls++;
|
||||
if (init.method === "HEAD") {
|
||||
return new Response(null, { status: 200, headers: { "Content-Length": String(total) } });
|
||||
}
|
||||
const m = /^bytes=(\d+)-(\d+)$/.exec(new Headers(init.headers).get("Range"));
|
||||
const a = Number(m[1]);
|
||||
const b = Math.min(Number(m[2]) + 1, total);
|
||||
const out = new Uint8Array(b - a);
|
||||
for (const [at, bytes] of [[0, buf], ...extra]) {
|
||||
const from = Math.max(a, at);
|
||||
const to = Math.min(b, at + bytes.length);
|
||||
if (from < to) out.set(bytes.subarray(from - at, to - at), from - a);
|
||||
}
|
||||
const headers = { "Content-Length": String(b - a) };
|
||||
if (!hide) headers["Content-Range"] = `bytes ${a}-${b - 1}/${total}`;
|
||||
return new Response(out, { status: 206, headers });
|
||||
};
|
||||
f.calls = 0;
|
||||
return f;
|
||||
}
|
||||
|
||||
// Sizes a hostile server or a large dataset can name are errors, never an
|
||||
// allocation that aborts the module (which would take every open file on
|
||||
// the page with it); see write_limits in make_fixture.py.
|
||||
async function limitTests() {
|
||||
// Read whole, /huge_u8 (2^28 + 1024 bytes) would take over 2 GiB while
|
||||
// decoding: refused before its chunks are fetched; a window reads.
|
||||
const n = 2 ** 28 + 1024;
|
||||
const limits = readFileSync(join(fixDir, "limits.h5"));
|
||||
const local = pkg.open(new Uint8Array(limits));
|
||||
await fails(() => local.read("/huge_u8"), /readHyperslab/, "huge read, in memory");
|
||||
const remote = await pkg.openUrl(`${base}/fix/limits.h5`);
|
||||
const before = remote.stats().requests;
|
||||
await fails(() => remote.read("/huge_u8"), /readHyperslab/, "huge read, openUrl");
|
||||
eq(remote.stats().requests, before, "huge read fetched nothing");
|
||||
eq(Array.from((await remote.readHyperslab("/huge_u8", [n - 4], [4])).data), [0, 0, 0, 7], "huge window");
|
||||
eq(Array.from(local.readHyperslab("/huge_u8", [n - 4], [4]).data), [0, 0, 0, 7], "huge window, in memory");
|
||||
local.free();
|
||||
|
||||
// A server claiming 3 GiB, a heap collection claiming 2 GiB + 4 KiB: an
|
||||
// error after a few requests (it used to fetch 2 GiB, then abort).
|
||||
const hostile = mockFetch(new Uint8Array(readFileSync(join(fixDir, "hostile_vl.h5"))), { total: 3 * 2 ** 30 });
|
||||
const h = await pkg.openUrl("http://hostile.invalid/h.h5", { fetch: hostile });
|
||||
await fails(() => h.read("/a"), /maxFetch/, "hostile collection size");
|
||||
assert.ok(hostile.calls <= 4, `hostile: ${hostile.calls} requests`);
|
||||
// The module survived: files open and read.
|
||||
eq((await (await pkg.openUrl(`${base}/fix/fixture.h5`)).read("/sensors/temp")).data[0], 21.5, "alive after hostile");
|
||||
|
||||
// A call that would fetch more than maxFetch fails before fetching it.
|
||||
await fails(async () => {
|
||||
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, maxFetch: 1024 });
|
||||
await f.read("/grid");
|
||||
}, /maxFetch/, "maxFetch");
|
||||
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxFetch: 2 ** 31 }), /maxFetch/, "maxFetch range");
|
||||
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxDownload: 2 ** 31 }), /maxDownload/, "maxDownload range");
|
||||
|
||||
// Offsets past 2 GiB work on wasm32: far.h5's data sits at 3 GiB.
|
||||
const far = JSON.parse(readFileSync(join(fixDir, "far.json"), "utf8"));
|
||||
const farBytes = new Uint8Array(readFileSync(join(fixDir, "far.h5")));
|
||||
const farData = farBytes.subarray(far.data_at, far.data_at + far.nbytes);
|
||||
const ff = await pkg.openUrl("http://far.invalid/far.h5", {
|
||||
fetch: mockFetch(farBytes, { total: far.length, extra: [[far.far_at, farData]] }),
|
||||
});
|
||||
eq(Array.from((await ff.read("/x")).data), far.values, "data past 2 GiB");
|
||||
eq(ff.stats().size, far.length, "size past 2 GiB");
|
||||
|
||||
// wasm32's reader holds file offsets in 32 bits: a file of 4 GiB or more
|
||||
// is refused at open, with the limit in the message (past 2^53 - 1 bytes
|
||||
// a JavaScript number is not even exact).
|
||||
for (const total of [2 ** 32, 2 ** 40, 2 ** 53 + 2]) {
|
||||
await fails(() => pkg.openUrl("http://huge.invalid/x.h5", { fetch: mockFetch(farBytes, { total }) }),
|
||||
total > 2 ** 53 ? /2\^53 - 1/ : /4 GiB/, `length ${total}`);
|
||||
}
|
||||
}
|
||||
|
||||
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
|
||||
// tells how many were.
|
||||
function flood(total) {
|
||||
let pulled = 0;
|
||||
const body = new ReadableStream({
|
||||
pull(c) {
|
||||
if (pulled >= total) return c.close();
|
||||
const n = Math.min(65536, total - pulled);
|
||||
pulled += n;
|
||||
c.enqueue(new Uint8Array(n));
|
||||
},
|
||||
}, { highWaterMark: 0 });
|
||||
return { body, pulled: () => pulled };
|
||||
}
|
||||
|
||||
// A 206 whose body runs past the range asked for is cut off as it arrives:
|
||||
// the page never reads (or holds) more than it asked for, whatever the
|
||||
// server sends. 64 MiB stands for "gigabytes".
|
||||
async function floodTests() {
|
||||
const FLOOD = 64 << 20;
|
||||
// The probe: asked for the first 1 MiB.
|
||||
let f = flood(FLOOD);
|
||||
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
|
||||
fetch: async () => new Response(f.body, { status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000" } }),
|
||||
}), /the server sent more/, "flooded probe");
|
||||
assert.ok(f.pulled() <= (1 << 20) + 65536, `probe: read ${f.pulled()} bytes`);
|
||||
// A range request after a correct probe.
|
||||
f = flood(FLOOD);
|
||||
const file = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||
blockSize: 512,
|
||||
fetch: async (url, init) => {
|
||||
const range = new Headers(init.headers).get("Range");
|
||||
if (range === "bytes=0-511") return fetch(url, init);
|
||||
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
|
||||
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
|
||||
},
|
||||
}).catch((e) => e);
|
||||
// The open itself may need a second range: flooded either way.
|
||||
const read = file instanceof Error ? Promise.reject(file) : file.read("/grid");
|
||||
await fails(() => read, /the server sent more/, "flooded range");
|
||||
assert.ok(f.pulled() <= 8 * 512 + 65536, `range: read ${f.pulled()} bytes`);
|
||||
// A declared Content-Length past the range is refused before reading.
|
||||
f = flood(FLOOD);
|
||||
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
|
||||
fetch: async () => new Response(f.body, {
|
||||
status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000", "Content-Length": String(FLOOD) },
|
||||
}),
|
||||
}), /the server sent more/, "declared too long");
|
||||
assert.ok(f.pulled() <= 65536, `declared: read ${f.pulled()} bytes`);
|
||||
}
|
||||
|
||||
function hdf5Files(dir, out) {
|
||||
for (const e of readdirSync(dir, { withFileTypes: true })) {
|
||||
const p = join(dir, e.name);
|
||||
let st;
|
||||
try {
|
||||
st = statSync(p);
|
||||
} catch {
|
||||
continue;
|
||||
}
|
||||
if (st.isDirectory()) hdf5Files(p, out);
|
||||
else if (st.size <= 16 << 20) {
|
||||
const b = readFileSync(p);
|
||||
const sig = [0x89, 0x48, 0x44, 0x46, 0x0d, 0x0a, 0x1a, 0x0a];
|
||||
for (let at = 0; at + 8 <= b.length; at = at === 0 ? 512 : at * 2) {
|
||||
if (sig.every((x, i) => b[at + i] === x)) { out.push(p); break; }
|
||||
}
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function show(v) {
|
||||
return JSON.stringify(v, (_, x) => {
|
||||
if (typeof x === "bigint") return `${x}n`;
|
||||
if (ArrayBuffer.isView(x)) return Array.from(x, (y) => (typeof y === "bigint" ? `${y}n` : Number.isNaN(y) ? "NaN" : y));
|
||||
return x;
|
||||
});
|
||||
}
|
||||
|
||||
async function transcript(f) {
|
||||
const out = [];
|
||||
const call = async (what, fn) => {
|
||||
try {
|
||||
out.push(`${what}: ${show(await fn())}`);
|
||||
} catch {
|
||||
out.push(`${what}: Err`);
|
||||
}
|
||||
};
|
||||
const todo = [["/", 0]];
|
||||
while (todo.length && out.length < 3000) {
|
||||
const [path, depth] = todo.pop();
|
||||
let kind = null;
|
||||
await call(`${path} kind`, async () => (kind = await f.kind(path)));
|
||||
await call(`${path} attrs`, () => f.attrs(path));
|
||||
if (kind === "group") {
|
||||
let list = [];
|
||||
await call(`${path} list`, async () => (list = await f.list(path)));
|
||||
if (depth < 12) for (const c of list.reverse()) todo.push([lib.joinPath(path, c.name), depth + 1]);
|
||||
} else if (kind === "dataset") {
|
||||
let info = null;
|
||||
await call(`${path} info`, async () => (info = await f.info(path)));
|
||||
if (info && [...info.shape, ...info.elementShape].reduce((a, b) => a * b, 1) <= 1 << 20) {
|
||||
await call(`${path} read`, () => f.read(path));
|
||||
}
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
async function corpusTests(corpus) {
|
||||
const files = hdf5Files(corpus, []).sort();
|
||||
let opened = 0;
|
||||
let bytes = 0;
|
||||
let size = 0;
|
||||
for (const p of files) {
|
||||
const rel = relative(corpus, p).split("/").map(encodeURIComponent).join("/");
|
||||
let local;
|
||||
try {
|
||||
local = pkg.open(new Uint8Array(readFileSync(p)));
|
||||
} catch {
|
||||
await fails(() => pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 }), /./, `${rel}: opens in neither`);
|
||||
continue;
|
||||
}
|
||||
const remote = await pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 });
|
||||
const want = await transcript(local);
|
||||
const got = await transcript(remote);
|
||||
// Errors are compared as errors: a malformed file can fail at another
|
||||
// check, with another message, when read by ranges.
|
||||
eq(got, want, `${rel}: transcript`);
|
||||
opened++;
|
||||
bytes += remote.stats().bytesFetched;
|
||||
size += remote.stats().size;
|
||||
local.free();
|
||||
remote.free();
|
||||
}
|
||||
console.log(`corpus: ${opened} of ${files.length} files agree over HTTP (${bytes} of ${size} bytes fetched)`);
|
||||
}
|
||||
|
||||
@@ -75,3 +75,24 @@ export function toRows(data, rows, cols, per = 1) {
|
||||
export function perElement(elementShape) {
|
||||
return elementShape.reduce((a, b) => a * b, 1);
|
||||
}
|
||||
|
||||
/** A byte count for people: `1.5 KiB`, `200 MiB`. */
|
||||
export function formatBytes(n) {
|
||||
const units = ["B", "KiB", "MiB", "GiB", "TiB"];
|
||||
let i = 0;
|
||||
let v = n;
|
||||
while (v >= 1024 && i < units.length - 1) {
|
||||
v /= 1024;
|
||||
i++;
|
||||
}
|
||||
const s = i === 0 ? String(v) : v < 10 ? String(Number(v.toFixed(1))) : String(Math.round(v));
|
||||
return `${s} ${units[i]}`;
|
||||
}
|
||||
|
||||
/** What reading a remote file has cost, from `RemoteFile.stats()`. */
|
||||
export function formatStats({ lazy, requests, bytesFetched, size }) {
|
||||
if (!lazy) return `downloaded whole (${formatBytes(size)}): the server does not support range requests`;
|
||||
const pct = size > 0 ? (100 * bytesFetched) / size : 0;
|
||||
return `${requests} request${requests === 1 ? "" : "s"}, ${formatBytes(bytesFetched)} of ` +
|
||||
`${formatBytes(size)} fetched (${pct.toFixed(1)}%)`;
|
||||
}
|
||||
|
||||
+9
-3
@@ -131,14 +131,17 @@ run_step "cargo clippy (h5rs remote)" cargo clippy \
|
||||
# js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C.
|
||||
# clawhdf5-remote is checked by default (plain HTTP) and with its
|
||||
# object-store feature, and h5rs with URL support (remote); the https
|
||||
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in.
|
||||
# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. The
|
||||
# Python bindings (clawhdf5-py, remote reads over plain HTTP) are checked too:
|
||||
# their https/s3/gcs/azure features are opt-in for the same reason.
|
||||
no_c_in_default_build() {
|
||||
local entry crate features found=0
|
||||
for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \
|
||||
clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \
|
||||
clawhdf5-tools \
|
||||
clawhdf5-wasm \
|
||||
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote; do
|
||||
clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote \
|
||||
clawhdf5-py; do
|
||||
crate=${entry%%:*}
|
||||
features=()
|
||||
[ "$entry" != "$crate" ] && features=(--features "${entry#*:}")
|
||||
@@ -269,7 +272,10 @@ python_package() {
|
||||
-i "$PYTHON" \
|
||||
--out "$out/wheel" || return 1
|
||||
"$PYTHON" -m pip install --quiet --no-deps --target "$out/site" "$out"/wheel/*.whl || return 1
|
||||
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
|
||||
# The editing tests run `h5rs check` on every file they edit.
|
||||
cargo build -q -p clawhdf5-tools || return 1
|
||||
CLAWHDF5_H5RS="${CARGO_TARGET_DIR:-$root/target}/debug/h5rs" \
|
||||
PYTHONPATH="$out/site" "$PYTHON" -m pytest -q -p no:cacheprovider \
|
||||
"$root/crates/clawhdf5-py/tests"
|
||||
}
|
||||
if "$PYTHON" -m maturin --version >/dev/null 2>&1 && "$PYTHON" -c "import pytest" >/dev/null 2>&1; then
|
||||
|
||||
Reference in New Issue
Block a user