Merge branch 'feat/p3-m4-wasm-lazy' into feat/p3-wasm-swmr-python
# Conflicts: # CHANGELOG.md # docs/design/range-reads.md # docs/known-issues.md
This commit is contained in:
@@ -214,6 +214,87 @@ Design: `docs/design/swmr.md`.
|
|||||||
that hangs up, 16 threads on one remote file, and a thread that keeps
|
that hangs up, 16 threads on one remote file, and a thread that keeps
|
||||||
running while a read waits on 0.2 s requests.
|
running while a read waits on 0.2 s requests.
|
||||||
|
|
||||||
|
### Range reads, milestone M4: remote files in the browser (2026-09-27)
|
||||||
|
- **`openUrl(url, opts)`** in `clawhdf5-wasm` opens an HDF5/NetCDF-4 file
|
||||||
|
on a web server without downloading it: the returned `RemoteFile` has
|
||||||
|
`H5File`'s methods (`kind`, `list`, `info`, `attrs`, `attrErrors`,
|
||||||
|
`read`, `readHyperslab`), each returning a promise, and `stats()`
|
||||||
|
(requests, bytes fetched, file size). Only the byte ranges a call needs
|
||||||
|
are fetched, with `fetch` and `Range` headers, in 1 MiB blocks
|
||||||
|
(`blockSize`) kept in a 64 MiB cache (`cacheSize`). `open(bytes)` is
|
||||||
|
unchanged.
|
||||||
|
- **How, on the browser's main thread:** the design's restartable
|
||||||
|
"NeedBytes" mode (`clawhdf5_wasm::lazy::LazyStorage`). A call runs as a
|
||||||
|
pass over the blocks fetched so far; a read that misses records the
|
||||||
|
missing blocks and fails, the pass's result is dropped (even if a parser
|
||||||
|
caught the error and carried on), the blocks are fetched and the pass is
|
||||||
|
re-run. No Web Worker and no synchronous XHR (what h5wasm's lazy files
|
||||||
|
need). No block is evicted while a call is in flight, so every call ends;
|
||||||
|
the cache is trimmed between calls, raw-data blocks before metadata.
|
||||||
|
`clawhdf5-remote`'s `BlockCache` is not reused: it fetches by blocking,
|
||||||
|
and it evicts, and does not keep large reads, during a read, where a
|
||||||
|
restartable pass needs every block it read to stay until it finishes.
|
||||||
|
- **Every answer is checked** (`js/remote.js`): a `206` with exactly the
|
||||||
|
bytes asked for (`Content-Range`, when the page can see it, and the body
|
||||||
|
length), and the same `ETag`/`Last-Modified` and length as at open, or
|
||||||
|
the call fails: never data from another file or offset. A server that
|
||||||
|
answers `200` to a `Range` request is downloaded whole (up to
|
||||||
|
`maxDownload`, 512 MiB) unless `fallback: "error"`. Other options:
|
||||||
|
`headers`, `credentials`, `parallel` (6 requests at a time), `fetch`.
|
||||||
|
- **The viewer** (`examples/wasm-viewer`) has a URL box, opens
|
||||||
|
`?file=<url>` lazily, and shows the requests and bytes fetched.
|
||||||
|
- Counted on tank, 2026-09-27 (`WASM_BIG_MB=200 bash
|
||||||
|
examples/wasm-viewer/test/run.sh`, requests as `test/serve.py` counted
|
||||||
|
them): in a 200 MB h5py file, listing the root, reading two small
|
||||||
|
datasets, a group's attributes, the large dataset's shape and a
|
||||||
|
10-value window of it (25 million values) took 5 requests and 6 MiB.
|
||||||
|
With
|
||||||
|
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, the 622 corpus files
|
||||||
|
up to 16 MiB that open read over HTTP as from bytes: the same listings,
|
||||||
|
attributes and values (datasets up to 2^20 values), and an error
|
||||||
|
wherever bytes give one. Natively, `cargo test -p clawhdf5-wasm --test
|
||||||
|
lazy` with the same variable compares 656 files (up to 64 MiB, error
|
||||||
|
messages included, against the facade's range-storage path).
|
||||||
|
- `Reader::open_storage` (the wasm crate's core over any `Storage`), and
|
||||||
|
variable-length strings resolve through the file's storage rather than
|
||||||
|
`File::as_bytes`.
|
||||||
|
- The package is larger: the facade's `Storage` read path is now
|
||||||
|
reachable from JavaScript (it was compiled out before), and the promise
|
||||||
|
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
|
||||||
|
build machine was shared); the viewer README's size table predates M4.
|
||||||
|
- **Hardened after review (2026-09-27):**
|
||||||
|
- Sizes a server or a dataset names are errors, never an abort of the
|
||||||
|
wasm module (which took every open file on the page with it): a read
|
||||||
|
past 2 GiB aborted in the lazy cache, reachable by a hostile server
|
||||||
|
claiming a large file and a 2 GiB heap collection, and `read()` of a
|
||||||
|
256 MiB `u8` dataset aborted widening it to 64 bits. New option
|
||||||
|
`maxFetch` (512 MiB, at most 1 GiB): what one call may fetch, and the
|
||||||
|
longest single read, refused before fetching. `read()` refuses a
|
||||||
|
dataset that would take more than 1 GiB to decode, naming
|
||||||
|
`readHyperslab`. A file of 4 GiB or more is refused at open (wasm32
|
||||||
|
reads offsets as 32-bit); `maxDownload` is at most 1 GiB.
|
||||||
|
- Response bodies are read as they arrive and cut off at the length
|
||||||
|
asked for (`maxDownload` for a `200`): a `206` with a gigabyte body
|
||||||
|
was buffered whole before its length was checked.
|
||||||
|
- Listing a group reads every child's header, and every node of each
|
||||||
|
level of the group's index, in one pass: 3000 datasets (h5py, 198 MB)
|
||||||
|
listed in 6 passes and 73 requests at 1 MiB blocks instead of 185
|
||||||
|
passes and 184 serial requests (`libver="latest"`: 9 passes instead of
|
||||||
|
189). In `clawhdf5-format`, the B-tree v1/v2 collectors, the symbol
|
||||||
|
table node loop and the dense-link loop read (without using) the
|
||||||
|
siblings after the first that fails, then return that error: same
|
||||||
|
results and errors, more reads only on failure (free in memory). The
|
||||||
|
lazy cache no longer re-fetches a cached block to merge two requests.
|
||||||
|
- `headers` may be a `Headers` instance or `[name, value]` pairs (a
|
||||||
|
`Headers` was silently dropped); when one range request fails the
|
||||||
|
others in flight are aborted; `parallel` must be an integer from 1 to
|
||||||
|
1024.
|
||||||
|
- Tests: `test/serve.py` serves ranges without exposed `Content-Range`/
|
||||||
|
`ETag` (`/noexpose/`, and `/unexposed/` for a real cross-origin page
|
||||||
|
in Chromium), so the HEAD-length path runs end to end; hostile and
|
||||||
|
oversized files (`make_fixture.py`'s `write_limits`), flooding
|
||||||
|
bodies, aborted siblings.
|
||||||
|
|
||||||
### Range reads, milestone M3: remote files (2026-09-26)
|
### Range reads, milestone M3: remote files (2026-09-26)
|
||||||
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
||||||
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
|
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
|
||||||
|
|||||||
@@ -181,14 +181,19 @@ Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal F
|
|||||||
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
|
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
|
||||||
file over HTTP with `File::open`.
|
file over HTTP with `File::open`.
|
||||||
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
|
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
|
||||||
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only, file held in memory;
|
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
|
||||||
no Zstd/SZIP since they link C) and the `examples/wasm-viewer/` page.
|
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
|
||||||
`examples/wasm-viewer/test/run.sh` builds the package (needs the
|
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
|
||||||
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node
|
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
|
||||||
and headless Chromium (a Playwright download in `~/.cache/ms-playwright`
|
call is re-run after each wave of misses; no block evicted while a call
|
||||||
on tank); the CI container has neither, so CI runs the native
|
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
|
||||||
`clawhdf5-wasm` `h5py_interop` test on the same fixture. Size numbers are
|
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
|
||||||
in the example's README.
|
version) and tests it under Node and headless Chromium (a Playwright
|
||||||
|
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
|
||||||
|
(range server with request counts, 200 MB budget file); the CI container
|
||||||
|
has neither, so CI runs the native `h5py_interop` and `lazy` tests
|
||||||
|
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
|
||||||
|
numbers in the example's README predate `openUrl`.
|
||||||
- Python and Node.js bindings for cross-language use
|
- Python and Node.js bindings for cross-language use
|
||||||
- NetCDF-4 compatibility for scientific data interop
|
- NetCDF-4 compatibility for scientific data interop
|
||||||
|
|
||||||
|
|||||||
@@ -827,7 +827,7 @@ clawhdf5 workspace (19 crates, ~86K lines of Rust in src/, ~104K with tests
|
|||||||
├── Bindings
|
├── Bindings
|
||||||
│ ├── clawhdf5-py — Python (PyO3)
|
│ ├── clawhdf5-py — Python (PyO3)
|
||||||
│ ├── clawhdf5-napi — Node.js (napi-rs)
|
│ ├── clawhdf5-napi — Node.js (napi-rs)
|
||||||
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only)
|
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only; remote files by HTTP range requests)
|
||||||
│
|
│
|
||||||
└── Tooling
|
└── Tooling
|
||||||
├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check
|
├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check
|
||||||
|
|||||||
@@ -199,19 +199,32 @@ fn collect_symbol_table_nodes_inner<S: Storage + ?Sized>(
|
|||||||
// Leaf: children are SNOD addresses
|
// Leaf: children are SNOD addresses
|
||||||
Ok(node.children)
|
Ok(node.children)
|
||||||
} else {
|
} else {
|
||||||
// Internal: recurse into children
|
// Internal: recurse into children. After the first child that
|
||||||
|
// fails, the others are only read (as `storage::touch` does), not
|
||||||
|
// descended into; that error is returned.
|
||||||
let mut result = Vec::new();
|
let mut result = Vec::new();
|
||||||
|
let mut failed = None;
|
||||||
for &child_addr in &node.children {
|
for &child_addr in &node.children {
|
||||||
let child_snods = collect_symbol_table_nodes_inner(
|
if failed.is_some() {
|
||||||
|
// Parsing reads the node's header, then its body.
|
||||||
|
let _ = BTreeV1Node::parse_in(file, child_addr, offset_size, length_size);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
match collect_symbol_table_nodes_inner(
|
||||||
file,
|
file,
|
||||||
child_addr,
|
child_addr,
|
||||||
offset_size,
|
offset_size,
|
||||||
length_size,
|
length_size,
|
||||||
depth + 1,
|
depth + 1,
|
||||||
)?;
|
) {
|
||||||
result.extend(child_snods);
|
Ok(child_snods) => result.extend(child_snods),
|
||||||
|
Err(e) => failed = Some(e),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
match failed {
|
||||||
|
Some(e) => Err(e),
|
||||||
|
None => Ok(result),
|
||||||
}
|
}
|
||||||
Ok(result)
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -504,45 +504,63 @@ fn collect_internal_records<S: Storage + ?Sized>(
|
|||||||
|
|
||||||
// Interleave: child[0], record[0], child[1], record[1], ..., child[nr]
|
// Interleave: child[0], record[0], child[1], record[1], ..., child[nr]
|
||||||
// We collect child[0] records, then record[0], then child[1], etc.
|
// We collect child[0] records, then record[0], then child[1], etc.
|
||||||
|
// After the first child that fails, the others are only touched (see
|
||||||
|
// `storage::touch`); that error is returned.
|
||||||
|
let mut failed = None;
|
||||||
for (i, &(child_addr, child_nrec)) in node.children.iter().enumerate() {
|
for (i, &(child_addr, child_nrec)) in node.children.iter().enumerate() {
|
||||||
if child_depth == 0 {
|
if failed.is_some() {
|
||||||
// Before parsing, so a refused tree is not also a large allocation.
|
let len = usize::try_from(node_size)
|
||||||
spend(budget, usize::from(child_nrec))?;
|
.unwrap_or(usize::MAX)
|
||||||
let leaf_recs = parse_leaf_records(
|
.min(1 << 16);
|
||||||
file,
|
crate::storage::touch(file, child_addr, len);
|
||||||
to_usize(child_addr)?,
|
continue;
|
||||||
child_nrec,
|
|
||||||
record_size,
|
|
||||||
node_size,
|
|
||||||
)?;
|
|
||||||
out.extend(leaf_recs);
|
|
||||||
} else {
|
|
||||||
collect_internal_records(
|
|
||||||
file,
|
|
||||||
to_usize(child_addr)?,
|
|
||||||
child_nrec,
|
|
||||||
child_depth,
|
|
||||||
record_size,
|
|
||||||
node_size,
|
|
||||||
offset_size,
|
|
||||||
length_size,
|
|
||||||
max_leaf_nrec,
|
|
||||||
budget,
|
|
||||||
out,
|
|
||||||
)?;
|
|
||||||
}
|
}
|
||||||
|
if let Err(e) = (|| -> Result<(), FormatError> {
|
||||||
|
if child_depth == 0 {
|
||||||
|
// Before parsing, so a refused tree is not also a large allocation.
|
||||||
|
spend(budget, usize::from(child_nrec))?;
|
||||||
|
let leaf_recs = parse_leaf_records(
|
||||||
|
file,
|
||||||
|
to_usize(child_addr)?,
|
||||||
|
child_nrec,
|
||||||
|
record_size,
|
||||||
|
node_size,
|
||||||
|
)?;
|
||||||
|
out.extend(leaf_recs);
|
||||||
|
} else {
|
||||||
|
collect_internal_records(
|
||||||
|
file,
|
||||||
|
to_usize(child_addr)?,
|
||||||
|
child_nrec,
|
||||||
|
child_depth,
|
||||||
|
record_size,
|
||||||
|
node_size,
|
||||||
|
offset_size,
|
||||||
|
length_size,
|
||||||
|
max_leaf_nrec,
|
||||||
|
budget,
|
||||||
|
out,
|
||||||
|
)?;
|
||||||
|
}
|
||||||
|
|
||||||
// Add record[i] (except after the last child)
|
// Add record[i] (except after the last child)
|
||||||
if i < nr {
|
if i < nr {
|
||||||
let data = node.record(i, rs)?;
|
let data = node.record(i, rs)?;
|
||||||
spend(budget, 1)?;
|
spend(budget, 1)?;
|
||||||
out.push(BTreeV2Record {
|
out.push(BTreeV2Record {
|
||||||
data: data.to_vec(),
|
data: data.to_vec(),
|
||||||
});
|
});
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
})() {
|
||||||
|
failed = Some(e);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
Ok(())
|
match failed {
|
||||||
|
Some(e) => Err(e),
|
||||||
|
None => Ok(()),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The records of a B-tree v2 that fall in one key range, found by
|
/// The records of a B-tree v2 that fall in one key range, found by
|
||||||
|
|||||||
@@ -79,27 +79,51 @@ pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
|
|||||||
length_size,
|
length_size,
|
||||||
)?;
|
)?;
|
||||||
|
|
||||||
|
// The names are read one by one from the heap's data segment; read
|
||||||
|
// (up to 1 MiB of) it first, so a storage that records what it lacks
|
||||||
|
// asks for it at once (see `storage::touch`).
|
||||||
|
if !snod_addrs.is_empty() {
|
||||||
|
let len = usize::try_from(heap.data_segment_size).map_or(1 << 20, |n| n.min(1 << 20));
|
||||||
|
crate::storage::touch(file_data, heap.data_segment_address, len);
|
||||||
|
}
|
||||||
|
|
||||||
let mut entries = Vec::new();
|
let mut entries = Vec::new();
|
||||||
let mut heap_checked = false;
|
let mut heap_checked = false;
|
||||||
|
// After the first node that fails, the others are only read (as
|
||||||
|
// `storage::touch` does); that error is returned.
|
||||||
|
let mut failed = None;
|
||||||
for snod_addr in snod_addrs {
|
for snod_addr in snod_addrs {
|
||||||
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
|
if failed.is_some() {
|
||||||
for entry in &snod.entries {
|
let _ = SymbolTableNode::parse_in(file_data, snod_addr, offset_size);
|
||||||
// Like libhdf5, look at the heap's free list only once a name is
|
continue;
|
||||||
// needed: an empty group with a damaged heap still lists.
|
}
|
||||||
if !heap_checked {
|
let mut node = || -> Result<(), FormatError> {
|
||||||
heap.validate_free_list_in(file_data, length_size)?;
|
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
|
||||||
heap_checked = true;
|
for entry in &snod.entries {
|
||||||
|
// Like libhdf5, look at the heap's free list only once a name
|
||||||
|
// is needed: an empty group with a damaged heap still lists.
|
||||||
|
if !heap_checked {
|
||||||
|
heap.validate_free_list_in(file_data, length_size)?;
|
||||||
|
heap_checked = true;
|
||||||
|
}
|
||||||
|
let name = heap.read_string_in(file_data, entry.link_name_offset)?;
|
||||||
|
entries.push(GroupEntry {
|
||||||
|
name,
|
||||||
|
object_header_address: entry.object_header_address,
|
||||||
|
cache_type: entry.cache_type,
|
||||||
|
});
|
||||||
}
|
}
|
||||||
let name = heap.read_string_in(file_data, entry.link_name_offset)?;
|
Ok(())
|
||||||
entries.push(GroupEntry {
|
};
|
||||||
name,
|
if let Err(e) = node() {
|
||||||
object_header_address: entry.object_header_address,
|
failed = Some(e);
|
||||||
cache_type: entry.cache_type,
|
|
||||||
});
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
Ok(entries)
|
match failed {
|
||||||
|
Some(e) => Err(e),
|
||||||
|
None => Ok(entries),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Symbol table cache type for a soft link: the scratch pad's first four bytes
|
/// Symbol table cache type for a soft link: the scratch pad's first four bytes
|
||||||
|
|||||||
@@ -126,6 +126,9 @@ fn for_each_dense_link<S: Storage + ?Sized>(
|
|||||||
)?;
|
)?;
|
||||||
let records = collect_btree_v2_records_in(file_data, &btree_hdr, offset_size, length_size)?;
|
let records = collect_btree_v2_records_in(file_data, &btree_hdr, offset_size, length_size)?;
|
||||||
|
|
||||||
|
// After the first link that fails, the others are only read, not
|
||||||
|
// visited (a touch, see `storage::touch`); that error is returned.
|
||||||
|
let mut failed = None;
|
||||||
for record in &records {
|
for record in &records {
|
||||||
// For type 5 (name index): hash(4) + heap_id(heap_id_length)
|
// For type 5 (name index): hash(4) + heap_id(heap_id_length)
|
||||||
// For type 6 (creation order): creation_order(8) + heap_id(heap_id_length)
|
// For type 6 (creation order): creation_order(8) + heap_id(heap_id_length)
|
||||||
@@ -141,12 +144,20 @@ fn for_each_dense_link<S: Storage + ?Sized>(
|
|||||||
let id_bytes = &record.data[id_offset..id_offset + fh.heap_id_length as usize];
|
let id_bytes = &record.data[id_offset..id_offset + fh.heap_id_length as usize];
|
||||||
|
|
||||||
// Read managed object from fractal heap
|
// Read managed object from fractal heap
|
||||||
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size)?;
|
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size);
|
||||||
if let Some(link) = parse_link(&link_data, offset_size)? {
|
if failed.is_some() {
|
||||||
visit(link);
|
continue;
|
||||||
|
}
|
||||||
|
match link_data.and_then(|d| parse_link(&d, offset_size)) {
|
||||||
|
Ok(Some(link)) => visit(link),
|
||||||
|
Ok(None) => {}
|
||||||
|
Err(e) => failed = Some(e),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
Ok(())
|
match failed {
|
||||||
|
Some(e) => Err(e),
|
||||||
|
None => Ok(()),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Resolve entries from dense storage (fractal heap + B-tree v2).
|
/// Resolve entries from dense storage (fractal heap + B-tree v2).
|
||||||
|
|||||||
@@ -202,6 +202,19 @@ pub(crate) fn len_usize<S: Storage + ?Sized>(file: &S) -> usize {
|
|||||||
usize::try_from(file.len()).unwrap_or(usize::MAX)
|
usize::try_from(file.len()).unwrap_or(usize::MAX)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Read `len` bytes at `offset` and drop them, ignoring any error.
|
||||||
|
///
|
||||||
|
/// For a traversal that has failed on one sibling (a B-tree child, a
|
||||||
|
/// symbol table node, a heap object) and would stop there: it first
|
||||||
|
/// touches the siblings it did not get to, so a storage that records what
|
||||||
|
/// it lacks — the browser's restartable reader, which fetches over the
|
||||||
|
/// network between attempts — learns about all of them in one attempt
|
||||||
|
/// instead of one per attempt. Results and errors are unchanged (the first
|
||||||
|
/// error is still the one returned); an in-memory read is free.
|
||||||
|
pub fn touch<S: Storage + ?Sized>(file: &S, offset: u64, len: usize) {
|
||||||
|
let _ = file.read_at(offset, len);
|
||||||
|
}
|
||||||
|
|
||||||
/// Bytes `[offset, offset + len)`, all of them.
|
/// Bytes `[offset, offset + len)`, all of them.
|
||||||
///
|
///
|
||||||
/// A range that runs past the end of the storage is
|
/// A range that runs past the end of the storage is
|
||||||
|
|||||||
@@ -24,6 +24,8 @@ clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
|
|||||||
# Must match the wasm-bindgen CLI exactly; build.sh checks.
|
# Must match the wasm-bindgen CLI exactly; build.sh checks.
|
||||||
wasm-bindgen = "0.2.129"
|
wasm-bindgen = "0.2.129"
|
||||||
js-sys = "0.3.106"
|
js-sys = "0.3.106"
|
||||||
|
# Promises for openUrl and RemoteFile (pure Rust over js-sys).
|
||||||
|
wasm-bindgen-futures = "0.4.79"
|
||||||
|
|
||||||
[dev-dependencies]
|
[dev-dependencies]
|
||||||
serde_json = "1"
|
serde_json = "1"
|
||||||
|
|||||||
@@ -0,0 +1,235 @@
|
|||||||
|
// HTTP for clawhdf5-wasm's openUrl (see src/lib.rs and src/lazy.rs).
|
||||||
|
//
|
||||||
|
// The Rust side decides which byte ranges a read needs; this file fetches
|
||||||
|
// them with `fetch` and `Range` headers and checks every answer, so a server
|
||||||
|
// that ignores the range, answers with other bytes, or serves a file that
|
||||||
|
// changed since it was opened is an error, never data. wasm-bindgen copies
|
||||||
|
// it into the package (pkg/snippets/...).
|
||||||
|
|
||||||
|
const DEFAULT_MAX_DOWNLOAD = 512 * 1024 * 1024;
|
||||||
|
const DEFAULT_PARALLEL = 6;
|
||||||
|
|
||||||
|
function fetcher(opts) {
|
||||||
|
const f = opts?.fetch ?? globalThis.fetch;
|
||||||
|
if (typeof f !== "function") {
|
||||||
|
throw new Error("openUrl: no fetch() in this environment (pass opts.fetch)");
|
||||||
|
}
|
||||||
|
return f;
|
||||||
|
}
|
||||||
|
|
||||||
|
// The caller's headers (`opts.headers`: a Headers, [name, value] pairs or a
|
||||||
|
// plain object, as fetch takes them; names come out in lower case) with
|
||||||
|
// `extra` over them. A Range of the caller's is dropped: this file asks for
|
||||||
|
// the ranges.
|
||||||
|
function init(opts, extra, method = "GET", signal = undefined) {
|
||||||
|
const headers = {};
|
||||||
|
if (opts?.headers != null) {
|
||||||
|
for (const [name, value] of new Headers(opts.headers)) {
|
||||||
|
if (name !== "range") headers[name] = value;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Object.assign(headers, extra);
|
||||||
|
return { method, headers, credentials: opts?.credentials, signal };
|
||||||
|
}
|
||||||
|
|
||||||
|
// `opts.parallel`: range requests in flight at once.
|
||||||
|
function parallelism(opts) {
|
||||||
|
const p = opts?.parallel ?? DEFAULT_PARALLEL;
|
||||||
|
if (!Number.isSafeInteger(p) || p < 1) {
|
||||||
|
throw new Error(`openUrl: parallel must be a positive integer, got ${String(p)}`);
|
||||||
|
}
|
||||||
|
return p;
|
||||||
|
}
|
||||||
|
|
||||||
|
// "bytes a-b/total" -> { start, end (exclusive), total | null }; null when
|
||||||
|
// the page cannot see the header (cross-origin, not exposed).
|
||||||
|
function contentRange(resp, url) {
|
||||||
|
const v = resp.headers.get("Content-Range");
|
||||||
|
if (v === null) return null;
|
||||||
|
const m = /^bytes (\d+)-(\d+)\/(\d+|\*)$/.exec(v.trim());
|
||||||
|
if (!m) throw new Error(`${url}: the server sent an unusable Content-Range: ${v}`);
|
||||||
|
return { start: Number(m[1]), end: Number(m[2]) + 1, total: m[3] === "*" ? null : Number(m[3]) };
|
||||||
|
}
|
||||||
|
|
||||||
|
// What pins the file: its ETag, else its Last-Modified (null if neither is
|
||||||
|
// visible to this page).
|
||||||
|
function validatorOf(resp) {
|
||||||
|
return resp.headers.get("ETag") ?? resp.headers.get("Last-Modified");
|
||||||
|
}
|
||||||
|
|
||||||
|
async function discard(resp) {
|
||||||
|
try {
|
||||||
|
await resp.body?.cancel();
|
||||||
|
} catch {
|
||||||
|
// Nothing to release.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The body, refusing more than `limit` bytes as they arrive: it is piped
|
||||||
|
// through a TransformStream that errors the moment the count passes the
|
||||||
|
// limit, which cancels the body and so aborts the request. Whatever the
|
||||||
|
// server declares or sends, the page never holds more than `limit` bytes of
|
||||||
|
// it. (`tooBig(n)` makes the error; n is the count so far.) A reader loop
|
||||||
|
// would do the same, but it stalls on small bodies in headless Chromium
|
||||||
|
// under --virtual-time-budget, which the page test uses; a pipe does not.
|
||||||
|
async function readCapped(resp, limit, tooBig) {
|
||||||
|
const declared = resp.headers.get("Content-Length");
|
||||||
|
if (declared !== null && Number(declared) > limit) {
|
||||||
|
await discard(resp);
|
||||||
|
throw tooBig(declared);
|
||||||
|
}
|
||||||
|
if (!resp.body) {
|
||||||
|
// No stream to read from (some fetch implementations): all at once.
|
||||||
|
const all = new Uint8Array(await resp.arrayBuffer());
|
||||||
|
if (all.length > limit) throw tooBig(all.length);
|
||||||
|
return all;
|
||||||
|
}
|
||||||
|
let n = 0;
|
||||||
|
let over = null;
|
||||||
|
const capped = resp.body.pipeThrough(new TransformStream({
|
||||||
|
transform(chunk, ctl) {
|
||||||
|
n += chunk.length;
|
||||||
|
if (n > limit) {
|
||||||
|
over = tooBig(`over ${limit}`);
|
||||||
|
ctl.error(over);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
ctl.enqueue(chunk);
|
||||||
|
},
|
||||||
|
}));
|
||||||
|
try {
|
||||||
|
return new Uint8Array(await new Response(capped).arrayBuffer());
|
||||||
|
} catch (e) {
|
||||||
|
throw over ?? e;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The whole body of a 200 answer (a server without range support), at
|
||||||
|
// most `limit` (maxDownload) bytes.
|
||||||
|
function readAll(resp, limit, url) {
|
||||||
|
return readCapped(resp, limit, (n) =>
|
||||||
|
new Error(`${url} is ${n} bytes, more than maxDownload (${limit}); ` +
|
||||||
|
"the server does not support range requests, so the whole file would have to be downloaded"));
|
||||||
|
}
|
||||||
|
|
||||||
|
// The body of a 206 answer, which may not be longer than the `limit` bytes
|
||||||
|
// asked for at `start` (the caller checks the exact length).
|
||||||
|
function readLimited(resp, limit, url, start) {
|
||||||
|
return readCapped(resp, limit, () =>
|
||||||
|
new Error(`${url}: asked for ${limit} bytes at offset ${start}, the server sent more`));
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Ask for the file's first `firstLen` bytes. A server that honours the
|
||||||
|
* range (206) gives `{ length, first, validator, requests }`; one that
|
||||||
|
* answers 200 sends the whole file, which is kept (`{ whole, requests }`)
|
||||||
|
* when `opts.fallback` is "download" (the default) and the file is at most
|
||||||
|
* `opts.maxDownload` bytes, and is an error otherwise.
|
||||||
|
*/
|
||||||
|
export async function probe(url, firstLen, opts) {
|
||||||
|
const f = fetcher(opts);
|
||||||
|
parallelism(opts);
|
||||||
|
const resp = await f(url, init(opts, { Range: `bytes=0-${firstLen - 1}` }));
|
||||||
|
if (resp.status === 206) {
|
||||||
|
const cr = contentRange(resp, url);
|
||||||
|
if (cr && cr.start !== 0) {
|
||||||
|
await discard(resp);
|
||||||
|
throw new Error(`${url}: asked for bytes from 0, the server sent bytes from ${cr.start}`);
|
||||||
|
}
|
||||||
|
const first = await readLimited(resp, firstLen, url, 0);
|
||||||
|
let length = cr?.total ?? null;
|
||||||
|
let requests = 1;
|
||||||
|
if (length === null) {
|
||||||
|
// Content-Range is not readable here: a cross-origin server that does
|
||||||
|
// not list it in Access-Control-Expose-Headers. Content-Length of a
|
||||||
|
// HEAD request is always readable.
|
||||||
|
const head = await f(url, init(opts, {}, "HEAD"));
|
||||||
|
requests++;
|
||||||
|
const cl = head.headers.get("Content-Length");
|
||||||
|
if (!head.ok || cl === null) {
|
||||||
|
throw new Error(`${url}: cannot learn the file's size (a cross-origin server must send ` +
|
||||||
|
"Access-Control-Expose-Headers: Content-Range, or answer HEAD with Content-Length)");
|
||||||
|
}
|
||||||
|
length = Number(cl);
|
||||||
|
}
|
||||||
|
if (!Number.isSafeInteger(length) || length < 0) {
|
||||||
|
throw new Error(`${url}: the server gave a file size of ${length} bytes; openUrl reads files ` +
|
||||||
|
"of up to 2^53 - 1 bytes (the largest offset a JavaScript number holds exactly)");
|
||||||
|
}
|
||||||
|
if (first.length !== Math.min(firstLen, length)) {
|
||||||
|
throw new Error(`${url}: asked for the first ${firstLen} bytes of ${length}, got ${first.length}`);
|
||||||
|
}
|
||||||
|
return { length, first, validator: validatorOf(resp), requests };
|
||||||
|
}
|
||||||
|
if (resp.status === 200) {
|
||||||
|
if ((opts?.fallback ?? "download") !== "download") {
|
||||||
|
await discard(resp);
|
||||||
|
throw new Error(`${url}: the server does not support HTTP range requests (it answered 200 ` +
|
||||||
|
"to a Range request); open it with { fallback: \"download\" } to download the whole file");
|
||||||
|
}
|
||||||
|
const whole = await readAll(resp, opts?.maxDownload ?? DEFAULT_MAX_DOWNLOAD, url);
|
||||||
|
return { whole, requests: 1 };
|
||||||
|
}
|
||||||
|
await discard(resp);
|
||||||
|
throw new Error(`${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Fetch `ranges` ([start0, end0, start1, end1, ...], ends exclusive) of a
|
||||||
|
* file opened by `probe`, at most `opts.parallel` (default 6) at a time.
|
||||||
|
* Every answer must be a 206 with exactly the bytes asked for, from the same
|
||||||
|
* file (validator and length). When one request fails, the others in
|
||||||
|
* flight are aborted and no more are made; that failure is the error.
|
||||||
|
*/
|
||||||
|
export async function fetchRanges(url, ranges, opts, validator, length) {
|
||||||
|
const f = fetcher(opts);
|
||||||
|
const parallel = parallelism(opts);
|
||||||
|
const n = ranges.length / 2;
|
||||||
|
const out = new Array(n);
|
||||||
|
const abort = new AbortController();
|
||||||
|
let next = 0;
|
||||||
|
async function one(i) {
|
||||||
|
const start = ranges[2 * i];
|
||||||
|
const end = ranges[2 * i + 1];
|
||||||
|
const resp = await f(url, init(opts, { Range: `bytes=${start}-${end - 1}` }, "GET", abort.signal));
|
||||||
|
if (resp.status !== 206) {
|
||||||
|
await discard(resp);
|
||||||
|
throw new Error(resp.status === 200
|
||||||
|
? `${url}: the server stopped honouring range requests`
|
||||||
|
: `${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
|
||||||
|
}
|
||||||
|
const cr = contentRange(resp, url);
|
||||||
|
const v = validatorOf(resp);
|
||||||
|
if ((validator != null && v !== null && v !== validator) ||
|
||||||
|
(cr?.total != null && cr.total !== length)) {
|
||||||
|
await discard(resp);
|
||||||
|
throw new Error(`${url} changed on the server since it was opened`);
|
||||||
|
}
|
||||||
|
if (cr && (cr.start !== start || cr.end !== end)) {
|
||||||
|
await discard(resp);
|
||||||
|
throw new Error(`${url}: asked for bytes ${start}-${end - 1}, the server sent ${cr.start}-${cr.end - 1}`);
|
||||||
|
}
|
||||||
|
const body = await readLimited(resp, end - start, url, start);
|
||||||
|
if (body.length !== end - start) {
|
||||||
|
throw new Error(`${url}: asked for ${end - start} bytes at offset ${start}, got ${body.length}`);
|
||||||
|
}
|
||||||
|
out[i] = body;
|
||||||
|
}
|
||||||
|
async function worker() {
|
||||||
|
while (next < n && !abort.signal.aborted) {
|
||||||
|
try {
|
||||||
|
await one(next++);
|
||||||
|
} catch (e) {
|
||||||
|
// The first failure stops the rest: requests in flight are aborted
|
||||||
|
// (their AbortErrors are not reported) and no new ones start.
|
||||||
|
if (!abort.signal.aborted) {
|
||||||
|
abort.abort();
|
||||||
|
throw e;
|
||||||
|
}
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
await Promise.all(Array.from({ length: Math.min(parallel, n) }, worker));
|
||||||
|
return out;
|
||||||
|
}
|
||||||
@@ -6,11 +6,25 @@
|
|||||||
//! with no typed-array mapping (compound, reference, opaque, ...) is refused
|
//! with no typed-array mapping (compound, reference, opaque, ...) is refused
|
||||||
//! with a message naming it, never returned as reinterpreted bytes.
|
//! with a message naming it, never returned as reinterpreted bytes.
|
||||||
|
|
||||||
|
use std::sync::Arc;
|
||||||
|
|
||||||
use clawhdf5::{AttrValue, File, Selection};
|
use clawhdf5::{AttrValue, File, Selection};
|
||||||
use clawhdf5_format::data_read;
|
use clawhdf5_format::data_read;
|
||||||
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
||||||
|
use clawhdf5_format::message_type::MessageType;
|
||||||
|
use clawhdf5_format::object_header::ObjectHeader;
|
||||||
|
use clawhdf5_format::storage::Storage;
|
||||||
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
||||||
|
|
||||||
|
/// The most memory one read may use while it decodes: the stored bytes,
|
||||||
|
/// the values at 64 bits (integers are widened first) and the values
|
||||||
|
/// returned. A larger read fails with an error naming `readHyperslab`,
|
||||||
|
/// before anything is read: on wasm32 a buffer past 2 GiB cannot be
|
||||||
|
/// allocated at all, and failing to allocate aborts the module (every open
|
||||||
|
/// file on the page with it). 1 GiB leaves room in wasm32's 4 GiB for the
|
||||||
|
/// file's cached blocks and the JavaScript copy of the result.
|
||||||
|
pub const MAX_READ_BYTES: u64 = 1 << 30;
|
||||||
|
|
||||||
/// Errors are reported to JavaScript as messages.
|
/// Errors are reported to JavaScript as messages.
|
||||||
pub type Result<T> = std::result::Result<T, String>;
|
pub type Result<T> = std::result::Result<T, String>;
|
||||||
|
|
||||||
@@ -119,7 +133,9 @@ pub struct Attr {
|
|||||||
pub value: AttrValue,
|
pub value: AttrValue,
|
||||||
}
|
}
|
||||||
|
|
||||||
/// An open file, held in memory.
|
/// An open file: held in memory ([`Reader::open`]) or read through a
|
||||||
|
/// [`Storage`] ([`Reader::open_storage`], such as a
|
||||||
|
/// [`LazyStorage`](crate::lazy::LazyStorage)).
|
||||||
pub struct Reader {
|
pub struct Reader {
|
||||||
file: File,
|
file: File,
|
||||||
}
|
}
|
||||||
@@ -132,6 +148,14 @@ impl Reader {
|
|||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Open a file read through `storage` (the file's bytes from offset 0,
|
||||||
|
/// user block included, as [`File::open_storage`] takes them).
|
||||||
|
pub fn open_storage(storage: Arc<dyn Storage + Send + Sync>) -> Result<Self> {
|
||||||
|
Ok(Self {
|
||||||
|
file: File::open_storage(storage).map_err(err)?,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
/// Whether `path` names a group or a dataset.
|
/// Whether `path` names a group or a dataset.
|
||||||
pub fn kind(&self, path: &str) -> Result<Kind> {
|
pub fn kind(&self, path: &str) -> Result<Kind> {
|
||||||
match self.file.dataset(path) {
|
match self.file.dataset(path) {
|
||||||
@@ -144,31 +168,54 @@ impl Reader {
|
|||||||
/// The groups, then the datasets, in the group at `path` (`/` is the
|
/// The groups, then the datasets, in the group at `path` (`/` is the
|
||||||
/// root). Soft links are listed as their targets; external and dangling
|
/// root). Soft links are listed as their targets; external and dangling
|
||||||
/// links, and named datatypes, are left out.
|
/// links, and named datatypes, are left out.
|
||||||
|
///
|
||||||
|
/// What [`Group::groups`](clawhdf5::Group::groups) and `datasets` list,
|
||||||
|
/// but every child's object header is read before an error ends the
|
||||||
|
/// listing (the first error, in listing order, is the one returned, as
|
||||||
|
/// there). Over a [`LazyStorage`](crate::lazy::LazyStorage) that makes
|
||||||
|
/// one pass ask for all the headers it is missing at once, instead of
|
||||||
|
/// one pass, and one round trip, per header.
|
||||||
pub fn list(&self, path: &str) -> Result<Vec<Child>> {
|
pub fn list(&self, path: &str) -> Result<Vec<Child>> {
|
||||||
if self.kind(path)? != Kind::Group {
|
if self.kind(path)? != Kind::Group {
|
||||||
return Err(format!("not a group: {path}"));
|
return Err(format!("not a group: {path}"));
|
||||||
}
|
}
|
||||||
let group = self.file.group(path).map_err(err)?;
|
let group = self.file.group(path).map_err(err)?;
|
||||||
let mut out: Vec<Child> = group
|
let entries = group.entries().map_err(err)?;
|
||||||
.groups()
|
let sb = self.file.superblock();
|
||||||
.map_err(err)?
|
let storage = self.file.storage();
|
||||||
.into_iter()
|
let mut groups = Vec::new();
|
||||||
.map(|name| Child {
|
let mut datasets = Vec::new();
|
||||||
name,
|
let mut first_error = None;
|
||||||
kind: Kind::Group,
|
for (name, address) in entries {
|
||||||
})
|
match ObjectHeader::parse_in(storage, address, sb.offset_size, sb.length_size) {
|
||||||
.collect();
|
Ok(header) => {
|
||||||
out.extend(
|
let has = |t: MessageType| header.messages.iter().any(|m| m.msg_type == t);
|
||||||
group
|
if has(MessageType::LinkInfo)
|
||||||
.datasets()
|
|| has(MessageType::Link)
|
||||||
.map_err(err)?
|
|| has(MessageType::SymbolTable)
|
||||||
.into_iter()
|
{
|
||||||
.map(|name| Child {
|
groups.push(Child {
|
||||||
name,
|
name: name.clone(),
|
||||||
kind: Kind::Dataset,
|
kind: Kind::Group,
|
||||||
}),
|
});
|
||||||
);
|
}
|
||||||
Ok(out)
|
if has(MessageType::DataLayout) {
|
||||||
|
datasets.push(Child {
|
||||||
|
name,
|
||||||
|
kind: Kind::Dataset,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Err(e) => {
|
||||||
|
first_error.get_or_insert(e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if let Some(e) = first_error {
|
||||||
|
return Err(err(clawhdf5::Error::from(e)));
|
||||||
|
}
|
||||||
|
groups.extend(datasets);
|
||||||
|
Ok(groups)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Shape, max shape and datatype of the dataset at `path`.
|
/// Shape, max shape and datatype of the dataset at `path`.
|
||||||
@@ -228,13 +275,27 @@ impl Reader {
|
|||||||
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
|
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
|
||||||
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
|
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
|
||||||
}
|
}
|
||||||
let raw = ds.read_selection(&selection).map_err(err)?;
|
|
||||||
let data = self.decode(&raw, &dt)?;
|
|
||||||
out_shape.extend(element_shape(&dt));
|
out_shape.extend(element_shape(&dt));
|
||||||
let expected = out_shape
|
let expected = out_shape
|
||||||
.iter()
|
.iter()
|
||||||
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
|
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
|
||||||
.ok_or("selection size overflows")?;
|
.ok_or("selection size overflows")?;
|
||||||
|
let cost = expected.saturating_mul(bytes_per_value(&dt));
|
||||||
|
if cost > MAX_READ_BYTES {
|
||||||
|
return Err(format!(
|
||||||
|
"reading {path}{} would take about {} MiB of memory, more than the {} MiB \
|
||||||
|
one read may use; read it in parts (readHyperslab)",
|
||||||
|
if slab.is_some() {
|
||||||
|
" (this selection)"
|
||||||
|
} else {
|
||||||
|
" whole"
|
||||||
|
},
|
||||||
|
cost >> 20,
|
||||||
|
MAX_READ_BYTES >> 20
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let raw = ds.read_selection(&selection).map_err(err)?;
|
||||||
|
let data = self.decode(&raw, &dt)?;
|
||||||
if data.len() as u64 != expected {
|
if data.len() as u64 != expected {
|
||||||
return Err(format!(
|
return Err(format!(
|
||||||
"read {} values for shape {out_shape:?} ({expected} expected)",
|
"read {} values for shape {out_shape:?} ({expected} expected)",
|
||||||
@@ -281,11 +342,14 @@ impl Reader {
|
|||||||
// string ends at its first NUL and a heap object of the
|
// string ends at its first NUL and a heap object of the
|
||||||
// wrong size is an error, as in libhdf5 and h5py.
|
// wrong size is an error, as in libhdf5 and h5py.
|
||||||
let sb = self.file.superblock();
|
let sb = self.file.superblock();
|
||||||
Data::Strings(
|
let strings = match self.file.contiguous_bytes() {
|
||||||
VlResolver::new(self.file.as_bytes(), sb.offset_size, sb.length_size)
|
Some(bytes) => {
|
||||||
.strings(raw)
|
VlResolver::new(bytes, sb.offset_size, sb.length_size).strings(raw)
|
||||||
.map_err(err)?,
|
}
|
||||||
)
|
None => VlResolver::new_in(self.file.storage(), sb.offset_size, sb.length_size)
|
||||||
|
.strings(raw),
|
||||||
|
};
|
||||||
|
Data::Strings(strings.map_err(err)?)
|
||||||
}
|
}
|
||||||
Datatype::Enumeration { .. } if !is_array => {
|
Datatype::Enumeration { .. } if !is_array => {
|
||||||
Data::Strings(data_read::read_enum_names(raw, dt).map_err(err)?)
|
Data::Strings(data_read::read_enum_names(raw, dt).map_err(err)?)
|
||||||
@@ -300,6 +364,27 @@ impl Reader {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Memory one value of type `dt` takes while [`Reader::read`] decodes it
|
||||||
|
/// (an array type's elements count as values): its stored bytes, plus what
|
||||||
|
/// [`Reader::decode`] builds from them. A string counts its `String` (24
|
||||||
|
/// bytes on 64-bit targets, less on wasm32) and, for a fixed-length one,
|
||||||
|
/// its text; a variable-length string's text lives in the heap and is
|
||||||
|
/// bounded by the storage's own read limit.
|
||||||
|
fn bytes_per_value(dt: &Datatype) -> u64 {
|
||||||
|
let base = array_base(dt);
|
||||||
|
let stored = u64::from(base.type_size());
|
||||||
|
stored
|
||||||
|
+ match base {
|
||||||
|
Datatype::FloatingPoint { size, .. } if *size <= 4 => 4,
|
||||||
|
Datatype::FloatingPoint { .. } => 8,
|
||||||
|
// Widened to 64 bits, then narrowed to a new vector.
|
||||||
|
Datatype::FixedPoint { .. } => 8 + stored,
|
||||||
|
Datatype::String { .. } => 24 + stored,
|
||||||
|
Datatype::VariableLength { .. } | Datatype::Enumeration { .. } => 24,
|
||||||
|
_ => 0,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Narrow integers read at 64 bits to the dataset's own width. The source is
|
/// Narrow integers read at 64 bits to the dataset's own width. The source is
|
||||||
/// that width, so this cannot fail on correct input; it is checked anyway.
|
/// that width, so this cannot fail on correct input; it is checked anyway.
|
||||||
fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T>> {
|
fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T>> {
|
||||||
|
|||||||
@@ -0,0 +1,815 @@
|
|||||||
|
//! Reading a file that is not all here, when no read may wait for the
|
||||||
|
//! network: the restartable "NeedBytes" mode of `docs/design/range-reads.md`
|
||||||
|
//! (milestone M4).
|
||||||
|
//!
|
||||||
|
//! A browser's main thread cannot block on `fetch`, and the parsers are
|
||||||
|
//! synchronous. So an operation (open, list a group, read a dataset) runs
|
||||||
|
//! as a *pass* over a [`LazyStorage`] that holds the blocks fetched so far:
|
||||||
|
//!
|
||||||
|
//! 1. [`LazyStorage::attempt`] runs the operation. A read whose blocks are
|
||||||
|
//! all present is served; a read that misses records the missing blocks
|
||||||
|
//! and fails with a storage error.
|
||||||
|
//! 2. If the pass missed anything, its result is thrown away — whatever it
|
||||||
|
//! is, since a parser may have caught the error and carried on (a
|
||||||
|
//! listing skips a link it cannot resolve) — and the caller gets the
|
||||||
|
//! byte ranges to fetch ([`Step::Need`]).
|
||||||
|
//! 3. The caller fetches them (asynchronously, with HTTP `Range` requests),
|
||||||
|
//! hands them over with [`LazyStorage::supply`] and runs the operation
|
||||||
|
//! again.
|
||||||
|
//!
|
||||||
|
//! A pass is pure over the storage: the facade only caches what completed
|
||||||
|
//! reads decoded (its chunk cache), so re-running it is safe. Every pass
|
||||||
|
//! that does not finish asks for at least one block not yet present, and no
|
||||||
|
//! block is evicted while an operation is in flight
|
||||||
|
//! ([`LazyStorage::operation`]), so an operation finishes after at most one
|
||||||
|
//! pass per block it needs. In practice it is one pass per *wave* of
|
||||||
|
//! misses: a chunked read asks for all the chunks of a batch at once.
|
||||||
|
//!
|
||||||
|
//! Blocks are kept in an LRU cache with a byte budget, trimmed only when no
|
||||||
|
//! operation is in flight. Blocks fetched for bulk reads (raw data: a
|
||||||
|
//! `read_ranges` call, or a read longer than a block) go first, so reading
|
||||||
|
//! a large dataset does not evict the metadata.
|
||||||
|
|
||||||
|
use std::borrow::Cow;
|
||||||
|
use std::collections::{BTreeSet, HashMap};
|
||||||
|
use std::ops::Range;
|
||||||
|
use std::sync::{Arc, Mutex, MutexGuard};
|
||||||
|
|
||||||
|
use clawhdf5_format::error::FormatError;
|
||||||
|
use clawhdf5_format::storage::Storage;
|
||||||
|
|
||||||
|
/// Default block size: 1 MiB, as `clawhdf5-remote`'s block cache (the size
|
||||||
|
/// `docs/design/range-reads.md` §2 measured).
|
||||||
|
pub const DEFAULT_BLOCK_SIZE: u64 = 1 << 20;
|
||||||
|
|
||||||
|
/// Default of [`LazyConfig::max_fetch`]: 512 MiB, the same as `openUrl`'s
|
||||||
|
/// `maxDownload` for a server without range support.
|
||||||
|
pub const DEFAULT_MAX_FETCH: u64 = 512 << 20;
|
||||||
|
|
||||||
|
/// The message of the error a read that misses returns. It never reaches
|
||||||
|
/// the caller of [`LazyStorage::attempt`]: a pass that missed is re-run.
|
||||||
|
pub const NEED_BYTES: &str = "bytes not fetched yet (restartable read)";
|
||||||
|
|
||||||
|
/// Settings of a [`LazyStorage`].
|
||||||
|
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||||
|
pub struct LazyConfig {
|
||||||
|
/// Size of a block in bytes (at least 512); fetches are whole, aligned
|
||||||
|
/// blocks (the file's last block is shorter).
|
||||||
|
pub block_size: u64,
|
||||||
|
/// Byte budget of cached blocks between operations. An operation keeps
|
||||||
|
/// every block it needs until it finishes, whatever the budget.
|
||||||
|
pub capacity: u64,
|
||||||
|
/// Largest single range asked for, in bytes (whole blocks, at least
|
||||||
|
/// one); longer runs are split so they can be fetched in parallel.
|
||||||
|
pub max_request: u64,
|
||||||
|
/// Most bytes one operation may fetch (at least one block), and so the
|
||||||
|
/// longest single read: a read longer than this fails at once, before
|
||||||
|
/// anything is fetched, and so does an operation whose passes would
|
||||||
|
/// fetch more. The file's length comes from the server, so without
|
||||||
|
/// this a hostile file (a heap "collection" claiming 2 GiB) makes the
|
||||||
|
/// reader fetch and hold whatever it names; on wasm32 a buffer past
|
||||||
|
/// 2 GiB cannot even be allocated.
|
||||||
|
pub max_fetch: u64,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Default for LazyConfig {
|
||||||
|
fn default() -> Self {
|
||||||
|
LazyConfig {
|
||||||
|
block_size: DEFAULT_BLOCK_SIZE,
|
||||||
|
capacity: 64 << 20,
|
||||||
|
max_request: 8 << 20,
|
||||||
|
max_fetch: DEFAULT_MAX_FETCH,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// What a [`LazyStorage`] has done so far.
|
||||||
|
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
|
||||||
|
pub struct LazyStats {
|
||||||
|
/// Passes run by [`LazyStorage::attempt`].
|
||||||
|
pub passes: u64,
|
||||||
|
/// Ranges handed to [`LazyStorage::supply`]: one HTTP request each.
|
||||||
|
pub requests: u64,
|
||||||
|
/// Bytes handed to [`LazyStorage::supply`].
|
||||||
|
pub bytes_fetched: u64,
|
||||||
|
/// Blocks evicted to stay within the budget.
|
||||||
|
pub evictions: u64,
|
||||||
|
/// Bytes cached now.
|
||||||
|
pub cached_bytes: u64,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The outcome of one pass.
|
||||||
|
#[derive(Debug)]
|
||||||
|
pub enum Step<T> {
|
||||||
|
/// The pass read only bytes that were present: its result stands.
|
||||||
|
Done(T),
|
||||||
|
/// The pass missed: fetch these byte ranges (sorted, disjoint, block
|
||||||
|
/// aligned), [`supply`](LazyStorage::supply) them and run it again.
|
||||||
|
Need(Vec<Range<u64>>),
|
||||||
|
}
|
||||||
|
|
||||||
|
struct Block {
|
||||||
|
data: Arc<[u8]>,
|
||||||
|
/// Eviction order: bulk blocks (`false`) before metadata (`true`),
|
||||||
|
/// then least recently used first.
|
||||||
|
key: (bool, u64),
|
||||||
|
}
|
||||||
|
|
||||||
|
#[derive(Default)]
|
||||||
|
struct State {
|
||||||
|
blocks: HashMap<u64, Block>,
|
||||||
|
/// `(metadata?, tick, block index)`, in eviction order.
|
||||||
|
order: BTreeSet<(bool, u64, u64)>,
|
||||||
|
tick: u64,
|
||||||
|
bytes: u64,
|
||||||
|
/// Blocks the current pass missed, and whether a small read wanted
|
||||||
|
/// them (metadata).
|
||||||
|
missing: HashMap<u64, bool>,
|
||||||
|
/// Blocks a bulk read missed that have not been supplied yet: kept
|
||||||
|
/// as bulk when they arrive.
|
||||||
|
bulk_pending: BTreeSet<u64>,
|
||||||
|
/// Operations in flight: no eviction while any is.
|
||||||
|
active: u32,
|
||||||
|
stats: LazyStats,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A [`Storage`] over the blocks of a file fetched so far; a read of
|
||||||
|
/// anything else fails and is recorded, so the pass can be re-run once the
|
||||||
|
/// bytes arrive. See the [module documentation](self).
|
||||||
|
pub struct LazyStorage {
|
||||||
|
len: u64,
|
||||||
|
config: LazyConfig,
|
||||||
|
state: Mutex<State>,
|
||||||
|
}
|
||||||
|
|
||||||
|
fn lock(m: &Mutex<State>) -> MutexGuard<'_, State> {
|
||||||
|
m.lock().unwrap_or_else(std::sync::PoisonError::into_inner)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Keeps an operation's blocks cached until it is dropped; see
|
||||||
|
/// [`LazyStorage::operation`].
|
||||||
|
pub struct Operation<'a> {
|
||||||
|
storage: &'a LazyStorage,
|
||||||
|
/// Bytes fetched for this operation so far.
|
||||||
|
fetched: std::cell::Cell<u64>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Operation<'_> {
|
||||||
|
/// Count `ranges` against the operation's budget
|
||||||
|
/// ([`LazyConfig::max_fetch`]) before they are fetched: an error, and
|
||||||
|
/// nothing counted, if they would take it past the budget.
|
||||||
|
pub fn charge(&self, ranges: &[Range<u64>]) -> Result<(), String> {
|
||||||
|
let max = self.storage.config.max_fetch;
|
||||||
|
let total = ranges.iter().fold(self.fetched.get(), |n, r| {
|
||||||
|
n.saturating_add(r.end.saturating_sub(r.start))
|
||||||
|
});
|
||||||
|
if total > max {
|
||||||
|
return Err(format!(
|
||||||
|
"this call would fetch more than {max} bytes of the file (the maxFetch limit); \
|
||||||
|
read less at a time (readHyperslab) or raise maxFetch"
|
||||||
|
));
|
||||||
|
}
|
||||||
|
self.fetched.set(total);
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Drop for Operation<'_> {
|
||||||
|
fn drop(&mut self) {
|
||||||
|
let mut st = lock(&self.storage.state);
|
||||||
|
st.active = st.active.saturating_sub(1);
|
||||||
|
if st.active == 0 {
|
||||||
|
self.storage.evict(&mut st);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl LazyStorage {
|
||||||
|
/// An empty cache for a file of `len` bytes.
|
||||||
|
pub fn new(len: u64, mut config: LazyConfig) -> Self {
|
||||||
|
config.block_size = config.block_size.max(512);
|
||||||
|
config.max_request = (config.max_request / config.block_size).max(1) * config.block_size;
|
||||||
|
config.max_fetch = config.max_fetch.max(config.block_size);
|
||||||
|
LazyStorage {
|
||||||
|
len,
|
||||||
|
config,
|
||||||
|
state: Mutex::new(State::default()),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The settings in use (after rounding).
|
||||||
|
pub fn config(&self) -> &LazyConfig {
|
||||||
|
&self.config
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Counters since the storage was made.
|
||||||
|
pub fn stats(&self) -> LazyStats {
|
||||||
|
let st = lock(&self.state);
|
||||||
|
LazyStats {
|
||||||
|
cached_bytes: st.bytes,
|
||||||
|
..st.stats
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Mark an operation in flight until the guard is dropped: no block is
|
||||||
|
/// evicted meanwhile, so re-running its passes always makes progress.
|
||||||
|
/// Hold it across every pass of one operation.
|
||||||
|
pub fn operation(&self) -> Operation<'_> {
|
||||||
|
lock(&self.state).active += 1;
|
||||||
|
Operation {
|
||||||
|
storage: self,
|
||||||
|
fetched: std::cell::Cell::new(0),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Run one pass of `f` over this storage. `Done` when `f` read nothing
|
||||||
|
/// that is missing; otherwise `Need` with the ranges to fetch, and `f`'s
|
||||||
|
/// result is dropped (it may be an error caused by the miss, or a
|
||||||
|
/// result built around one).
|
||||||
|
pub fn attempt<T>(&self, f: impl FnOnce() -> T) -> Step<T> {
|
||||||
|
{
|
||||||
|
let mut st = lock(&self.state);
|
||||||
|
st.missing.clear();
|
||||||
|
st.stats.passes += 1;
|
||||||
|
}
|
||||||
|
let out = f();
|
||||||
|
let missing = std::mem::take(&mut lock(&self.state).missing);
|
||||||
|
if missing.is_empty() {
|
||||||
|
return Step::Done(out);
|
||||||
|
}
|
||||||
|
drop(out);
|
||||||
|
Step::Need(self.runs(missing))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The bytes of the file at `offset`, fetched for a range a pass asked
|
||||||
|
/// for. `offset` must be block aligned and the bytes whole blocks (the
|
||||||
|
/// last block of the file may be short) inside the file, or this is an
|
||||||
|
/// error and nothing is kept. Blocks already present are left alone.
|
||||||
|
pub fn supply(&self, offset: u64, bytes: &[u8]) -> Result<(), String> {
|
||||||
|
let bs = self.config.block_size;
|
||||||
|
let end = offset
|
||||||
|
.checked_add(bytes.len() as u64)
|
||||||
|
.filter(|&e| e <= self.len)
|
||||||
|
.ok_or_else(|| {
|
||||||
|
format!(
|
||||||
|
"{} bytes at offset {offset} run past the end of the {}-byte file",
|
||||||
|
bytes.len(),
|
||||||
|
self.len
|
||||||
|
)
|
||||||
|
})?;
|
||||||
|
if !offset.is_multiple_of(bs) || (!end.is_multiple_of(bs) && end != self.len) {
|
||||||
|
return Err(format!(
|
||||||
|
"{} bytes at offset {offset} are not whole {bs}-byte blocks",
|
||||||
|
bytes.len()
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let mut st = lock(&self.state);
|
||||||
|
st.stats.requests += 1;
|
||||||
|
st.stats.bytes_fetched += bytes.len() as u64;
|
||||||
|
let mut start = offset;
|
||||||
|
while start < end {
|
||||||
|
let i = start / bs;
|
||||||
|
let stop = (start + bs).min(end);
|
||||||
|
if !st.blocks.contains_key(&i) {
|
||||||
|
let rel = (start - offset) as usize..(stop - offset) as usize;
|
||||||
|
let metadata = !st.bulk_pending.remove(&i);
|
||||||
|
self.keep(&mut st, i, Arc::from(&bytes[rel]), metadata);
|
||||||
|
}
|
||||||
|
start = stop;
|
||||||
|
}
|
||||||
|
if st.active == 0 {
|
||||||
|
self.evict(&mut st);
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// [`supply`](Self::supply) the bytes fetched for `range`, one of the
|
||||||
|
/// ranges a [`Step::Need`] asked for: anything but exactly its length
|
||||||
|
/// (a server that answered with more or less) is an error.
|
||||||
|
pub fn supply_range(&self, range: &Range<u64>, bytes: &[u8]) -> Result<(), String> {
|
||||||
|
let want = range.end.saturating_sub(range.start);
|
||||||
|
if bytes.len() as u64 != want {
|
||||||
|
return Err(format!(
|
||||||
|
"asked for {want} bytes at offset {}, got {}",
|
||||||
|
range.start,
|
||||||
|
bytes.len()
|
||||||
|
));
|
||||||
|
}
|
||||||
|
self.supply(range.start, bytes)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Run `f` to completion, fetching what its passes miss with `fetch`
|
||||||
|
/// (a byte range to its bytes). The blocking driver, for native code
|
||||||
|
/// and tests; the browser's is the same loop with an `await` between
|
||||||
|
/// passes.
|
||||||
|
pub fn run_blocking<T>(
|
||||||
|
&self,
|
||||||
|
mut f: impl FnMut() -> T,
|
||||||
|
mut fetch: impl FnMut(Range<u64>) -> Result<Vec<u8>, String>,
|
||||||
|
) -> Result<T, String> {
|
||||||
|
let op = self.operation();
|
||||||
|
loop {
|
||||||
|
match self.attempt(&mut f) {
|
||||||
|
Step::Done(v) => return Ok(v),
|
||||||
|
Step::Need(ranges) => {
|
||||||
|
op.charge(&ranges)?;
|
||||||
|
for r in ranges {
|
||||||
|
let bytes = fetch(r.clone())?;
|
||||||
|
self.supply_range(&r, &bytes)?;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Cache block `i`.
|
||||||
|
fn keep(&self, st: &mut State, i: u64, data: Arc<[u8]>, metadata: bool) {
|
||||||
|
st.tick += 1;
|
||||||
|
let key = (metadata, st.tick);
|
||||||
|
st.bytes += <[u8]>::len(&data) as u64;
|
||||||
|
st.order.insert((key.0, key.1, i));
|
||||||
|
if let Some(old) = st.blocks.insert(i, Block { data, key }) {
|
||||||
|
st.order.remove(&(old.key.0, old.key.1, i));
|
||||||
|
st.bytes -= <[u8]>::len(&old.data) as u64;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn evict(&self, st: &mut State) {
|
||||||
|
while st.bytes > self.config.capacity {
|
||||||
|
let Some((_, _, i)) = st.order.pop_first() else {
|
||||||
|
break;
|
||||||
|
};
|
||||||
|
if let Some(b) = st.blocks.remove(&i) {
|
||||||
|
st.bytes -= <[u8]>::len(&b.data) as u64;
|
||||||
|
st.stats.evictions += 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Byte ranges covering the missing blocks: runs of consecutive
|
||||||
|
/// blocks, a one-block hole between two runs filled so they merge
|
||||||
|
/// (unless the hole is cached: it would be fetched again), each at
|
||||||
|
/// most `max_request` long.
|
||||||
|
fn runs(&self, missing: HashMap<u64, bool>) -> Vec<Range<u64>> {
|
||||||
|
let bs = self.config.block_size;
|
||||||
|
let mut wanted: Vec<u64> = missing.keys().copied().collect();
|
||||||
|
wanted.sort_unstable();
|
||||||
|
let mut st = lock(&self.state);
|
||||||
|
// Remember which blocks only bulk reads asked for: they are kept
|
||||||
|
// as bulk once supplied.
|
||||||
|
for (&i, &metadata) in &missing {
|
||||||
|
if metadata {
|
||||||
|
st.bulk_pending.remove(&i);
|
||||||
|
} else {
|
||||||
|
st.bulk_pending.insert(i);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let per_request = self.config.max_request / bs;
|
||||||
|
let mut runs: Vec<(u64, u64)> = Vec::new();
|
||||||
|
for i in wanted {
|
||||||
|
match runs.last_mut() {
|
||||||
|
Some((first, last))
|
||||||
|
if (i == *last + 1
|
||||||
|
|| (i == *last + 2 && !st.blocks.contains_key(&(i - 1))))
|
||||||
|
&& i - *first < per_request =>
|
||||||
|
{
|
||||||
|
*last = i
|
||||||
|
}
|
||||||
|
_ => runs.push((i, i)),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
drop(st);
|
||||||
|
runs.into_iter()
|
||||||
|
.map(|(a, b)| a * bs..((b + 1) * bs).min(self.len))
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Block indices covering `[offset, offset + len)`, clamped to the file.
|
||||||
|
fn span(&self, offset: u64, len: u64) -> Option<Range<u64>> {
|
||||||
|
let end = offset.saturating_add(len).min(self.len);
|
||||||
|
if offset >= end {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
let bs = self.config.block_size;
|
||||||
|
Some(offset / bs..(end - 1) / bs + 1)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The blocks of `spans` if all are present (touching them), else
|
||||||
|
/// record the missing ones and fail.
|
||||||
|
fn blocks(
|
||||||
|
&self,
|
||||||
|
spans: &[Range<u64>],
|
||||||
|
metadata: bool,
|
||||||
|
) -> Result<HashMap<u64, Arc<[u8]>>, FormatError> {
|
||||||
|
let mut st = lock(&self.state);
|
||||||
|
let mut have = HashMap::new();
|
||||||
|
let mut missed = false;
|
||||||
|
for span in spans {
|
||||||
|
for i in span.clone() {
|
||||||
|
if have.contains_key(&i) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
match st.blocks.get(&i) {
|
||||||
|
Some(b) => {
|
||||||
|
have.insert(i, b.data.clone());
|
||||||
|
}
|
||||||
|
None => {
|
||||||
|
missed = true;
|
||||||
|
let m = st.missing.entry(i).or_insert(metadata);
|
||||||
|
*m |= metadata;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if missed {
|
||||||
|
return Err(FormatError::Storage(NEED_BYTES.into()));
|
||||||
|
}
|
||||||
|
// Touch: most recently used last; a small read promotes a bulk
|
||||||
|
// block to metadata.
|
||||||
|
for &i in have.keys() {
|
||||||
|
st.tick += 1;
|
||||||
|
let tick = st.tick;
|
||||||
|
let Some(b) = st.blocks.get_mut(&i) else {
|
||||||
|
continue;
|
||||||
|
};
|
||||||
|
let old = b.key;
|
||||||
|
b.key = (old.0 || metadata, tick);
|
||||||
|
let new = b.key;
|
||||||
|
st.order.remove(&(old.0, old.1, i));
|
||||||
|
st.order.insert((new.0, new.1, i));
|
||||||
|
}
|
||||||
|
Ok(have)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Refuse a read of `n` bytes longer than an operation may fetch
|
||||||
|
/// ([`LazyConfig::max_fetch`]), before its blocks are asked for.
|
||||||
|
fn check_len(&self, n: u64) -> Result<(), FormatError> {
|
||||||
|
let max = self.config.max_fetch;
|
||||||
|
if n > max {
|
||||||
|
return Err(FormatError::Storage(format!(
|
||||||
|
"a read of {n} bytes is more than one call may fetch ({max} bytes, the maxFetch limit)"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The bytes `offset..end` from `blocks`, which hold every block of
|
||||||
|
/// that span. The buffer is reserved fallibly: a length the address
|
||||||
|
/// space cannot hold (past `isize::MAX` on wasm32) is an error, never
|
||||||
|
/// an abort.
|
||||||
|
fn assemble(
|
||||||
|
&self,
|
||||||
|
offset: u64,
|
||||||
|
end: u64,
|
||||||
|
blocks: &HashMap<u64, Arc<[u8]>>,
|
||||||
|
) -> Result<Vec<u8>, FormatError> {
|
||||||
|
let bs = self.config.block_size;
|
||||||
|
let n = end - offset;
|
||||||
|
let too_long =
|
||||||
|
|| FormatError::Storage(format!("cannot hold a read of {n} bytes in memory"));
|
||||||
|
let mut out = Vec::new();
|
||||||
|
out.try_reserve_exact(usize::try_from(n).map_err(|_| too_long())?)
|
||||||
|
.map_err(|_| too_long())?;
|
||||||
|
let mut pos = offset;
|
||||||
|
while pos < end {
|
||||||
|
let i = pos / bs;
|
||||||
|
let block = &blocks[&i];
|
||||||
|
let from = (pos - i * bs) as usize;
|
||||||
|
let to = ((end - i * bs) as usize).min(<[u8]>::len(block));
|
||||||
|
out.extend_from_slice(&block[from..to]);
|
||||||
|
pos = i * bs + to as u64;
|
||||||
|
}
|
||||||
|
Ok(out)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Storage for LazyStorage {
|
||||||
|
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
|
||||||
|
let Some(span) = self.span(offset, len as u64) else {
|
||||||
|
return Ok(Cow::Owned(Vec::new()));
|
||||||
|
};
|
||||||
|
let end = offset.saturating_add(len as u64).min(self.len);
|
||||||
|
self.check_len(end - offset)?;
|
||||||
|
let metadata = len as u64 <= self.config.block_size;
|
||||||
|
let blocks = self.blocks(std::slice::from_ref(&span), metadata)?;
|
||||||
|
Ok(Cow::Owned(self.assemble(offset, end, &blocks)?))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn len(&self) -> u64 {
|
||||||
|
self.len
|
||||||
|
}
|
||||||
|
|
||||||
|
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
|
||||||
|
let mut spans = Vec::with_capacity(ranges.len());
|
||||||
|
let mut total = 0u64;
|
||||||
|
for r in ranges {
|
||||||
|
if r.end < r.start {
|
||||||
|
return Err(FormatError::Storage(
|
||||||
|
"read range ends before it starts".into(),
|
||||||
|
));
|
||||||
|
}
|
||||||
|
total = total.saturating_add(r.end.min(self.len).saturating_sub(r.start));
|
||||||
|
spans.extend(self.span(r.start, r.end - r.start));
|
||||||
|
}
|
||||||
|
self.check_len(total)?;
|
||||||
|
let blocks = self.blocks(&spans, false)?;
|
||||||
|
ranges
|
||||||
|
.iter()
|
||||||
|
.map(|r| {
|
||||||
|
let end = r.end.min(self.len);
|
||||||
|
if r.start >= end {
|
||||||
|
Ok(Cow::Owned(Vec::new()))
|
||||||
|
} else {
|
||||||
|
self.assemble(r.start, end, &blocks).map(Cow::Owned)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
#[allow(clippy::single_range_in_vec_init)]
|
||||||
|
mod tests {
|
||||||
|
use super::*;
|
||||||
|
|
||||||
|
fn file(n: usize) -> Vec<u8> {
|
||||||
|
(0..n).map(|i| (i * 7 + i / 251) as u8).collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn config(block: u64, capacity: u64) -> LazyConfig {
|
||||||
|
LazyConfig {
|
||||||
|
block_size: block,
|
||||||
|
capacity,
|
||||||
|
max_request: 4 * block,
|
||||||
|
max_fetch: DEFAULT_MAX_FETCH,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Supply every range of `need` from `data`.
|
||||||
|
fn serve(s: &LazyStorage, data: &[u8], need: &[Range<u64>]) {
|
||||||
|
for r in need {
|
||||||
|
s.supply_range(r, &data[r.start as usize..r.end as usize])
|
||||||
|
.unwrap();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn owned(r: Result<Cow<'_, [u8]>, FormatError>) -> Result<Vec<u8>, FormatError> {
|
||||||
|
r.map(Cow::into_owned)
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_miss_asks_for_whole_blocks_then_the_rerun_reads_them() {
|
||||||
|
let data = file(10_000);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||||
|
let Step::Need(need) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
|
||||||
|
panic!("nothing is cached yet");
|
||||||
|
};
|
||||||
|
assert_eq!(need, vec![1024..3072]);
|
||||||
|
serve(&s, &data, &need);
|
||||||
|
let Step::Done(got) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
|
||||||
|
panic!("the blocks were supplied");
|
||||||
|
};
|
||||||
|
assert_eq!(got.unwrap(), &data[1500..2500]);
|
||||||
|
// Past the end: short, then empty, as for a slice.
|
||||||
|
serve(&s, &data, &[9216..10_000]);
|
||||||
|
let Step::Done(tail) = s.attempt(|| owned(s.read_at(9_990, 100))) else {
|
||||||
|
panic!("the last block was supplied");
|
||||||
|
};
|
||||||
|
assert_eq!(tail.unwrap(), &data[9_990..]);
|
||||||
|
assert!(matches!(
|
||||||
|
s.attempt(|| s.read_at(20_000, 10).map(|c| c.len())),
|
||||||
|
Step::Done(Ok(0))
|
||||||
|
));
|
||||||
|
let st = s.stats();
|
||||||
|
assert_eq!(
|
||||||
|
(st.passes, st.requests, st.bytes_fetched),
|
||||||
|
(4, 2, 2048 + 784)
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_pass_that_swallowed_the_miss_is_still_rerun() {
|
||||||
|
// A parser that catches the error and returns something anyway
|
||||||
|
// (a listing skipping a link it cannot resolve) must not have its
|
||||||
|
// result used.
|
||||||
|
let data = file(4096);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||||
|
let step = s.attempt(|| s.read_at(0, 4).map(|b| b.len()).unwrap_or(0));
|
||||||
|
assert!(
|
||||||
|
matches!(step, Step::Need(ref n) if n == &vec![0..1024]),
|
||||||
|
"{step:?}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn read_ranges_asks_for_every_missing_block_in_one_pass() {
|
||||||
|
let data = file(64 * 1024);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||||
|
let ranges = [
|
||||||
|
100..200,
|
||||||
|
5000..5100,
|
||||||
|
5200..5300,
|
||||||
|
30_000..33_000,
|
||||||
|
60_000..60_010,
|
||||||
|
];
|
||||||
|
let read = || {
|
||||||
|
s.read_ranges(&ranges)
|
||||||
|
.map(|v| v.into_iter().map(Cow::into_owned).collect::<Vec<_>>())
|
||||||
|
};
|
||||||
|
let Step::Need(need) = s.attempt(read) else {
|
||||||
|
panic!("nothing is cached yet");
|
||||||
|
};
|
||||||
|
// 5000..5300 is blocks 4 and 5; 30_000..33_000 is blocks 29..=32.
|
||||||
|
assert_eq!(
|
||||||
|
need,
|
||||||
|
vec![
|
||||||
|
0..1024,
|
||||||
|
4096..6144,
|
||||||
|
29 * 1024..33 * 1024,
|
||||||
|
58 * 1024..59 * 1024
|
||||||
|
]
|
||||||
|
);
|
||||||
|
serve(&s, &data, &need);
|
||||||
|
let Step::Done(got) = s.attempt(read) else {
|
||||||
|
panic!("one pass fetched everything");
|
||||||
|
};
|
||||||
|
for (r, g) in ranges.iter().zip(&got.unwrap()) {
|
||||||
|
assert_eq!(g, &data[r.start as usize..r.end as usize]);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn runs_merge_one_block_holes_and_split_long_runs() {
|
||||||
|
let data = file(32 * 1024);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||||
|
// Blocks 0 and 2 (a hole of one: merged), 5..=14 (split in fours).
|
||||||
|
let Step::Need(need) = s.attempt(|| {
|
||||||
|
let _ = s.read_at(0, 10);
|
||||||
|
let _ = s.read_at(2048, 10);
|
||||||
|
s.read_ranges(&[5120..15 * 1024]).map(|_| ())
|
||||||
|
}) else {
|
||||||
|
panic!("nothing is cached yet");
|
||||||
|
};
|
||||||
|
assert_eq!(
|
||||||
|
need,
|
||||||
|
vec![0..3072, 5120..9216, 9216..13_312, 13_312..15_360]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_cached_hole_is_not_fetched_again() {
|
||||||
|
let data = file(8 * 1024);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||||
|
serve(&s, &data, &[1024..2048]);
|
||||||
|
// Blocks 0 and 2 missing, 1 cached: two requests, not 0..3072.
|
||||||
|
let Step::Need(need) = s.attempt(|| {
|
||||||
|
let _ = s.read_at(0, 10);
|
||||||
|
s.read_at(2048, 10).map(|_| ())
|
||||||
|
}) else {
|
||||||
|
panic!("blocks 0 and 2 are missing");
|
||||||
|
};
|
||||||
|
assert_eq!(need, vec![0..1024, 2048..3072]);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn supply_refuses_what_was_not_asked_for() {
|
||||||
|
let data = file(10_000);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||||
|
assert!(s.supply(1, &data[1..1025]).unwrap_err().contains("whole"));
|
||||||
|
assert!(s.supply(0, &data[..1000]).unwrap_err().contains("whole"));
|
||||||
|
assert!(
|
||||||
|
s.supply(9216, &[0u8; 1024])
|
||||||
|
.unwrap_err()
|
||||||
|
.contains("past the end")
|
||||||
|
);
|
||||||
|
for wrong in [&data[..1023], &data[..2048]] {
|
||||||
|
assert!(
|
||||||
|
s.supply_range(&(0..1024), wrong)
|
||||||
|
.unwrap_err()
|
||||||
|
.contains("asked for 1024 bytes")
|
||||||
|
);
|
||||||
|
}
|
||||||
|
assert_eq!(s.stats().cached_bytes, 0);
|
||||||
|
// The file's last, short block is whole.
|
||||||
|
s.supply(9216, &data[9216..]).unwrap();
|
||||||
|
assert_eq!(s.stats().cached_bytes, 784);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn an_operation_keeps_its_blocks_whatever_the_budget() {
|
||||||
|
// A budget of one block, an operation that needs eight: without
|
||||||
|
// the operation guard every pass would evict what the last one
|
||||||
|
// fetched and never finish.
|
||||||
|
let data = file(8 * 1024);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1024));
|
||||||
|
let got = s
|
||||||
|
.run_blocking(
|
||||||
|
|| {
|
||||||
|
(0..8)
|
||||||
|
.map(|i| owned(s.read_at(i * 1024 + 3, 10)))
|
||||||
|
.collect::<Result<Vec<_>, _>>()
|
||||||
|
},
|
||||||
|
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
|
||||||
|
)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap();
|
||||||
|
for (i, g) in got.iter().enumerate() {
|
||||||
|
assert_eq!(g, &data[i * 1024 + 3..i * 1024 + 13]);
|
||||||
|
}
|
||||||
|
// Trimmed to the budget once the operation is over.
|
||||||
|
let st = s.stats();
|
||||||
|
assert_eq!(st.cached_bytes, 1024);
|
||||||
|
assert_eq!(st.evictions, 7);
|
||||||
|
assert_eq!(st.passes, 9, "one pass per block, then the one that ends");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn bulk_blocks_are_evicted_before_metadata() {
|
||||||
|
let data = file(16 * 1024);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 4 * 1024));
|
||||||
|
let fetch = |r: Range<u64>| Ok(data[r.start as usize..r.end as usize].to_vec());
|
||||||
|
// Metadata: a small read of block 0.
|
||||||
|
s.run_blocking(|| s.read_at(0, 16).map(|_| ()), fetch)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap();
|
||||||
|
// Bulk: raw data over blocks 4..12, more than the budget.
|
||||||
|
s.run_blocking(|| s.read_ranges(&[4096..12 * 1024]).map(|_| ()), fetch)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap();
|
||||||
|
// The metadata block survived: reading it again is a hit.
|
||||||
|
let before = s.stats();
|
||||||
|
assert!(before.cached_bytes <= 4 * 1024);
|
||||||
|
assert!(matches!(
|
||||||
|
s.attempt(|| s.read_at(0, 16).map(|_| ())),
|
||||||
|
Step::Done(Ok(()))
|
||||||
|
));
|
||||||
|
assert_eq!(s.stats().requests, before.requests);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_read_longer_than_max_fetch_fails_without_fetching() {
|
||||||
|
// A hostile file names a 2 GiB heap collection in a "file" the
|
||||||
|
// server claims is 1 TiB: the read is refused before any block is
|
||||||
|
// asked for (on wasm32 its buffer could not even be allocated).
|
||||||
|
let s = LazyStorage::new(1 << 40, LazyConfig::default());
|
||||||
|
let step = s.attempt(|| s.read_at(4096, (1usize << 31) + 4096).map(|b| b.len()));
|
||||||
|
match step {
|
||||||
|
Step::Done(Err(e)) => assert!(e.to_string().contains("maxFetch"), "{e}"),
|
||||||
|
other => panic!("expected a refusal, got {other:?}"),
|
||||||
|
}
|
||||||
|
let step = s.attempt(|| s.read_ranges(&[0..(600 << 20)]).map(|v| v.len()));
|
||||||
|
assert!(matches!(step, Step::Done(Err(_))), "{step:?}");
|
||||||
|
assert_eq!(s.stats().requests, 0);
|
||||||
|
// At the limit it is an ordinary miss.
|
||||||
|
let s = LazyStorage::new(1 << 40, config(1024, 1 << 20));
|
||||||
|
let step = s.attempt(|| s.read_at(0, DEFAULT_MAX_FETCH as usize).map(|b| b.len()));
|
||||||
|
assert!(matches!(step, Step::Need(_)), "{step:?}");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn an_operation_stops_at_its_fetch_budget() {
|
||||||
|
// Many small reads, none over the limit, that together would fetch
|
||||||
|
// more than the budget: the operation fails before fetching past it.
|
||||||
|
let data = file(64 * 1024);
|
||||||
|
let mut c = config(1024, 1 << 20);
|
||||||
|
c.max_fetch = 8 * 1024;
|
||||||
|
let s = LazyStorage::new(data.len() as u64, c);
|
||||||
|
let e = s
|
||||||
|
.run_blocking(
|
||||||
|
|| {
|
||||||
|
(0..64)
|
||||||
|
.map(|i| owned(s.read_at(i * 1024, 8)))
|
||||||
|
.collect::<Result<Vec<_>, _>>()
|
||||||
|
},
|
||||||
|
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
|
||||||
|
)
|
||||||
|
.unwrap_err();
|
||||||
|
assert!(e.contains("maxFetch"), "{e}");
|
||||||
|
assert!(s.stats().bytes_fetched <= 8 * 1024, "{:?}", s.stats());
|
||||||
|
// Within the budget it completes, and the budget is per operation.
|
||||||
|
for _ in 0..3 {
|
||||||
|
s.run_blocking(
|
||||||
|
|| owned(s.read_at(10 * 1024, 3000)),
|
||||||
|
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
|
||||||
|
)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_failed_fetch_is_an_error_not_data() {
|
||||||
|
let data = file(4096);
|
||||||
|
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
|
||||||
|
let e = s
|
||||||
|
.run_blocking(|| owned(s.read_at(0, 8)), |_| Err("HTTP 500".into()))
|
||||||
|
.unwrap_err();
|
||||||
|
assert_eq!(e, "HTTP 500");
|
||||||
|
let e = s
|
||||||
|
.run_blocking(|| owned(s.read_at(0, 8)), |_| Ok(vec![0; 10]))
|
||||||
|
.unwrap_err();
|
||||||
|
assert!(e.contains("asked for 1024 bytes"), "{e}");
|
||||||
|
assert_eq!(s.stats().cached_bytes, 0);
|
||||||
|
drop(data);
|
||||||
|
}
|
||||||
|
}
|
||||||
+476
-62
@@ -1,7 +1,7 @@
|
|||||||
//! clawhdf5's HDF5 reader for JavaScript, via `wasm-bindgen`.
|
//! clawhdf5's HDF5 reader for JavaScript, via `wasm-bindgen`.
|
||||||
//!
|
//!
|
||||||
//! ```js
|
//! ```js
|
||||||
//! import init, { open } from "./pkg/clawhdf5_wasm.js";
|
//! import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
|
||||||
//! await init();
|
//! await init();
|
||||||
//! const file = open(new Uint8Array(await blob.arrayBuffer()));
|
//! const file = open(new Uint8Array(await blob.arrayBuffer()));
|
||||||
//! file.list("/"); // [{ name, kind: "group" | "dataset" }]
|
//! file.list("/"); // [{ name, kind: "group" | "dataset" }]
|
||||||
@@ -10,6 +10,12 @@
|
|||||||
//! file.read("/x"); // { shape, dtype, data: Float64Array | ... | string[] }
|
//! file.read("/x"); // { shape, dtype, data: Float64Array | ... | string[] }
|
||||||
//! file.readHyperslab("/x", [0, 0], [10, 10]); // stride, block optional
|
//! file.readHyperslab("/x", [0, 0], [10, 10]); // stride, block optional
|
||||||
//! file.free();
|
//! file.free();
|
||||||
|
//!
|
||||||
|
//! // A file on a web server, read by HTTP range requests as needed: the
|
||||||
|
//! // same methods, returning promises.
|
||||||
|
//! const remote = await openUrl("https://example.org/data.h5");
|
||||||
|
//! await remote.list("/");
|
||||||
|
//! remote.stats(); // { requests, bytesFetched, size, ... }
|
||||||
//! ```
|
//! ```
|
||||||
//!
|
//!
|
||||||
//! Numeric data comes back in the typed array of the stored width
|
//! Numeric data comes back in the typed array of the stored width
|
||||||
@@ -18,15 +24,23 @@
|
|||||||
//! Anything else is a thrown `Error` naming the datatype. Only the reader is
|
//! Anything else is a thrown `Error` naming the datatype. Only the reader is
|
||||||
//! exposed: nothing here writes files.
|
//! exposed: nothing here writes files.
|
||||||
//!
|
//!
|
||||||
//! The logic lives in [`core`], which is plain Rust and tested natively.
|
//! The logic lives in [`core`] and [`lazy`], which are plain Rust and tested
|
||||||
|
//! natively; `js/remote.js` does the HTTP.
|
||||||
|
|
||||||
pub mod core;
|
pub mod core;
|
||||||
|
pub mod lazy;
|
||||||
|
|
||||||
|
use std::ops::Range;
|
||||||
|
use std::rc::Rc;
|
||||||
|
use std::sync::Arc;
|
||||||
|
|
||||||
use clawhdf5::AttrValue;
|
use clawhdf5::AttrValue;
|
||||||
use js_sys::{Array, Object, Reflect};
|
use js_sys::{Array, Object, Promise, Reflect, Uint8Array};
|
||||||
use wasm_bindgen::prelude::*;
|
use wasm_bindgen::prelude::*;
|
||||||
|
use wasm_bindgen_futures::future_to_promise;
|
||||||
|
|
||||||
use crate::core::{Data, Hyperslab, Reader};
|
use crate::core::{Attr, Child, Data, DatasetInfo, Hyperslab, Reader};
|
||||||
|
use crate::lazy::{LazyConfig, LazyStorage, Step};
|
||||||
|
|
||||||
/// JavaScript numbers are exact up to 2^53.
|
/// JavaScript numbers are exact up to 2^53.
|
||||||
const MAX_SAFE_INTEGER: f64 = 9_007_199_254_740_991.0;
|
const MAX_SAFE_INTEGER: f64 = 9_007_199_254_740_991.0;
|
||||||
@@ -35,11 +49,30 @@ fn js_err(msg: String) -> JsError {
|
|||||||
JsError::new(&msg)
|
JsError::new(&msg)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// A JavaScript exception (from `fetch`, or `js/remote.js`) as a `JsError`
|
||||||
|
/// with its message.
|
||||||
|
fn js_exception(e: JsValue) -> JsError {
|
||||||
|
let msg = e
|
||||||
|
.dyn_ref::<js_sys::Error>()
|
||||||
|
.map(|e| String::from(e.message()))
|
||||||
|
.or_else(|| e.as_string())
|
||||||
|
.unwrap_or_else(|| format!("{e:?}"));
|
||||||
|
JsError::new(&msg)
|
||||||
|
}
|
||||||
|
|
||||||
fn set(obj: &Object, key: &str, value: impl Into<JsValue>) {
|
fn set(obj: &Object, key: &str, value: impl Into<JsValue>) {
|
||||||
// Defining a property on a fresh plain object cannot fail.
|
// Defining a property on a fresh plain object cannot fail.
|
||||||
Reflect::set(obj, &JsValue::from_str(key), &value.into()).unwrap_throw();
|
Reflect::set(obj, &JsValue::from_str(key), &value.into()).unwrap_throw();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn get(obj: &JsValue, key: &str) -> JsValue {
|
||||||
|
if obj.is_object() {
|
||||||
|
Reflect::get(obj, &JsValue::from_str(key)).unwrap_or(JsValue::UNDEFINED)
|
||||||
|
} else {
|
||||||
|
JsValue::UNDEFINED
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
fn shape_to_js(shape: &[u64]) -> Array {
|
fn shape_to_js(shape: &[u64]) -> Array {
|
||||||
shape.iter().map(|&d| JsValue::from_f64(d as f64)).collect()
|
shape.iter().map(|&d| JsValue::from_f64(d as f64)).collect()
|
||||||
}
|
}
|
||||||
@@ -58,6 +91,20 @@ fn indices_from_js(what: &str, v: &[f64]) -> Result<Vec<u64>, JsError> {
|
|||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn slab_from_js(
|
||||||
|
start: &[f64],
|
||||||
|
count: &[f64],
|
||||||
|
stride: Option<Vec<f64>>,
|
||||||
|
block: Option<Vec<f64>>,
|
||||||
|
) -> Result<Hyperslab, JsError> {
|
||||||
|
Ok(Hyperslab {
|
||||||
|
start: indices_from_js("start", start)?,
|
||||||
|
count: indices_from_js("count", count)?,
|
||||||
|
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
|
||||||
|
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
fn data_to_js(data: Data) -> JsValue {
|
fn data_to_js(data: Data) -> JsValue {
|
||||||
match data {
|
match data {
|
||||||
Data::F32(v) => js_sys::Float32Array::from(&v[..]).into(),
|
Data::F32(v) => js_sys::Float32Array::from(&v[..]).into(),
|
||||||
@@ -110,6 +157,72 @@ fn attr_to_js(value: AttrValue) -> (JsValue, Option<String>) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn list_to_js(children: Vec<Child>) -> Array {
|
||||||
|
children
|
||||||
|
.into_iter()
|
||||||
|
.map(|c| {
|
||||||
|
let o = Object::new();
|
||||||
|
set(&o, "name", c.name);
|
||||||
|
set(&o, "kind", c.kind.as_str());
|
||||||
|
JsValue::from(o)
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn info_to_js(i: DatasetInfo) -> Object {
|
||||||
|
let o = Object::new();
|
||||||
|
set(&o, "shape", shape_to_js(&i.shape));
|
||||||
|
let max: JsValue = match i.maxshape {
|
||||||
|
None => JsValue::NULL,
|
||||||
|
Some(dims) => dims
|
||||||
|
.into_iter()
|
||||||
|
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
|
||||||
|
.collect::<Array>()
|
||||||
|
.into(),
|
||||||
|
};
|
||||||
|
set(&o, "maxshape", max);
|
||||||
|
set(&o, "dtype", i.dtype);
|
||||||
|
set(&o, "elementShape", shape_to_js(&i.element_shape));
|
||||||
|
o
|
||||||
|
}
|
||||||
|
|
||||||
|
fn attrs_to_js(attrs: Vec<Attr>) -> Array {
|
||||||
|
attrs
|
||||||
|
.into_iter()
|
||||||
|
.map(|a| {
|
||||||
|
let o = Object::new();
|
||||||
|
set(&o, "name", a.name);
|
||||||
|
let (value, dtype) = attr_to_js(a.value);
|
||||||
|
set(&o, "value", value);
|
||||||
|
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
|
||||||
|
JsValue::from(o)
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn errors_to_js(errors: Vec<String>) -> Array {
|
||||||
|
errors.into_iter().map(JsValue::from).collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A read's values with the dataset's datatype.
|
||||||
|
fn values_to_js((dtype, a): (String, core::Array)) -> Object {
|
||||||
|
let o = Object::new();
|
||||||
|
set(&o, "shape", shape_to_js(&a.shape));
|
||||||
|
set(&o, "dtype", dtype);
|
||||||
|
set(&o, "data", data_to_js(a.data));
|
||||||
|
o
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read the dataset at `path` (whole, or `slab`) with its datatype.
|
||||||
|
fn read_values(
|
||||||
|
r: &Reader,
|
||||||
|
path: &str,
|
||||||
|
slab: Option<&Hyperslab>,
|
||||||
|
) -> core::Result<(String, core::Array)> {
|
||||||
|
let dtype = r.info(path)?.dtype;
|
||||||
|
Ok((dtype, r.read(path, slab)?))
|
||||||
|
}
|
||||||
|
|
||||||
/// An open HDF5 (or NetCDF-4) file.
|
/// An open HDF5 (or NetCDF-4) file.
|
||||||
#[wasm_bindgen]
|
#[wasm_bindgen]
|
||||||
pub struct H5File {
|
pub struct H5File {
|
||||||
@@ -145,38 +258,13 @@ impl H5File {
|
|||||||
|
|
||||||
/// The group's members: `[{ name, kind }]`, groups first.
|
/// The group's members: `[{ name, kind }]`, groups first.
|
||||||
pub fn list(&self, path: &str) -> Result<Array, JsError> {
|
pub fn list(&self, path: &str) -> Result<Array, JsError> {
|
||||||
Ok(self
|
Ok(list_to_js(self.inner.list(path).map_err(js_err)?))
|
||||||
.inner
|
|
||||||
.list(path)
|
|
||||||
.map_err(js_err)?
|
|
||||||
.into_iter()
|
|
||||||
.map(|c| {
|
|
||||||
let o = Object::new();
|
|
||||||
set(&o, "name", c.name);
|
|
||||||
set(&o, "kind", c.kind.as_str());
|
|
||||||
JsValue::from(o)
|
|
||||||
})
|
|
||||||
.collect())
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// `{ shape, maxshape, dtype, elementShape }`. `maxshape` is `null`
|
/// `{ shape, maxshape, dtype, elementShape }`. `maxshape` is `null`
|
||||||
/// when not recorded, with `null` for each unlimited dimension.
|
/// when not recorded, with `null` for each unlimited dimension.
|
||||||
pub fn info(&self, path: &str) -> Result<Object, JsError> {
|
pub fn info(&self, path: &str) -> Result<Object, JsError> {
|
||||||
let i = self.inner.info(path).map_err(js_err)?;
|
Ok(info_to_js(self.inner.info(path).map_err(js_err)?))
|
||||||
let o = Object::new();
|
|
||||||
set(&o, "shape", shape_to_js(&i.shape));
|
|
||||||
let max: JsValue = match i.maxshape {
|
|
||||||
None => JsValue::NULL,
|
|
||||||
Some(dims) => dims
|
|
||||||
.into_iter()
|
|
||||||
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
|
|
||||||
.collect::<Array>()
|
|
||||||
.into(),
|
|
||||||
};
|
|
||||||
set(&o, "maxshape", max);
|
|
||||||
set(&o, "dtype", i.dtype);
|
|
||||||
set(&o, "elementShape", shape_to_js(&i.element_shape));
|
|
||||||
Ok(o)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// `[{ name, value, dtype }]`, sorted by name. Scalars are `number`
|
/// `[{ name, value, dtype }]`, sorted by name. Scalars are `number`
|
||||||
@@ -185,31 +273,21 @@ impl H5File {
|
|||||||
/// and its `dtype`; one that could not be read at all is reported by
|
/// and its `dtype`; one that could not be read at all is reported by
|
||||||
/// [`attrErrors`](Self::attr_errors).
|
/// [`attrErrors`](Self::attr_errors).
|
||||||
pub fn attrs(&self, path: &str) -> Result<Array, JsError> {
|
pub fn attrs(&self, path: &str) -> Result<Array, JsError> {
|
||||||
let (attrs, _) = self.inner.attrs(path).map_err(js_err)?;
|
Ok(attrs_to_js(self.inner.attrs(path).map_err(js_err)?.0))
|
||||||
Ok(attrs
|
|
||||||
.into_iter()
|
|
||||||
.map(|a| {
|
|
||||||
let o = Object::new();
|
|
||||||
set(&o, "name", a.name);
|
|
||||||
let (value, dtype) = attr_to_js(a.value);
|
|
||||||
set(&o, "value", value);
|
|
||||||
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
|
|
||||||
JsValue::from(o)
|
|
||||||
})
|
|
||||||
.collect())
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Messages for attributes that could not be read.
|
/// Messages for attributes that could not be read.
|
||||||
#[wasm_bindgen(js_name = attrErrors)]
|
#[wasm_bindgen(js_name = attrErrors)]
|
||||||
pub fn attr_errors(&self, path: &str) -> Result<Array, JsError> {
|
pub fn attr_errors(&self, path: &str) -> Result<Array, JsError> {
|
||||||
let (_, errors) = self.inner.attrs(path).map_err(js_err)?;
|
Ok(errors_to_js(self.inner.attrs(path).map_err(js_err)?.1))
|
||||||
Ok(errors.into_iter().map(JsValue::from).collect())
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The whole dataset: `{ shape, dtype, data }`, `data` in row-major
|
/// The whole dataset: `{ shape, dtype, data }`, `data` in row-major
|
||||||
/// order.
|
/// order.
|
||||||
pub fn read(&self, path: &str) -> Result<Object, JsError> {
|
pub fn read(&self, path: &str) -> Result<Object, JsError> {
|
||||||
self.read_impl(path, None)
|
Ok(values_to_js(
|
||||||
|
read_values(&self.inner, path, None).map_err(js_err)?,
|
||||||
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// A regular hyperslab (`H5Sselect_hyperslab`): `stride` and `block`
|
/// A regular hyperslab (`H5Sselect_hyperslab`): `stride` and `block`
|
||||||
@@ -223,22 +301,358 @@ impl H5File {
|
|||||||
stride: Option<Vec<f64>>,
|
stride: Option<Vec<f64>>,
|
||||||
block: Option<Vec<f64>>,
|
block: Option<Vec<f64>>,
|
||||||
) -> Result<Object, JsError> {
|
) -> Result<Object, JsError> {
|
||||||
let slab = Hyperslab {
|
let slab = slab_from_js(&start, &count, stride, block)?;
|
||||||
start: indices_from_js("start", &start)?,
|
Ok(values_to_js(
|
||||||
count: indices_from_js("count", &count)?,
|
read_values(&self.inner, path, Some(&slab)).map_err(js_err)?,
|
||||||
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
|
))
|
||||||
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
|
}
|
||||||
};
|
}
|
||||||
self.read_impl(path, Some(&slab))
|
|
||||||
}
|
// ---------------------------------------------------------------------------
|
||||||
|
// Remote files: openUrl.
|
||||||
fn read_impl(&self, path: &str, slab: Option<&Hyperslab>) -> Result<Object, JsError> {
|
|
||||||
let dtype = self.inner.info(path).map_err(js_err)?.dtype;
|
#[wasm_bindgen(module = "/js/remote.js")]
|
||||||
let a = self.inner.read(path, slab).map_err(js_err)?;
|
extern "C" {
|
||||||
|
#[wasm_bindgen(catch)]
|
||||||
|
async fn probe(url: &str, first_len: f64, opts: &JsValue) -> Result<JsValue, JsValue>;
|
||||||
|
|
||||||
|
#[wasm_bindgen(catch, js_name = fetchRanges)]
|
||||||
|
async fn fetch_ranges(
|
||||||
|
url: &str,
|
||||||
|
ranges: Vec<f64>,
|
||||||
|
opts: &JsValue,
|
||||||
|
validator: &JsValue,
|
||||||
|
length: f64,
|
||||||
|
) -> Result<JsValue, JsValue>;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Where a remote file's bytes come from.
|
||||||
|
struct Http {
|
||||||
|
url: String,
|
||||||
|
opts: JsValue,
|
||||||
|
/// ETag or Last-Modified at open (`null` if the server sent neither).
|
||||||
|
validator: JsValue,
|
||||||
|
length: u64,
|
||||||
|
/// Requests the probe made (1, or 2 with a HEAD for the length).
|
||||||
|
probe_requests: u64,
|
||||||
|
}
|
||||||
|
|
||||||
|
enum Source {
|
||||||
|
/// Read by range requests through a restartable cache.
|
||||||
|
Lazy {
|
||||||
|
http: Http,
|
||||||
|
storage: Arc<LazyStorage>,
|
||||||
|
reader: Reader,
|
||||||
|
},
|
||||||
|
/// The server ignored `Range`: the whole file, downloaded at open.
|
||||||
|
Whole {
|
||||||
|
reader: Reader,
|
||||||
|
size: u64,
|
||||||
|
requests: u64,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Http {
|
||||||
|
/// Fetch `ranges` and hand them to `storage`.
|
||||||
|
async fn fetch(&self, storage: &LazyStorage, ranges: &[Range<u64>]) -> Result<(), JsError> {
|
||||||
|
let flat: Vec<f64> = ranges
|
||||||
|
.iter()
|
||||||
|
.flat_map(|r| [r.start as f64, r.end as f64])
|
||||||
|
.collect();
|
||||||
|
let got = fetch_ranges(
|
||||||
|
&self.url,
|
||||||
|
flat,
|
||||||
|
&self.opts,
|
||||||
|
&self.validator,
|
||||||
|
self.length as f64,
|
||||||
|
)
|
||||||
|
.await
|
||||||
|
.map_err(js_exception)?;
|
||||||
|
let got = Array::from(&got);
|
||||||
|
if got.length() as usize != ranges.len() {
|
||||||
|
return Err(js_err(format!(
|
||||||
|
"fetchRanges returned {} ranges for {}",
|
||||||
|
got.length(),
|
||||||
|
ranges.len()
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
for (r, bytes) in ranges.iter().zip(got.iter()) {
|
||||||
|
let bytes = Uint8Array::new(&bytes).to_vec();
|
||||||
|
storage.supply_range(r, &bytes).map_err(js_err)?;
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Run `f` over `storage` until it has every byte it reads.
|
||||||
|
async fn drive<T>(
|
||||||
|
&self,
|
||||||
|
storage: &LazyStorage,
|
||||||
|
mut f: impl FnMut() -> T,
|
||||||
|
) -> Result<T, JsError> {
|
||||||
|
let op = storage.operation();
|
||||||
|
loop {
|
||||||
|
match storage.attempt(&mut f) {
|
||||||
|
Step::Done(v) => return Ok(v),
|
||||||
|
Step::Need(ranges) => {
|
||||||
|
op.charge(&ranges).map_err(js_err)?;
|
||||||
|
self.fetch(storage, &ranges).await?
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Source {
|
||||||
|
async fn run<T>(&self, op: impl Fn(&Reader) -> core::Result<T>) -> Result<T, JsError> {
|
||||||
|
match self {
|
||||||
|
Source::Whole { reader, .. } => op(reader).map_err(js_err),
|
||||||
|
Source::Lazy {
|
||||||
|
http,
|
||||||
|
storage,
|
||||||
|
reader,
|
||||||
|
} => http.drive(storage, || op(reader)).await?.map_err(js_err),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A non-negative integer option, or `None` when not given.
|
||||||
|
fn int_opt(opts: &JsValue, key: &str, min: f64, max: f64) -> Result<Option<u64>, JsError> {
|
||||||
|
let v = get(opts, key);
|
||||||
|
if v.is_undefined() || v.is_null() {
|
||||||
|
return Ok(None);
|
||||||
|
}
|
||||||
|
match v.as_f64() {
|
||||||
|
Some(x) if x.fract() == 0.0 && (min..=max).contains(&x) => Ok(Some(x as u64)),
|
||||||
|
_ => Err(js_err(format!(
|
||||||
|
"openUrl: {key} must be an integer from {min} to {max}"
|
||||||
|
))),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Most bytes `maxFetch` and `maxDownload` may allow: 1 GiB. wasm32 has
|
||||||
|
/// 4 GiB of memory and no buffer past 2 GiB, and what is fetched is held
|
||||||
|
/// while it is decoded.
|
||||||
|
const MAX_FETCH_LIMIT: u64 = 1 << 30;
|
||||||
|
|
||||||
|
/// The largest file `openUrl` reads by ranges: on wasm32, 4 GiB - 1 bytes.
|
||||||
|
/// The format code turns file offsets into `usize` to use them (with a
|
||||||
|
/// clean error past it, see scripts/check-32bit-casts.sh), so on a 32-bit
|
||||||
|
/// target nothing at 4 GiB or beyond can be read; a larger file is refused
|
||||||
|
/// at open rather than failing on whichever read reaches past 4 GiB. On
|
||||||
|
/// 64-bit targets it is 2^53 - 1, the largest offset a JavaScript number
|
||||||
|
/// holds exactly.
|
||||||
|
const MAX_REMOTE_LENGTH: u64 = if (usize::MAX as u64) < MAX_SAFE_INTEGER as u64 {
|
||||||
|
usize::MAX as u64
|
||||||
|
} else {
|
||||||
|
MAX_SAFE_INTEGER as u64
|
||||||
|
};
|
||||||
|
|
||||||
|
fn config_from(opts: &JsValue) -> Result<LazyConfig, JsError> {
|
||||||
|
let mut c = LazyConfig::default();
|
||||||
|
if let Some(b) = int_opt(opts, "blockSize", 512.0, (64u64 << 20) as f64)? {
|
||||||
|
c.block_size = b;
|
||||||
|
}
|
||||||
|
if let Some(n) = int_opt(opts, "cacheSize", 0.0, MAX_SAFE_INTEGER)? {
|
||||||
|
c.capacity = n;
|
||||||
|
}
|
||||||
|
if let Some(n) = int_opt(opts, "maxFetch", 512.0, MAX_FETCH_LIMIT as f64)? {
|
||||||
|
c.max_fetch = n;
|
||||||
|
}
|
||||||
|
// Read by remote.js; checked here so a value wasm32 cannot hold is an
|
||||||
|
// option error rather than a download that cannot be kept.
|
||||||
|
int_opt(opts, "maxDownload", 0.0, MAX_FETCH_LIMIT as f64)?;
|
||||||
|
// Also read by remote.js (which checks it too, for direct callers).
|
||||||
|
int_opt(opts, "parallel", 1.0, 1024.0)?;
|
||||||
|
Ok(c)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Open the HDF5 file at `url` without downloading it: its bytes are
|
||||||
|
/// fetched with HTTP `Range` requests as the methods of the returned
|
||||||
|
/// [`RemoteFile`] need them, through a block cache.
|
||||||
|
///
|
||||||
|
/// `opts` (all optional):
|
||||||
|
/// - `blockSize` — bytes per request block, 512 to 64 MiB (default 1 MiB);
|
||||||
|
/// - `cacheSize` — bytes of blocks kept between calls (default 64 MiB);
|
||||||
|
/// - `maxFetch` — most bytes one call may fetch, and so the longest single
|
||||||
|
/// read, up to 1 GiB (default 512 MiB): a call that would fetch more
|
||||||
|
/// fails before fetching it;
|
||||||
|
/// - `fallback` — `"download"` (default) reads the whole file when the
|
||||||
|
/// server ignores `Range` (answers 200), up to `maxDownload` bytes
|
||||||
|
/// (default 512 MiB, at most 1 GiB); `"error"` refuses such a server;
|
||||||
|
/// - `headers`, `credentials` — passed to every `fetch` (`headers` as
|
||||||
|
/// `fetch` takes them: a `Headers`, `[name, value]` pairs or an object);
|
||||||
|
/// - `parallel` — range requests in flight at once, 1 to 1024 (default
|
||||||
|
/// 6); when one fails the others are aborted;
|
||||||
|
/// - `fetch` — a `fetch`-compatible function to use instead of the global.
|
||||||
|
///
|
||||||
|
/// Cross-origin servers must allow CORS and expose `Content-Range` (or
|
||||||
|
/// answer `HEAD` with `Content-Length`). A file may be up to 4 GiB - 1
|
||||||
|
/// bytes long (wasm32 offsets); a longer one is refused at open. A whole-dataset `read` that would use more than 1 GiB of
|
||||||
|
/// memory ([`core::MAX_READ_BYTES`]) is refused: read it in parts with
|
||||||
|
/// `readHyperslab`.
|
||||||
|
#[wasm_bindgen(js_name = openUrl)]
|
||||||
|
pub async fn open_url(url: String, opts: JsValue) -> Result<RemoteFile, JsError> {
|
||||||
|
let config = config_from(&opts)?;
|
||||||
|
let p = probe(&url, config.block_size as f64, &opts)
|
||||||
|
.await
|
||||||
|
.map_err(js_exception)?;
|
||||||
|
let requests = get(&p, "requests").as_f64().unwrap_or(1.0) as u64;
|
||||||
|
let whole = get(&p, "whole");
|
||||||
|
if !whole.is_undefined() {
|
||||||
|
let bytes = Uint8Array::new(&whole).to_vec();
|
||||||
|
let size = bytes.len() as u64;
|
||||||
|
let reader = Reader::open(bytes).map_err(js_err)?;
|
||||||
|
return Ok(RemoteFile {
|
||||||
|
inner: Rc::new(Source::Whole {
|
||||||
|
reader,
|
||||||
|
size,
|
||||||
|
requests,
|
||||||
|
}),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
let length = get(&p, "length")
|
||||||
|
.as_f64()
|
||||||
|
.filter(|x| x.fract() == 0.0 && (0.0..=MAX_SAFE_INTEGER).contains(x))
|
||||||
|
.ok_or_else(|| js_err(format!("{url}: the server gave no usable file size")))?;
|
||||||
|
if length as u64 > MAX_REMOTE_LENGTH {
|
||||||
|
return Err(js_err(format!(
|
||||||
|
"{url} is {length} bytes; openUrl reads files of up to {MAX_REMOTE_LENGTH} bytes \
|
||||||
|
(4 GiB - 1: the WebAssembly reader addresses a file with 32-bit offsets)"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
let http = Http {
|
||||||
|
url,
|
||||||
|
opts,
|
||||||
|
validator: get(&p, "validator"),
|
||||||
|
length: length as u64,
|
||||||
|
probe_requests: requests,
|
||||||
|
};
|
||||||
|
let storage = Arc::new(LazyStorage::new(http.length, config));
|
||||||
|
let first = Uint8Array::new(&get(&p, "first")).to_vec();
|
||||||
|
storage.supply(0, &first).map_err(js_err)?;
|
||||||
|
let s = storage.clone();
|
||||||
|
let reader = http
|
||||||
|
.drive(&storage, || Reader::open_storage(s.clone()))
|
||||||
|
.await?
|
||||||
|
.map_err(js_err)?;
|
||||||
|
Ok(RemoteFile {
|
||||||
|
inner: Rc::new(Source::Lazy {
|
||||||
|
http,
|
||||||
|
storage,
|
||||||
|
reader,
|
||||||
|
}),
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A file opened with [`openUrl`](open_url): the methods of [`H5File`],
|
||||||
|
/// each returning a `Promise` (it may have to fetch bytes first).
|
||||||
|
#[wasm_bindgen]
|
||||||
|
pub struct RemoteFile {
|
||||||
|
inner: Rc<Source>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl RemoteFile {
|
||||||
|
/// Run `op` (fetching what it needs) and convert its result.
|
||||||
|
fn call<T: 'static>(
|
||||||
|
&self,
|
||||||
|
op: impl Fn(&Reader) -> core::Result<T> + 'static,
|
||||||
|
to_js: impl FnOnce(T) -> JsValue + 'static,
|
||||||
|
) -> Promise {
|
||||||
|
let inner = self.inner.clone();
|
||||||
|
future_to_promise(async move {
|
||||||
|
let v = inner.run(op).await.map_err(JsValue::from)?;
|
||||||
|
Ok(to_js(v))
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[wasm_bindgen]
|
||||||
|
impl RemoteFile {
|
||||||
|
/// `"group"` or `"dataset"`.
|
||||||
|
#[wasm_bindgen(unchecked_return_type = "Promise<string>")]
|
||||||
|
pub fn kind(&self, path: String) -> Promise {
|
||||||
|
self.call(move |r| r.kind(&path), |k| k.as_str().into())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The group's members: `[{ name, kind }]`, groups first.
|
||||||
|
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
|
||||||
|
pub fn list(&self, path: String) -> Promise {
|
||||||
|
self.call(move |r| r.list(&path), |c| list_to_js(c).into())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `{ shape, maxshape, dtype, elementShape }`, as [`H5File::info`].
|
||||||
|
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
|
||||||
|
pub fn info(&self, path: String) -> Promise {
|
||||||
|
self.call(move |r| r.info(&path), |i| info_to_js(i).into())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `[{ name, value, dtype }]`, as [`H5File::attrs`].
|
||||||
|
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
|
||||||
|
pub fn attrs(&self, path: String) -> Promise {
|
||||||
|
self.call(move |r| r.attrs(&path), |a| attrs_to_js(a.0).into())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Messages for attributes that could not be read.
|
||||||
|
#[wasm_bindgen(js_name = attrErrors, unchecked_return_type = "Promise<Array<string>>")]
|
||||||
|
pub fn attr_errors(&self, path: String) -> Promise {
|
||||||
|
self.call(move |r| r.attrs(&path), |a| errors_to_js(a.1).into())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The whole dataset: `{ shape, dtype, data }`, as [`H5File::read`].
|
||||||
|
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
|
||||||
|
pub fn read(&self, path: String) -> Promise {
|
||||||
|
self.call(
|
||||||
|
move |r| read_values(r, &path, None),
|
||||||
|
|v| values_to_js(v).into(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A regular hyperslab, as [`H5File::read_hyperslab`]. Only the chunks
|
||||||
|
/// (or the contiguous runs) the selection touches are fetched.
|
||||||
|
#[wasm_bindgen(js_name = readHyperslab, unchecked_return_type = "Promise<any>")]
|
||||||
|
pub fn read_hyperslab(
|
||||||
|
&self,
|
||||||
|
path: String,
|
||||||
|
start: Vec<f64>,
|
||||||
|
count: Vec<f64>,
|
||||||
|
stride: Option<Vec<f64>>,
|
||||||
|
block: Option<Vec<f64>>,
|
||||||
|
) -> Result<Promise, JsError> {
|
||||||
|
let slab = slab_from_js(&start, &count, stride, block)?;
|
||||||
|
Ok(self.call(
|
||||||
|
move |r| read_values(r, &path, Some(&slab)),
|
||||||
|
|v| values_to_js(v).into(),
|
||||||
|
))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// What reading this file has cost so far: `{ lazy, size, requests,
|
||||||
|
/// bytesFetched, cachedBytes, passes }`. `lazy` is false when the
|
||||||
|
/// server ignored `Range` and the file was downloaded whole.
|
||||||
|
pub fn stats(&self) -> Object {
|
||||||
let o = Object::new();
|
let o = Object::new();
|
||||||
set(&o, "shape", shape_to_js(&a.shape));
|
match &*self.inner {
|
||||||
set(&o, "dtype", dtype);
|
Source::Lazy { http, storage, .. } => {
|
||||||
set(&o, "data", data_to_js(a.data));
|
let st = storage.stats();
|
||||||
Ok(o)
|
set(&o, "lazy", true);
|
||||||
|
set(&o, "size", http.length as f64);
|
||||||
|
set(
|
||||||
|
&o,
|
||||||
|
"requests",
|
||||||
|
(st.requests.saturating_sub(1) + http.probe_requests) as f64,
|
||||||
|
);
|
||||||
|
set(&o, "bytesFetched", st.bytes_fetched as f64);
|
||||||
|
set(&o, "cachedBytes", st.cached_bytes as f64);
|
||||||
|
set(&o, "passes", st.passes as f64);
|
||||||
|
}
|
||||||
|
Source::Whole { size, requests, .. } => {
|
||||||
|
set(&o, "lazy", false);
|
||||||
|
set(&o, "size", *size as f64);
|
||||||
|
set(&o, "requests", *requests as f64);
|
||||||
|
set(&o, "bytesFetched", *size as f64);
|
||||||
|
set(&o, "cachedBytes", *size as f64);
|
||||||
|
set(&o, "passes", 0.0);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
o
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,578 @@
|
|||||||
|
//! The restartable ("NeedBytes") reader against the in-memory one: every
|
||||||
|
//! file must list, describe and read the same through a [`LazyStorage`]
|
||||||
|
//! that starts empty and is fed only the ranges its passes ask for, as the
|
||||||
|
//! browser's `openUrl` feeds it from HTTP range requests.
|
||||||
|
//!
|
||||||
|
//! - Files written here with `FileBuilder`, at several block sizes (512 B
|
||||||
|
//! blocks make almost every structure read a miss).
|
||||||
|
//! - The h5py/netCDF4 fixture of `examples/wasm-viewer/test/make_fixture.py`
|
||||||
|
//! (skipped without h5py, unless `CLAWHDF5_REQUIRE_INTEROP=1`;
|
||||||
|
//! `CLAWHDF5_PYTHON` names the interpreter).
|
||||||
|
//! - `CLAWHDF5_WASM_CORPUS=dir[:dir...]`: every HDF5 file under those
|
||||||
|
//! directories up to 64 MiB (e.g. `conformance/.cache/corpus`).
|
||||||
|
//!
|
||||||
|
//! Also the request budget: listing and reading one small dataset of a large
|
||||||
|
//! file fetches a few blocks, not the file.
|
||||||
|
|
||||||
|
use std::ops::Range;
|
||||||
|
use std::path::{Path, PathBuf};
|
||||||
|
use std::process::Command;
|
||||||
|
use std::sync::Arc;
|
||||||
|
|
||||||
|
use clawhdf5::{AttrValue, FileBuilder};
|
||||||
|
use clawhdf5_format::storage::CountingStorage;
|
||||||
|
use clawhdf5_wasm::core::{Hyperslab, Kind, Reader};
|
||||||
|
use clawhdf5_wasm::lazy::{LazyConfig, LazyStorage};
|
||||||
|
|
||||||
|
/// The API the JavaScript side calls, one operation at a time.
|
||||||
|
trait Api {
|
||||||
|
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T;
|
||||||
|
}
|
||||||
|
|
||||||
|
struct Local(Reader);
|
||||||
|
|
||||||
|
impl Api for Local {
|
||||||
|
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
|
||||||
|
op(&self.0)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A lazily read file and the "server" it fetches from.
|
||||||
|
struct Lazy {
|
||||||
|
data: Arc<Vec<u8>>,
|
||||||
|
storage: Arc<LazyStorage>,
|
||||||
|
reader: Reader,
|
||||||
|
}
|
||||||
|
|
||||||
|
fn fetch(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
|
||||||
|
Ok(data[r.start as usize..r.end as usize].to_vec())
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Lazy {
|
||||||
|
/// Open as `openUrl` does: the first block comes with the probe that
|
||||||
|
/// learns the length, then the open is run until it has its bytes.
|
||||||
|
fn open(data: Vec<u8>, config: LazyConfig) -> Result<Lazy, String> {
|
||||||
|
let data = Arc::new(data);
|
||||||
|
let storage = Arc::new(LazyStorage::new(data.len() as u64, config));
|
||||||
|
let first = (storage.config().block_size as usize).min(data.len());
|
||||||
|
storage.supply(0, &data[..first])?;
|
||||||
|
let s = storage.clone();
|
||||||
|
let reader =
|
||||||
|
storage.run_blocking(|| Reader::open_storage(s.clone()), |r| fetch(&data, r))??;
|
||||||
|
Ok(Lazy {
|
||||||
|
data,
|
||||||
|
storage,
|
||||||
|
reader,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Api for Lazy {
|
||||||
|
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
|
||||||
|
self.storage
|
||||||
|
.run_blocking(|| op(&self.reader), |r| fetch(&self.data, r))
|
||||||
|
.expect("serving from memory cannot fail")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Everything the viewer can show of a file, as text: each object's kind,
|
||||||
|
/// listing, attributes (and attribute errors), dataset info, whole value
|
||||||
|
/// and a hyperslab — or the error each gives.
|
||||||
|
fn transcript(api: &impl Api) -> Vec<String> {
|
||||||
|
let mut out = Vec::new();
|
||||||
|
let mut todo = vec![("/".to_string(), 0usize)];
|
||||||
|
while let Some((path, depth)) = todo.pop() {
|
||||||
|
if out.len() > 4000 {
|
||||||
|
out.push("... (truncated)".into());
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
let kind = api.call(|r| r.kind(&path));
|
||||||
|
out.push(format!("{path}: {kind:?}"));
|
||||||
|
out.push(format!("{path} attrs: {:?}", api.call(|r| r.attrs(&path))));
|
||||||
|
match kind {
|
||||||
|
Ok(Kind::Group) => {
|
||||||
|
let list = api.call(|r| r.list(&path));
|
||||||
|
out.push(format!("{path} list: {list:?}"));
|
||||||
|
if let Ok(children) = list
|
||||||
|
&& depth < 12
|
||||||
|
{
|
||||||
|
for c in children.into_iter().rev() {
|
||||||
|
let child = if path == "/" {
|
||||||
|
format!("/{}", c.name)
|
||||||
|
} else {
|
||||||
|
format!("{path}/{}", c.name)
|
||||||
|
};
|
||||||
|
todo.push((child, depth + 1));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Ok(Kind::Dataset) => {
|
||||||
|
let info = api.call(|r| r.info(&path));
|
||||||
|
out.push(format!("{path} info: {info:?}"));
|
||||||
|
let Ok(info) = info else { continue };
|
||||||
|
let n = info
|
||||||
|
.shape
|
||||||
|
.iter()
|
||||||
|
.chain(&info.element_shape)
|
||||||
|
.try_fold(1u64, |a, &d| a.checked_mul(d));
|
||||||
|
if n.is_none_or(|n| n > 4_000_000) {
|
||||||
|
out.push(format!("{path}: not read ({n:?} values)"));
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
out.push(format!(
|
||||||
|
"{path} read: {:?}",
|
||||||
|
api.call(|r| r.read(&path, None))
|
||||||
|
));
|
||||||
|
if !info.shape.is_empty() && info.shape.iter().all(|&d| d > 1) {
|
||||||
|
let slab = Hyperslab {
|
||||||
|
start: info.shape.iter().map(|_| 1).collect(),
|
||||||
|
count: info.shape.iter().map(|&d| d / 2).collect(),
|
||||||
|
stride: None,
|
||||||
|
block: None,
|
||||||
|
};
|
||||||
|
let part = api.call(|r| r.read(&path, Some(&slab)));
|
||||||
|
out.push(format!("{path} slab: {part:?}"));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Err(_) => {}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
out
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The lazy transcript of `data` at `block` bytes per block equals the
|
||||||
|
/// transcript of the same file through a range storage that has every byte
|
||||||
|
/// (`CountingStorage`: the facade's `Storage` path, the one the lazy reader
|
||||||
|
/// takes), and agrees with the in-memory one: the same values, and an error
|
||||||
|
/// wherever it has one (a malformed file can fail at a different check,
|
||||||
|
/// with a different message, when read by ranges). Returns what the lazy
|
||||||
|
/// reader fetched and its transcript.
|
||||||
|
fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec<String>) {
|
||||||
|
let ctx = format!("{name} (blocks of {block} B)");
|
||||||
|
let ranged = Reader::open_storage(Arc::new(CountingStorage::new(data.to_vec())));
|
||||||
|
let local = Reader::open(data.to_vec());
|
||||||
|
let lazy = Lazy::open(data.to_vec(), config(block));
|
||||||
|
let (ranged, local, lazy) = match (ranged, local, lazy) {
|
||||||
|
(Ok(r), Ok(l), Ok(z)) => (r, l, z),
|
||||||
|
(Err(r), Err(_), Err(z)) => {
|
||||||
|
assert_eq!(z, r, "{ctx}: open error");
|
||||||
|
return (0, 0, Vec::new());
|
||||||
|
}
|
||||||
|
(r, l, z) => panic!(
|
||||||
|
"{ctx}: opens differently: ranged {:?}, in memory {:?}, lazily {:?}",
|
||||||
|
r.err(),
|
||||||
|
l.err(),
|
||||||
|
z.err()
|
||||||
|
),
|
||||||
|
};
|
||||||
|
let got = transcript(&lazy);
|
||||||
|
let want = transcript(&Local(ranged));
|
||||||
|
for (i, (w, g)) in want.iter().zip(&got).enumerate() {
|
||||||
|
assert_eq!(g, w, "{ctx}, line {i}");
|
||||||
|
}
|
||||||
|
assert_eq!(got.len(), want.len(), "{ctx}: transcript length");
|
||||||
|
let local = transcript(&Local(local));
|
||||||
|
for (i, (l, g)) in local.iter().zip(&got).enumerate() {
|
||||||
|
let both_errors = match (l.split_once("Err("), g.split_once("Err(")) {
|
||||||
|
(Some((a, _)), Some((b, _))) => a == b,
|
||||||
|
_ => false,
|
||||||
|
};
|
||||||
|
assert!(
|
||||||
|
l == g || both_errors,
|
||||||
|
"{ctx}, line {i}: in memory\n {l}\nlazily\n {g}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
assert_eq!(got.len(), local.len(), "{ctx}: transcript length");
|
||||||
|
let st = lazy.storage.stats();
|
||||||
|
(st.requests, st.bytes_fetched, got)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn config(block: u64) -> LazyConfig {
|
||||||
|
LazyConfig {
|
||||||
|
block_size: block,
|
||||||
|
// A small budget, so eviction between operations is exercised.
|
||||||
|
capacity: 16 * block,
|
||||||
|
max_request: 8 * block,
|
||||||
|
..LazyConfig::default()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn builder_file() -> Vec<u8> {
|
||||||
|
let mut b = FileBuilder::new();
|
||||||
|
b.create_dataset("grid")
|
||||||
|
.with_f64_data(&(0..20_000).map(f64::from).collect::<Vec<_>>())
|
||||||
|
.with_shape(&[100, 200])
|
||||||
|
.with_chunks(&[10, 25])
|
||||||
|
.with_deflate(4);
|
||||||
|
b.create_dataset("contiguous")
|
||||||
|
.with_i32_data(&(0..50_000).collect::<Vec<_>>());
|
||||||
|
b.create_dataset("bytes").with_u8_data(&[1, 2, 250]);
|
||||||
|
let mut g = b.create_group("sensors");
|
||||||
|
for i in 0..40 {
|
||||||
|
g.create_dataset(&format!("t{i}"))
|
||||||
|
.with_f32_data(&[i as f32, 1.5, -2.25]);
|
||||||
|
}
|
||||||
|
g.set_attr("location", AttrValue::String("lab".into()));
|
||||||
|
b.add_group(g.finish());
|
||||||
|
b.set_attr("version", AttrValue::I64(3));
|
||||||
|
b.set_attr("scale", AttrValue::F64Array(vec![0.5, 2.0]));
|
||||||
|
b.finish().unwrap()
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn builder_files_read_the_same_at_every_block_size() {
|
||||||
|
let data = builder_file();
|
||||||
|
for block in [512, 4096, 1 << 20] {
|
||||||
|
let (requests, _, lines) = check_equal("builder", &data, block);
|
||||||
|
assert!(requests > 0);
|
||||||
|
// The transcript covers every object, values included.
|
||||||
|
assert!(lines.iter().any(|l| l.starts_with("/grid read: Ok")));
|
||||||
|
assert!(lines.iter().any(|l| l.starts_with("/grid slab: Ok")));
|
||||||
|
assert!(lines.iter().any(|l| l.starts_with("/sensors/t39 read: Ok")));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn garbage_fails_to_open_as_in_memory() {
|
||||||
|
check_equal("zeros", &[0u8; 5000], 512);
|
||||||
|
check_equal("empty", &[], 512);
|
||||||
|
let mut cut = builder_file();
|
||||||
|
cut.truncate(cut.len() / 3);
|
||||||
|
check_equal("truncated", &cut, 512);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Listing a large file and reading one small dataset fetches a few blocks,
|
||||||
|
/// not the file.
|
||||||
|
#[test]
|
||||||
|
fn a_small_read_of_a_large_file_fetches_a_few_blocks() {
|
||||||
|
let mut b = FileBuilder::new();
|
||||||
|
b.create_dataset("small").with_f64_data(&[1.0, 2.0, 3.0]);
|
||||||
|
// 48 MB of raw data, written after the small dataset's metadata.
|
||||||
|
b.create_dataset("big")
|
||||||
|
.with_f64_data(&(0..6_000_000).map(f64::from).collect::<Vec<_>>());
|
||||||
|
let mut g = b.create_group("group");
|
||||||
|
g.create_dataset("inner").with_i32_data(&[7, 8]);
|
||||||
|
b.add_group(g.finish());
|
||||||
|
let data = b.finish().unwrap();
|
||||||
|
let lazy = Lazy::open(data.clone(), LazyConfig::default()).unwrap();
|
||||||
|
let list = lazy.call(|r| r.list("/")).unwrap();
|
||||||
|
assert_eq!(list.len(), 3);
|
||||||
|
assert_eq!(
|
||||||
|
format!("{:?}", lazy.call(|r| r.read("/small", None)).unwrap().data),
|
||||||
|
"F64([1.0, 2.0, 3.0])"
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
format!(
|
||||||
|
"{:?}",
|
||||||
|
lazy.call(|r| r.read("/group/inner", None)).unwrap().data
|
||||||
|
),
|
||||||
|
"I32([7, 8])"
|
||||||
|
);
|
||||||
|
// A window of the big dataset reads only its block(s).
|
||||||
|
let slab = Hyperslab {
|
||||||
|
start: vec![3_000_000],
|
||||||
|
count: vec![4],
|
||||||
|
stride: None,
|
||||||
|
block: None,
|
||||||
|
};
|
||||||
|
assert_eq!(
|
||||||
|
format!(
|
||||||
|
"{:?}",
|
||||||
|
lazy.call(|r| r.read("/big", Some(&slab))).unwrap().data
|
||||||
|
),
|
||||||
|
"F64([3000000.0, 3000001.0, 3000002.0, 3000003.0])"
|
||||||
|
);
|
||||||
|
let st = lazy.storage.stats();
|
||||||
|
eprintln!("{} bytes: {st:?}", data.len());
|
||||||
|
assert!(st.requests <= 6, "{st:?}");
|
||||||
|
assert!(st.bytes_fetched <= 6 << 20, "{st:?}");
|
||||||
|
assert!(st.bytes_fetched * 8 < data.len() as u64, "{st:?}");
|
||||||
|
}
|
||||||
|
|
||||||
|
fn python() -> String {
|
||||||
|
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn python_available() -> bool {
|
||||||
|
Command::new(python())
|
||||||
|
.args(["-c", "import h5py, netCDF4, numpy"])
|
||||||
|
.output()
|
||||||
|
.is_ok_and(|o| o.status.success())
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn h5py_and_netcdf4_files_read_the_same_lazily() {
|
||||||
|
if !python_available() {
|
||||||
|
assert!(
|
||||||
|
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||||
|
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
|
||||||
|
python()
|
||||||
|
);
|
||||||
|
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = fixture_dir();
|
||||||
|
for name in ["fixture.h5", "fixture.nc"] {
|
||||||
|
let data = std::fs::read(dir.path().join(name)).unwrap();
|
||||||
|
for block in [512, 64 * 1024] {
|
||||||
|
let (_, _, lines) = check_equal(name, &data, block);
|
||||||
|
assert!(lines.iter().filter(|l| l.contains(" read: Ok")).count() >= 2);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The bytes of `data` at `r`, zero past its end: a server that claims
|
||||||
|
/// the file is longer than it is.
|
||||||
|
fn fetch_padded(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
|
||||||
|
let mut out = vec![0u8; (r.end - r.start) as usize];
|
||||||
|
let len = data.len() as u64;
|
||||||
|
if r.start < len {
|
||||||
|
let end = r.end.min(len);
|
||||||
|
out[..(end - r.start) as usize].copy_from_slice(&data[r.start as usize..end as usize]);
|
||||||
|
}
|
||||||
|
Ok(out)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Sizes a hostile server or a large dataset can name are errors, never
|
||||||
|
/// allocations that abort the wasm module: make_fixture.py's limits.h5 and
|
||||||
|
/// hostile_vl.h5 (see write_limits there).
|
||||||
|
#[test]
|
||||||
|
fn size_limits_are_errors_not_aborts() {
|
||||||
|
if !python_available() {
|
||||||
|
assert!(
|
||||||
|
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||||
|
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
|
||||||
|
python()
|
||||||
|
);
|
||||||
|
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = fixture_dir();
|
||||||
|
|
||||||
|
// Read whole, /huge_u8 would widen 2^28 values to 64 bits (2 GiB): an
|
||||||
|
// error naming readHyperslab, before its chunks are read. A window of
|
||||||
|
// it reads.
|
||||||
|
let data = std::fs::read(dir.path().join("limits.h5")).unwrap();
|
||||||
|
let n = (1u64 << 28) + 1024;
|
||||||
|
let window = Hyperslab {
|
||||||
|
start: vec![n - 4],
|
||||||
|
count: vec![4],
|
||||||
|
stride: None,
|
||||||
|
block: None,
|
||||||
|
};
|
||||||
|
let local = Reader::open(data.clone()).unwrap();
|
||||||
|
let lazy = Lazy::open(data, LazyConfig::default()).unwrap();
|
||||||
|
let before = lazy.storage.stats().requests;
|
||||||
|
for e in [
|
||||||
|
local.read("/huge_u8", None).unwrap_err(),
|
||||||
|
lazy.call(|r| r.read("/huge_u8", None)).unwrap_err(),
|
||||||
|
] {
|
||||||
|
assert!(e.contains("readHyperslab"), "{e}");
|
||||||
|
}
|
||||||
|
assert_eq!(lazy.storage.stats().requests, before, "nothing fetched");
|
||||||
|
for part in [
|
||||||
|
local.read("/huge_u8", Some(&window)).unwrap(),
|
||||||
|
lazy.call(|r| r.read("/huge_u8", Some(&window))).unwrap(),
|
||||||
|
] {
|
||||||
|
assert_eq!(format!("{:?}", part.data), "U8([0, 0, 0, 7])");
|
||||||
|
}
|
||||||
|
|
||||||
|
// A server that claims 3 GiB and a heap collection of 2 GiB + 4 KiB:
|
||||||
|
// reading the strings fails at once, fetching a few blocks.
|
||||||
|
let data = std::fs::read(dir.path().join("hostile_vl.h5")).unwrap();
|
||||||
|
let storage = Arc::new(LazyStorage::new(3 << 30, LazyConfig::default()));
|
||||||
|
let s = storage.clone();
|
||||||
|
let reader = storage
|
||||||
|
.run_blocking(
|
||||||
|
|| Reader::open_storage(s.clone()),
|
||||||
|
|r| fetch_padded(&data, r),
|
||||||
|
)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap();
|
||||||
|
let e = storage
|
||||||
|
.run_blocking(|| reader.read("/a", None), |r| fetch_padded(&data, r))
|
||||||
|
.unwrap()
|
||||||
|
.unwrap_err();
|
||||||
|
assert!(e.contains("maxFetch"), "{e}");
|
||||||
|
let st = storage.stats();
|
||||||
|
assert!(st.requests <= 4 && st.bytes_fetched <= 4 << 20, "{st:?}");
|
||||||
|
}
|
||||||
|
|
||||||
|
/// make_fixture.py's files, written to a temporary directory.
|
||||||
|
fn fixture_dir() -> tempfile::TempDir {
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let generator = Path::new(env!("CARGO_MANIFEST_DIR"))
|
||||||
|
.join("../../examples/wasm-viewer/test/make_fixture.py");
|
||||||
|
let out = Command::new(python())
|
||||||
|
.arg(&generator)
|
||||||
|
.arg(dir.path())
|
||||||
|
.output()
|
||||||
|
.unwrap();
|
||||||
|
assert!(
|
||||||
|
out.status.success(),
|
||||||
|
"{}",
|
||||||
|
String::from_utf8_lossy(&out.stderr)
|
||||||
|
);
|
||||||
|
dir
|
||||||
|
}
|
||||||
|
|
||||||
|
fn hdf5_files(dir: &Path, out: &mut Vec<PathBuf>) {
|
||||||
|
let Ok(entries) = std::fs::read_dir(dir) else {
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
for e in entries.flatten() {
|
||||||
|
let p = e.path();
|
||||||
|
if p.is_dir() {
|
||||||
|
hdf5_files(&p, out);
|
||||||
|
} else if std::fs::read(&p)
|
||||||
|
.ok()
|
||||||
|
.is_some_and(|b| b.len() <= 64 << 20 && is_hdf5(&b))
|
||||||
|
{
|
||||||
|
out.push(p);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The HDF5 signature at 0 or a power-of-two user-block offset.
|
||||||
|
fn is_hdf5(b: &[u8]) -> bool {
|
||||||
|
const SIG: &[u8] = b"\x89HDF\r\n\x1a\n";
|
||||||
|
let mut at = 0usize;
|
||||||
|
loop {
|
||||||
|
if b.get(at..at + 8) == Some(SIG) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
at = if at == 0 { 512 } else { at * 2 };
|
||||||
|
if at >= b.len() {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn corpus_files_read_the_same_lazily() {
|
||||||
|
let Ok(dirs) = std::env::var("CLAWHDF5_WASM_CORPUS") else {
|
||||||
|
eprintln!("CLAWHDF5_WASM_CORPUS not set; skipping the corpus");
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
let mut files = Vec::new();
|
||||||
|
for d in std::env::split_paths(&dirs) {
|
||||||
|
hdf5_files(&d, &mut files);
|
||||||
|
}
|
||||||
|
files.sort();
|
||||||
|
assert!(!files.is_empty(), "no HDF5 files under {dirs}");
|
||||||
|
let (mut requests, mut bytes, mut total) = (0u64, 0u64, 0u64);
|
||||||
|
for f in &files {
|
||||||
|
let data = std::fs::read(f).unwrap();
|
||||||
|
total += data.len() as u64;
|
||||||
|
let (r, b, _) = check_equal(&f.display().to_string(), &data, 64 * 1024);
|
||||||
|
requests += r;
|
||||||
|
bytes += b;
|
||||||
|
}
|
||||||
|
eprintln!(
|
||||||
|
"{} files ({total} bytes): {requests} requests, {bytes} bytes fetched",
|
||||||
|
files.len()
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Passes and requests `list(path)` takes on a file opened lazily at
|
||||||
|
/// `block`-byte blocks (the open not counted), checking the listing against
|
||||||
|
/// the in-memory one.
|
||||||
|
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
|
||||||
|
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
|
||||||
|
let lazy = Lazy::open(
|
||||||
|
data.to_vec(),
|
||||||
|
LazyConfig {
|
||||||
|
block_size: block,
|
||||||
|
..LazyConfig::default()
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
let before = lazy.storage.stats();
|
||||||
|
assert_eq!(lazy.call(|r| r.list(path)).unwrap(), want);
|
||||||
|
let after = lazy.storage.stats();
|
||||||
|
(
|
||||||
|
after.passes - before.passes,
|
||||||
|
after.requests - before.requests,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Listing a group reads every child's object header, and its index (B-tree
|
||||||
|
/// and symbol table nodes, or B-tree v2 and heap blocks) before that. Each
|
||||||
|
/// pass asks for every node of a level it is missing, not the first one
|
||||||
|
/// only, so the passes (network round trips) grow with the depth of the
|
||||||
|
/// index, not with the number of children: 2000 children with headers
|
||||||
|
/// scattered over 512-byte blocks list in a handful of passes, where each
|
||||||
|
/// header block used to cost its own.
|
||||||
|
#[test]
|
||||||
|
fn listing_a_large_group_takes_a_few_passes() {
|
||||||
|
let mut b = FileBuilder::new();
|
||||||
|
let mut g = b.create_group("many");
|
||||||
|
for i in 0..600 {
|
||||||
|
g.create_dataset(&format!("d{i}")).with_i32_data(&[i; 64]);
|
||||||
|
}
|
||||||
|
b.add_group(g.finish());
|
||||||
|
let data = b.finish().unwrap();
|
||||||
|
let (passes, requests) = listing_cost(&data, "/many", 512);
|
||||||
|
eprintln!("FileBuilder, 600 children: {passes} passes, {requests} requests");
|
||||||
|
assert!(passes <= 6, "{passes} passes");
|
||||||
|
|
||||||
|
if !python_available() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
for libver in ["earliest", "latest"] {
|
||||||
|
let path = dir.path().join(format!("{libver}.h5"));
|
||||||
|
let script = format!(
|
||||||
|
"import h5py, numpy as np\n\
|
||||||
|
with h5py.File({:?}, 'w', libver='{libver}') as f:\n\
|
||||||
|
\x20 for i in range(2000):\n\
|
||||||
|
\x20 f.create_dataset('d%d' % i, data=np.full(256, i, np.float32))\n",
|
||||||
|
path.display().to_string()
|
||||||
|
);
|
||||||
|
let out = Command::new(python())
|
||||||
|
.args(["-c", &script])
|
||||||
|
.output()
|
||||||
|
.unwrap();
|
||||||
|
assert!(
|
||||||
|
out.status.success(),
|
||||||
|
"{}",
|
||||||
|
String::from_utf8_lossy(&out.stderr)
|
||||||
|
);
|
||||||
|
let data = std::fs::read(&path).unwrap();
|
||||||
|
let (passes, requests) = listing_cost(&data, "/", 512);
|
||||||
|
eprintln!("h5py libver={libver}, 2000 children: {passes} passes, {requests} requests");
|
||||||
|
assert!(passes <= 12, "{libver}: {passes} passes");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `CLAWHDF5_WASM_LIST_FILE=file.h5`: what listing the root group of that
|
||||||
|
/// file costs lazily, at 1 MiB and 64 KiB blocks (a measurement, printed).
|
||||||
|
#[test]
|
||||||
|
fn listing_cost_of_a_given_file() {
|
||||||
|
let Ok(path) = std::env::var("CLAWHDF5_WASM_LIST_FILE") else {
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
let data = std::fs::read(&path).unwrap();
|
||||||
|
for block in [1 << 20, 64 << 10] {
|
||||||
|
let lazy = Lazy::open(
|
||||||
|
data.clone(),
|
||||||
|
LazyConfig {
|
||||||
|
block_size: block,
|
||||||
|
..LazyConfig::default()
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
let open = lazy.storage.stats();
|
||||||
|
let n = lazy.call(|r| r.list("/")).unwrap().len();
|
||||||
|
let st = lazy.storage.stats();
|
||||||
|
eprintln!(
|
||||||
|
"{path} ({} bytes), {block}-byte blocks: open {} requests / {} passes; list('/') of {n}: {} passes, {} requests, {} bytes",
|
||||||
|
data.len(),
|
||||||
|
open.requests,
|
||||||
|
open.passes,
|
||||||
|
st.passes - open.passes,
|
||||||
|
st.requests - open.requests,
|
||||||
|
st.bytes_fetched - open.bytes_fetched
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -20,13 +20,20 @@ change. Progress: M1, first part (the `Storage` trait and the metadata
|
|||||||
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
||||||
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
||||||
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
|
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||||
|
|
||||||
|
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
|
||||||
|
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
|
||||||
|
reader, through the restartable `NeedBytes` mode (see the M4 status
|
||||||
|
below). M5 (SWMR) is not started.
|
||||||
|
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||||
converted code without changing the plan: object-header continuation chunks
|
converted code without changing the plan: object-header continuation chunks
|
||||||
are followed without recursion (still one bounded `read_at` per chunk), and
|
are followed without recursion (still one bounded `read_at` per chunk), and
|
||||||
implicit chunk indexes are addressed over the maximum chunk grid (in
|
implicit chunk indexes are addressed over the maximum chunk grid (in
|
||||||
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
|
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
|
||||||
working on the whole file in memory; it is not part of this design. Every count below was
|
working on the whole file in memory; it is not part of this design. Every
|
||||||
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
|
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
|
||||||
next to it. No timing numbers appear here on purpose: the machine was shared
|
every count in a milestone's status on the date it gives, with the
|
||||||
|
commands given next to it. No timing numbers appear here on purpose: the machine was shared
|
||||||
with other build jobs when this was written.
|
with other build jobs when this was written.
|
||||||
|
|
||||||
## The problem
|
## The problem
|
||||||
@@ -536,6 +543,70 @@ fast path within benchmark noise.
|
|||||||
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
|
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
|
||||||
to a whole download when the server does not answer 206.
|
to a whole download when the server does not answer 206.
|
||||||
- `examples/wasm-viewer`: open by URL.
|
- `examples/wasm-viewer`: open by URL.
|
||||||
|
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
|
||||||
|
with these choices:
|
||||||
|
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
|
||||||
|
`Storage` over the blocks fetched so far. A call (open, list, read)
|
||||||
|
runs as a pass; a read that misses records its blocks and fails with
|
||||||
|
a storage error. The error itself is not the signal: parsers catch
|
||||||
|
errors and carry on (a listing leaves out a link it cannot resolve),
|
||||||
|
so *any* pass that recorded a miss is thrown away, whatever it
|
||||||
|
returned, and re-run once the missing ranges have been fetched
|
||||||
|
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
|
||||||
|
synchronous XHR would have kept the parsers' single pass, but it
|
||||||
|
needs the page to run the reader off the main thread (a second
|
||||||
|
module, messages for every call and every typed array copied back),
|
||||||
|
since synchronous XHR can return binary data only there; the
|
||||||
|
restartable loop costs a
|
||||||
|
re-parse per wave of misses instead, which is CPU, not network.
|
||||||
|
- **Progress:** no block is evicted while a call is in flight
|
||||||
|
(`LazyStorage::operation`), so each pass that does not finish asks for
|
||||||
|
at least one new block, and a call ends after at most one pass per
|
||||||
|
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
|
||||||
|
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
|
||||||
|
evicted before metadata. The price: a call holds everything it reads
|
||||||
|
until it finishes.
|
||||||
|
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
|
||||||
|
backend by blocking, evicts during a read, and does not keep reads
|
||||||
|
that miss more than half its budget; a restartable pass needs every
|
||||||
|
block it has read to be there when it is re-run. The block
|
||||||
|
arithmetic and coalescing (runs of consecutive blocks, a one-block
|
||||||
|
hole filled, at most 8 MiB per request) follow it; the cache is
|
||||||
|
about 450 lines (with its documentation) in the wasm crate, with no
|
||||||
|
in-flight tracking or HTTP. (A hole already cached is not filled:
|
||||||
|
it would be fetched again.)
|
||||||
|
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
|
||||||
|
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
|
||||||
|
requests at a time, every answer checked (206, `Content-Range` when
|
||||||
|
visible, body length, the ETag/Last-Modified and length of the
|
||||||
|
first answer). A `200` to the first request is kept as the whole
|
||||||
|
file up to `maxDownload` (512 MiB), or refused with
|
||||||
|
`fallback: "error"`. The first block comes with the request that
|
||||||
|
learns the length, as in M3.
|
||||||
|
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
|
||||||
|
--test lazy` compares what the viewer shows of each file (kinds,
|
||||||
|
listings, attributes, info, whole reads, hyperslabs) read lazily at
|
||||||
|
512 B to 1 MiB blocks with the in-memory reader and the facade's
|
||||||
|
range-storage path, for built files, the h5py/netCDF4 fixture and,
|
||||||
|
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
|
||||||
|
files. The built package, under Node and headless Chromium
|
||||||
|
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
|
||||||
|
`test/serve.py`, which counts requests): every fixture check over
|
||||||
|
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
|
||||||
|
attributes and a 10-value window of the 25-million-value dataset
|
||||||
|
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
|
||||||
|
with `open(bytes)`.
|
||||||
|
- After review (2026-09-27): a call may fetch at most `maxFetch`
|
||||||
|
(512 MiB, at most 1 GiB) and a longer single read is refused before
|
||||||
|
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
|
||||||
|
more is refused at open on wasm32 (offsets become `usize`); every
|
||||||
|
response body is cut off at the length asked for. Parsers that walk
|
||||||
|
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
|
||||||
|
the wasm listing's child headers) read the siblings after the first
|
||||||
|
failure before returning it, so a pass asks for a whole level's
|
||||||
|
missing blocks: listing 3000 datasets went from 185 passes to 6.
|
||||||
|
The 32-bit risk below is covered by a Node test that reads data at
|
||||||
|
3 GiB from a mock server and is refused a 4 GiB file.
|
||||||
|
|
||||||
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
||||||
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
||||||
|
|||||||
+68
-3
@@ -948,6 +948,11 @@ cache, but:
|
|||||||
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
|
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
|
||||||
(`h5rs` and the Python bindings read through `File::storage`, and take
|
(`h5rs` and the Python bindings read through `File::storage`, and take
|
||||||
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
|
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
|
||||||
|
|
||||||
|
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
|
||||||
|
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
|
||||||
|
feature; the wasm reader's `openUrl` reads by range requests since
|
||||||
|
2026-09-27, its `open(bytes)` takes a whole file).
|
||||||
- The file's length is read once, at open: a growing file (SWMR) is not
|
- The file's length is read once, at open: a growing file (SWMR) is not
|
||||||
followed (milestone M5). A remote file is pinned at open, so one that
|
followed (milestone M5). A remote file is pinned at open, so one that
|
||||||
grows is `RemoteError::FileChanged`.
|
grows is `RemoteError::FileChanged`.
|
||||||
@@ -969,6 +974,10 @@ cache, but:
|
|||||||
built with `--features https` (rustls with ring, which compiles C), and
|
built with `--features https` (rustls with ring, which compiles C), and
|
||||||
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
|
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
|
||||||
The Python tests run against an in-process `http.server` only.
|
The Python tests run against an in-process `http.server` only.
|
||||||
|
|
||||||
|
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
|
||||||
|
`File::as_bytes`, which a remote file does not have. (The browser can
|
||||||
|
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
|
||||||
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
||||||
The design's policy of using a paged file's page size as the block size
|
The design's policy of using a paged file's page size as the block size
|
||||||
is not implemented, and only the first block is read ahead.
|
is not implemented, and only the first block is read ahead.
|
||||||
@@ -1007,10 +1016,66 @@ cache, but:
|
|||||||
|
|
||||||
## `clawhdf5-wasm` (browser) limits
|
## `clawhdf5-wasm` (browser) limits
|
||||||
|
|
||||||
**Status:** open (by design for now; added 2026-09-26).
|
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
|
||||||
|
|
||||||
- The whole file is held in memory: `open()` takes its bytes. There are no
|
- `open()` holds the whole file in memory (it takes its bytes), so a
|
||||||
HTTP range reads, so a multi-GB file does not fit a browser tab.
|
multi-GB local file does not fit a browser tab. A file on a web server
|
||||||
|
can be opened with `openUrl()` instead, which fetches only the byte
|
||||||
|
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
|
||||||
|
with these limits:
|
||||||
|
- **Round trips:** a call runs as passes over the blocks fetched so far
|
||||||
|
and is re-run after each wave of misses, so a call costs one round
|
||||||
|
trip per wave, not one for everything: a chunk index is walked a
|
||||||
|
level (or a node) per round trip, while the chunks of a read are
|
||||||
|
fetched together. Listing a group asks for every child's object
|
||||||
|
header, and every node of a level of the group's index, in one pass
|
||||||
|
(since 2026-09-27; it was one round trip per header block): 3000
|
||||||
|
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
|
||||||
|
`libver="latest"` file (dense links). Each pass re-parses what the
|
||||||
|
call reads (CPU, not network). With headers spread through the file
|
||||||
|
(h5py writes each next to its data) a listing still fetches most of
|
||||||
|
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
|
||||||
|
- **Memory:** a call keeps every block it reads until it finishes (the
|
||||||
|
cache budget applies between calls). It may fetch at most `maxFetch`
|
||||||
|
bytes (512 MiB by default, at most 1 GiB), and a single read longer
|
||||||
|
than that fails before anything is fetched: a hostile server cannot
|
||||||
|
make the page fetch or allocate what a file's lengths claim. `read()`
|
||||||
|
of a dataset that would take more than 1 GiB while it is decoded
|
||||||
|
(stored bytes, the values widened to 64 bits, the result) fails,
|
||||||
|
naming `readHyperslab`; read windows of large datasets. On wasm32 a
|
||||||
|
buffer past 2 GiB cannot exist and a failed allocation aborts the
|
||||||
|
whole module (every open file on the page), which these limits keep
|
||||||
|
from happening; before 2026-09-27 both did abort it.
|
||||||
|
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
|
||||||
|
open. The format code turns file offsets into `usize` to use them
|
||||||
|
(with a clean error past it), which is 32 bits on wasm32, so nothing
|
||||||
|
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
|
||||||
|
tested (with a mock server); files above 200 MB have not been served
|
||||||
|
for real.
|
||||||
|
- **Cross-origin servers** must allow CORS for the page's origin and
|
||||||
|
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
|
||||||
|
answer `HEAD` with `Content-Length`. The file is pinned at open by its
|
||||||
|
`ETag` or `Last-Modified` (and its length): when the page cannot see
|
||||||
|
either header, only a change of length is detected.
|
||||||
|
- **A server without range support** (it answers `200`) costs a whole
|
||||||
|
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
|
||||||
|
with `fallback: "error"`. Every body, this one and each `206`, is read
|
||||||
|
as it arrives and cut off past its limit (the range asked for, or
|
||||||
|
`maxDownload`): a server cannot make the page buffer more.
|
||||||
|
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
|
||||||
|
size is not used. No retries: a failed request fails the call (calling
|
||||||
|
again retries it; what was fetched stays cached).
|
||||||
|
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
||||||
|
a local server, cross-origin included (a page on 127.0.0.1 reading a
|
||||||
|
file from localhost, with and without exposed headers); not in Firefox
|
||||||
|
or Safari.
|
||||||
|
- The native corpus comparison (`tests/lazy.rs` with
|
||||||
|
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
|
||||||
|
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
|
||||||
|
and which chunk's error is reported depends on the iteration order of
|
||||||
|
the chunk index (a `HashMap`, seeded per process), so the lazy and
|
||||||
|
the range-storage reads can name different errors. Both are errors;
|
||||||
|
not specific to `openUrl` (it predates it).
|
||||||
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
||||||
refused with an error naming the type; attributes of those types come back
|
refused with an error naming the type; attributes of those types come back
|
||||||
as `value: null` with their `dtype`.
|
as `value: null` with their `dtype`.
|
||||||
|
|||||||
+100
-20
@@ -2,10 +2,12 @@
|
|||||||
|
|
||||||
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
|
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
|
||||||
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
|
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
|
||||||
file, browse its groups, and look at a dataset's type, shape, attributes
|
file or give a URL, browse its groups, and look at a dataset's type, shape,
|
||||||
and values (a 50 x 12 window at a time, read as a hyperslab, with the
|
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
|
||||||
leading dimensions of a 3-D+ dataset held at chosen indices). The file never
|
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
|
||||||
leaves the page.
|
local file never leaves the page. A file given by URL is not downloaded:
|
||||||
|
only the byte ranges each view needs are fetched (HTTP range requests),
|
||||||
|
and the header shows the requests and bytes that has cost so far.
|
||||||
|
|
||||||
## Build and open
|
## Build and open
|
||||||
|
|
||||||
@@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe
|
|||||||
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
|
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
|
||||||
```
|
```
|
||||||
|
|
||||||
Then open <http://localhost:8000/>. `?file=<url>&path=<object>` opens a
|
Then open <http://localhost:8000/>. The URL box, or
|
||||||
file from a URL (same origin, or one serving CORS headers) and selects an
|
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
|
||||||
object in it, e.g. `?file=data/run1.h5&path=/results/energy`.
|
sending CORS headers, see Limits) and selects an object in it, e.g.
|
||||||
|
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
|
||||||
|
answer range requests, so a file served by it is downloaded whole (the
|
||||||
|
header says so); `test/serve.py` does:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
|
||||||
|
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
|
||||||
|
```
|
||||||
|
|
||||||
## JavaScript API
|
## JavaScript API
|
||||||
|
|
||||||
```js
|
```js
|
||||||
import init, { open } from "./pkg/clawhdf5_wasm.js";
|
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
|
||||||
await init();
|
await init();
|
||||||
const f = open(new Uint8Array(await blob.arrayBuffer()));
|
const f = open(new Uint8Array(await blob.arrayBuffer()));
|
||||||
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
|
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
|
||||||
@@ -32,8 +42,41 @@ f.attrs("/grid"); // [{ name, value, dtype }]
|
|||||||
f.read("/grid"); // { shape, dtype, data }
|
f.read("/grid"); // { shape, dtype, data }
|
||||||
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
|
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
|
||||||
f.free();
|
f.free();
|
||||||
|
|
||||||
|
// A file on a web server, read by HTTP range requests as needed. The same
|
||||||
|
// methods, each returning a promise.
|
||||||
|
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
|
||||||
|
await r.list("/");
|
||||||
|
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
|
||||||
|
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
|
||||||
|
r.free();
|
||||||
```
|
```
|
||||||
|
|
||||||
|
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
|
||||||
|
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
|
||||||
|
between calls, default 64 MiB), `maxFetch` (bytes one call may fetch, and
|
||||||
|
so the longest single read, up to 1 GiB, default 512 MiB), `fallback`
|
||||||
|
(`"download"`, the default, reads the whole file when the server ignores
|
||||||
|
`Range`, up to `maxDownload` bytes, default 512 MiB, at most 1 GiB;
|
||||||
|
`"error"` refuses such a server), `headers` (a `Headers`, `[name, value]`
|
||||||
|
pairs or an object) and `credentials` (passed to every request),
|
||||||
|
`parallel` (requests in flight, 1 to 1024, default 6; when one fails the
|
||||||
|
others are aborted), `fetch` (a `fetch`-compatible function to use).
|
||||||
|
|
||||||
|
How it works: the reader is synchronous and a page cannot block on the
|
||||||
|
network, so each call runs as a *pass* over the blocks fetched so far. A
|
||||||
|
pass that needs a block not yet fetched is abandoned, the missing blocks
|
||||||
|
are fetched (in parallel, adjacent blocks in one request), and the pass is
|
||||||
|
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
||||||
|
costs one request (the first block, which also gives the file's size);
|
||||||
|
listing a group whose metadata is in blocks already fetched costs none,
|
||||||
|
and otherwise a round trip per level of the group's index plus one for
|
||||||
|
its children's headers, all fetched together;
|
||||||
|
reading a chunked dataset costs a round trip for its chunk index (a few
|
||||||
|
for a deep one) and one batch of requests for its chunks. Every answer is
|
||||||
|
checked — a `206` with exactly the bytes asked for, from the same file
|
||||||
|
(ETag or Last-Modified, and length) — or the call fails.
|
||||||
|
|
||||||
`data` is the typed array of the stored width (`Float64Array`,
|
`data` is the typed array of the stored width (`Float64Array`,
|
||||||
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
|
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
|
||||||
`BigUint64Array`), or an array of strings for fixed- and variable-length
|
`BigUint64Array`), or an array of strings for fixed- and variable-length
|
||||||
@@ -43,7 +86,17 @@ else throws an `Error` naming the type.
|
|||||||
|
|
||||||
## Limits
|
## Limits
|
||||||
|
|
||||||
- Read-only, and the whole file is held in memory (no range requests).
|
- Read-only. `open(bytes)` holds the whole file in memory.
|
||||||
|
- `openUrl`: a cross-origin server must allow CORS and expose
|
||||||
|
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
|
||||||
|
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
|
||||||
|
`Last-Modified` either, a file replaced on the server is detected only by
|
||||||
|
a change of length. A call keeps what it reads until it finishes, and
|
||||||
|
may fetch at most `maxFetch`; `read()` of a dataset that would take more
|
||||||
|
than 1 GiB to decode is refused (read windows of large datasets with
|
||||||
|
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
|
||||||
|
Every response body is cut off past the length asked for. More in
|
||||||
|
`docs/known-issues.md`.
|
||||||
- Compound, reference, opaque and variable-length-sequence datasets are
|
- Compound, reference, opaque and variable-length-sequence datasets are
|
||||||
refused with an error. Attributes of those types are listed with
|
refused with an error. Attributes of those types are listed with
|
||||||
`value: null` and their `dtype`.
|
`value: null` and their `dtype`.
|
||||||
@@ -56,28 +109,55 @@ else throws an `Error` naming the type.
|
|||||||
## Tests
|
## Tests
|
||||||
|
|
||||||
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
|
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
|
||||||
`fixture.nc` (netCDF4) with `test/make_fixture.py`, then:
|
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
|
||||||
|
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
|
||||||
|
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
|
||||||
|
counter, `/norange/...` for a server without range support,
|
||||||
|
`/noexpose/...` and `/unexposed/...` for one that does not expose its
|
||||||
|
headers to CORS), then:
|
||||||
|
|
||||||
- runs `test/test.mjs` under Node: every dataset (whole and a strided
|
- runs `test/test.mjs` under Node: every dataset (whole and a strided
|
||||||
hyperslab), listing and attribute is compared with what libhdf5 reads
|
hyperslab), listing and attribute is compared with what libhdf5 reads
|
||||||
back, error paths are checked, and so are the page's DOM-free helpers
|
back, error paths are checked, and so are the page's DOM-free helpers
|
||||||
(`viewer-lib.js`);
|
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
|
||||||
|
and 512 B blocks), calls in flight at once, the request budget of
|
||||||
|
`big.h5` (listing it and reading small things of it must take at most 8
|
||||||
|
requests and under 5% of the file), the download fallback and its
|
||||||
|
limit, and the errors: HTTP status, a file that changes, a server that
|
||||||
|
sends the wrong bytes or stops honouring `Range`. Also the cross-origin
|
||||||
|
path with no exposed headers (`/noexpose/`: length from `HEAD`), the
|
||||||
|
size limits on `make_fixture.py`'s `limits.h5`, `hostile_vl.h5` (a heap
|
||||||
|
collection claiming 2 GiB) and `far.h5` (data at 3 GiB, served by a
|
||||||
|
mock), bodies longer than asked for, `headers` forms, `parallel`, and
|
||||||
|
sibling requests aborted after a failure.
|
||||||
|
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
|
||||||
|
to 16 MiB) read by URL with the same file read from bytes;
|
||||||
- runs `test/browser.sh`: loads the page in headless Chromium with
|
- runs `test/browser.sh`: loads the page in headless Chromium with
|
||||||
`?file=fixture.h5&path=...` for eight objects and checks the rendered tree,
|
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
|
||||||
types, shapes, attribute and value cells, and the error shown for an
|
tree, types, shapes, attribute and value cells, the request counter, and
|
||||||
unsupported type. Skipped when no Chromium is found (`CHROME` names one;
|
the error shown for an unsupported type; then the file from another
|
||||||
a Playwright download under `~/.cache/ms-playwright` is picked up).
|
origin (localhost), with CORS exposing `Content-Range` and exposing
|
||||||
Drag-and-drop and the file picker are not driven by it; they share
|
nothing (`/unexposed/`, where the server must see a `HEAD`), a server
|
||||||
`load()` with the `?file=` path.
|
without range support, and `big.h5` (a small dataset and a window of the large one,
|
||||||
|
with a single-digit percentage of the file fetched). Skipped when no
|
||||||
|
Chromium is found (`CHROME` names one; a Playwright download under
|
||||||
|
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
|
||||||
|
and the URL box are not driven by it; they share `setFile()` with the
|
||||||
|
`?file=` path.
|
||||||
|
|
||||||
The same expectations are checked natively, without Node, by
|
The same expectations are checked natively, without Node, by
|
||||||
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI
|
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
|
||||||
container has no Node or browser).
|
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
|
||||||
|
corpus too), which is what CI runs (the CI container has no Node or
|
||||||
|
browser).
|
||||||
|
|
||||||
## Size
|
## Size
|
||||||
|
|
||||||
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
|
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
|
||||||
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`:
|
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
|
||||||
|
larger now and the table has not been re-measured: the reader has grown
|
||||||
|
since, and `openUrl` (2026-09-27) made the facade's range-read path
|
||||||
|
reachable from JavaScript and added promise glue and `remote.js`.
|
||||||
|
|
||||||
| | raw | gzip -9 |
|
| | raw | gzip -9 |
|
||||||
|---|---:|---:|
|
|---|---:|---:|
|
||||||
|
|||||||
@@ -56,27 +56,39 @@
|
|||||||
border: 1px solid var(--line); border-radius: 4px; }
|
border: 1px solid var(--line); border-radius: 4px; }
|
||||||
.error { color: var(--bad); font-family: var(--mono); white-space: pre-wrap; }
|
.error { color: var(--bad); font-family: var(--mono); white-space: pre-wrap; }
|
||||||
.muted { color: var(--muted); }
|
.muted { color: var(--muted); }
|
||||||
|
form.url { display: flex; gap: 6px; margin-left: auto; flex: 1 1 320px; max-width: 560px; }
|
||||||
|
form.url input { flex: 1; min-width: 0; font: 13px var(--mono); padding: 4px 8px; background: var(--panel);
|
||||||
|
color: var(--ink); border: 1px solid var(--line); border-radius: 6px; }
|
||||||
|
.netstats { flex-basis: 100%; color: var(--muted); font-size: 12.5px; font-family: var(--mono); }
|
||||||
|
.netstats:empty { display: none; }
|
||||||
</style>
|
</style>
|
||||||
</head>
|
</head>
|
||||||
<body>
|
<body>
|
||||||
<header>
|
<header>
|
||||||
<h1>HDF5 Viewer</h1>
|
<h1>HDF5 Viewer</h1>
|
||||||
<span class="file" id="filename">no file</span>
|
<span class="file" id="filename">no file</span>
|
||||||
|
<form class="url" id="urlform">
|
||||||
|
<input type="url" id="url" placeholder="https://…/file.h5 (read by HTTP range requests)" aria-label="File URL">
|
||||||
|
<button class="button" type="submit">Open URL</button>
|
||||||
|
</form>
|
||||||
<label class="button">Open file…<input type="file" id="picker" accept=".h5,.hdf5,.he5,.nc,.nc4,.cdf" hidden></label>
|
<label class="button">Open file…<input type="file" id="picker" accept=".h5,.hdf5,.he5,.nc,.nc4,.cdf" hidden></label>
|
||||||
|
<span class="netstats" id="netstats"></span>
|
||||||
</header>
|
</header>
|
||||||
<main>
|
<main>
|
||||||
<nav id="tree"></nav>
|
<nav id="tree"></nav>
|
||||||
<section id="detail">
|
<section id="detail">
|
||||||
<div class="drop">
|
<div class="drop">
|
||||||
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, or use “Open file…”.</p>
|
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, use “Open file…”, or give a URL.</p>
|
||||||
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; it is not uploaded anywhere.</p>
|
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; a local file is not uploaded anywhere.
|
||||||
|
A file given by URL is not downloaded: only the byte ranges each view needs are fetched (HTTP range requests),
|
||||||
|
and the requests and bytes it has cost are shown at the top.</p>
|
||||||
<p class="muted" id="version"></p>
|
<p class="muted" id="version"></p>
|
||||||
</div>
|
</div>
|
||||||
</section>
|
</section>
|
||||||
</main>
|
</main>
|
||||||
<script type="module">
|
<script type="module">
|
||||||
import init, { open, version } from "./pkg/clawhdf5_wasm.js";
|
import init, { open, openUrl, version } from "./pkg/clawhdf5_wasm.js";
|
||||||
import { joinPath, formatValue, viewWindow, toRows, perElement } from "./viewer-lib.js";
|
import { joinPath, formatValue, viewWindow, toRows, perElement, formatStats } from "./viewer-lib.js";
|
||||||
|
|
||||||
const $ = (id) => document.getElementById(id);
|
const $ = (id) => document.getElementById(id);
|
||||||
const el = (tag, props = {}, ...kids) => {
|
const el = (tag, props = {}, ...kids) => {
|
||||||
@@ -85,30 +97,58 @@ const el = (tag, props = {}, ...kids) => {
|
|||||||
return e;
|
return e;
|
||||||
};
|
};
|
||||||
|
|
||||||
|
// The open file: an H5File (a local file, read in memory; synchronous
|
||||||
|
// methods) or a RemoteFile (openUrl: bytes fetched by range requests as
|
||||||
|
// needed; methods return promises). Every call below is awaited, so both
|
||||||
|
// work the same.
|
||||||
let file = null;
|
let file = null;
|
||||||
let selectedNode = null;
|
let selectedNode = null;
|
||||||
|
// Bumped by every selection, so a slow read for an object no longer shown
|
||||||
|
// does not overwrite the current one.
|
||||||
|
let showSeq = 0;
|
||||||
|
|
||||||
await init();
|
await init();
|
||||||
$("version").textContent = `clawhdf5 ${version()}`;
|
$("version").textContent = `clawhdf5 ${version()}`;
|
||||||
|
|
||||||
async function load(blob) {
|
// Requests and bytes a remote file has cost so far.
|
||||||
const bytes = new Uint8Array(await blob.arrayBuffer());
|
function updateStats() {
|
||||||
|
$("netstats").textContent = file && file.stats ? formatStats(file.stats()) : "";
|
||||||
|
}
|
||||||
|
|
||||||
|
async function setFile(name, opener) {
|
||||||
if (file) file.free();
|
if (file) file.free();
|
||||||
file = null;
|
file = null;
|
||||||
$("filename").textContent = blob.name;
|
selectedNode = null;
|
||||||
|
showSeq++;
|
||||||
|
$("filename").textContent = name;
|
||||||
$("tree").replaceChildren();
|
$("tree").replaceChildren();
|
||||||
|
$("netstats").textContent = "";
|
||||||
|
$("detail").replaceChildren(el("p", { className: "muted", textContent: `Opening ${name}…` }));
|
||||||
try {
|
try {
|
||||||
file = open(bytes);
|
file = await opener();
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${blob.name}: ${e.message}` }));
|
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${name}: ${e.message}` }));
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
updateStats();
|
||||||
const root = el("ul", { className: "tree" });
|
const root = el("ul", { className: "tree" });
|
||||||
root.append(treeNode("/", "/", "group"));
|
root.append(treeNode("/", "/", "group"));
|
||||||
$("tree").append(root);
|
$("tree").append(root);
|
||||||
root.querySelector(".node").click();
|
await activate(root.querySelector(".node"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
async function load(blob) {
|
||||||
|
const bytes = new Uint8Array(await blob.arrayBuffer());
|
||||||
|
await setFile(blob.name, async () => open(bytes));
|
||||||
|
}
|
||||||
|
|
||||||
|
async function loadUrl(url) {
|
||||||
|
$("url").value = url;
|
||||||
|
await setFile(url, () => openUrl(new URL(url, location.href).href));
|
||||||
|
}
|
||||||
|
|
||||||
|
// A tree row: clicking it selects it (and expands a group); `activate`
|
||||||
|
// does the same and resolves once the listing and the detail are shown.
|
||||||
function treeNode(path, name, kind) {
|
function treeNode(path, name, kind) {
|
||||||
const li = el("li");
|
const li = el("li");
|
||||||
const icon = el("span", { className: "icon", textContent: kind === "group" ? "▸" : "·" });
|
const icon = el("span", { className: "icon", textContent: kind === "group" ? "▸" : "·" });
|
||||||
@@ -116,34 +156,39 @@ function treeNode(path, name, kind) {
|
|||||||
row.dataset.path = path;
|
row.dataset.path = path;
|
||||||
li.append(row);
|
li.append(row);
|
||||||
let children = null;
|
let children = null;
|
||||||
row.addEventListener("click", () => {
|
row.activate = async () => {
|
||||||
if (selectedNode) selectedNode.classList.remove("selected");
|
if (selectedNode) selectedNode.classList.remove("selected");
|
||||||
row.classList.add("selected");
|
row.classList.add("selected");
|
||||||
selectedNode = row;
|
selectedNode = row;
|
||||||
|
const shown = show(path, kind);
|
||||||
if (kind === "group") {
|
if (kind === "group") {
|
||||||
if (children) {
|
if (children) {
|
||||||
children.hidden = !children.hidden;
|
children.hidden = !children.hidden;
|
||||||
} else {
|
} else {
|
||||||
children = el("ul", { className: "tree" });
|
children = el("ul", { className: "tree" });
|
||||||
|
li.append(children);
|
||||||
try {
|
try {
|
||||||
for (const c of file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
|
for (const c of await file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
children.append(el("li", { className: "error", textContent: e.message }));
|
children.append(el("li", { className: "error", textContent: e.message }));
|
||||||
}
|
}
|
||||||
li.append(children);
|
updateStats();
|
||||||
}
|
}
|
||||||
icon.textContent = children.hidden ? "▸" : "▾";
|
icon.textContent = children.hidden ? "▸" : "▾";
|
||||||
}
|
}
|
||||||
show(path, kind);
|
await shown;
|
||||||
});
|
};
|
||||||
|
row.addEventListener("click", () => row.activate());
|
||||||
return li;
|
return li;
|
||||||
}
|
}
|
||||||
|
|
||||||
function attrsTable(path) {
|
const activate = (row) => row.activate();
|
||||||
|
|
||||||
|
async function attrsTable(path) {
|
||||||
let attrs, errors;
|
let attrs, errors;
|
||||||
try {
|
try {
|
||||||
attrs = file.attrs(path);
|
attrs = await file.attrs(path);
|
||||||
errors = file.attrErrors(path);
|
errors = await file.attrErrors(path);
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
return el("p", { className: "error", textContent: e.message });
|
return el("p", { className: "error", textContent: e.message });
|
||||||
}
|
}
|
||||||
@@ -157,14 +202,16 @@ function attrsTable(path) {
|
|||||||
return el("div", { className: "scroll" }, t);
|
return el("div", { className: "scroll" }, t);
|
||||||
}
|
}
|
||||||
|
|
||||||
function show(path, kind) {
|
async function show(path, kind) {
|
||||||
|
const seq = ++showSeq;
|
||||||
|
const current = () => seq === showSeq;
|
||||||
const out = [el("h2", { textContent: path })];
|
const out = [el("h2", { textContent: path })];
|
||||||
if (kind === "dataset") {
|
if (kind === "dataset") {
|
||||||
let info;
|
let info;
|
||||||
try {
|
try {
|
||||||
info = file.info(path);
|
info = await file.info(path);
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
$("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
|
if (current()) $("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
const max = info.maxshape === null ? "—" : `(${info.maxshape.map((d) => d ?? "∞").join(", ")})`;
|
const max = info.maxshape === null ? "—" : `(${info.maxshape.map((d) => d ?? "∞").join(", ")})`;
|
||||||
@@ -172,15 +219,16 @@ function show(path, kind) {
|
|||||||
el("dt", { textContent: "type" }), el("dd", { textContent: info.dtype }),
|
el("dt", { textContent: "type" }), el("dd", { textContent: info.dtype }),
|
||||||
el("dt", { textContent: "shape" }), el("dd", { textContent: `(${info.shape.join(", ")})` }),
|
el("dt", { textContent: "shape" }), el("dd", { textContent: `(${info.shape.join(", ")})` }),
|
||||||
el("dt", { textContent: "max shape" }), el("dd", { textContent: max })));
|
el("dt", { textContent: "max shape" }), el("dd", { textContent: max })));
|
||||||
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
|
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
|
||||||
out.push(el("h3", { textContent: "Values" }), valuesView(path, info));
|
out.push(el("h3", { textContent: "Values" }), await valuesView(path, info, current));
|
||||||
} else {
|
} else {
|
||||||
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
|
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
|
||||||
}
|
}
|
||||||
$("detail").replaceChildren(...out);
|
updateStats();
|
||||||
|
if (current()) $("detail").replaceChildren(...out);
|
||||||
}
|
}
|
||||||
|
|
||||||
function valuesView(path, info) {
|
async function valuesView(path, info, current) {
|
||||||
const shape = info.shape;
|
const shape = info.shape;
|
||||||
const per = perElement(info.elementShape);
|
const per = perElement(info.elementShape);
|
||||||
const state = { row: 0, col: 0, rows: 50, cols: 12, fixed: shape.slice(0, Math.max(0, shape.length - 2)).map(() => 0) };
|
const state = { row: 0, col: 0, rows: 50, cols: 12, fixed: shape.slice(0, Math.max(0, shape.length - 2)).map(() => 0) };
|
||||||
@@ -202,15 +250,19 @@ function valuesView(path, info) {
|
|||||||
if (shape.length >= 2) controls.append(num("column", "col"));
|
if (shape.length >= 2) controls.append(num("column", "col"));
|
||||||
state.fixed.forEach((_, i) => controls.append(num(`dim ${i}`, "fixed", i)));
|
state.fixed.forEach((_, i) => controls.append(num(`dim ${i}`, "fixed", i)));
|
||||||
|
|
||||||
function render() {
|
let renderSeq = 0;
|
||||||
|
async function render() {
|
||||||
|
const seq = ++renderSeq;
|
||||||
const w = viewWindow(shape, state);
|
const w = viewWindow(shape, state);
|
||||||
let res;
|
let res;
|
||||||
try {
|
try {
|
||||||
res = w === null ? file.read(path) : file.readHyperslab(path, w.start, w.count);
|
res = w === null ? await file.read(path) : await file.readHyperslab(path, w.start, w.count);
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
body.replaceChildren(el("p", { className: "error", textContent: e.message }));
|
if (seq === renderSeq) body.replaceChildren(el("p", { className: "error", textContent: e.message }));
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
updateStats();
|
||||||
|
if (seq !== renderSeq || !current()) return;
|
||||||
if (w === null) {
|
if (w === null) {
|
||||||
body.replaceChildren(el("pre", { textContent: toRows(res.data, 1, 1, per)[0][0] }));
|
body.replaceChildren(el("pre", { textContent: toRows(res.data, 1, 1, per)[0][0] }));
|
||||||
return;
|
return;
|
||||||
@@ -229,41 +281,39 @@ function valuesView(path, info) {
|
|||||||
(shape.length >= 2 ? `, columns ${w.col}–${w.col + w.cols - 1}` : "") + ` of ${total} values` });
|
(shape.length >= 2 ? `, columns ${w.col}–${w.col + w.cols - 1}` : "") + ` of ${total} values` });
|
||||||
body.replaceChildren(t, note);
|
body.replaceChildren(t, note);
|
||||||
}
|
}
|
||||||
render();
|
await render();
|
||||||
return box;
|
return box;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Expand the tree down to `path` and select it.
|
// Expand the tree down to `path` and select it.
|
||||||
function reveal(path) {
|
async function reveal(path) {
|
||||||
const rowFor = (p) => [...document.querySelectorAll(".node")].find((n) => n.dataset.path === p);
|
const rowFor = (p) => [...document.querySelectorAll(".node")].find((n) => n.dataset.path === p);
|
||||||
let cur = "/";
|
let cur = "/";
|
||||||
for (const part of path.split("/").filter(Boolean)) {
|
for (const part of path.split("/").filter(Boolean)) {
|
||||||
const row = rowFor(cur);
|
const row = rowFor(cur);
|
||||||
if (!row) return;
|
if (!row) return;
|
||||||
const kids = row.parentElement.querySelector(":scope > ul");
|
const kids = row.parentElement.querySelector(":scope > ul");
|
||||||
if (!kids || kids.hidden) row.click();
|
if (!kids || kids.hidden) await activate(row);
|
||||||
cur = joinPath(cur, part);
|
cur = joinPath(cur, part);
|
||||||
}
|
}
|
||||||
const target = rowFor(cur);
|
const target = rowFor(cur);
|
||||||
if (target && target !== selectedNode) target.click();
|
if (target && target !== selectedNode) await activate(target);
|
||||||
}
|
}
|
||||||
|
|
||||||
// ?file=<url>&path=<object> opens a file from a URL (same origin, or one
|
// ?file=<url>&path=<object> opens a file from a URL (same origin, or one
|
||||||
// that allows CORS) and selects an object in it.
|
// that allows CORS) by range requests and selects an object in it.
|
||||||
const params = new URLSearchParams(location.search);
|
const params = new URLSearchParams(location.search);
|
||||||
if (params.get("file")) {
|
if (params.get("file")) {
|
||||||
const url = params.get("file");
|
await loadUrl(params.get("file"));
|
||||||
try {
|
if (file && params.get("path")) await reveal(params.get("path"));
|
||||||
const resp = await fetch(url);
|
document.body.dataset.ready = "1";
|
||||||
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
|
|
||||||
const blob = await resp.blob();
|
|
||||||
await load(new File([blob], url.split("/").pop()));
|
|
||||||
if (file && params.get("path")) reveal(params.get("path"));
|
|
||||||
} catch (e) {
|
|
||||||
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot fetch ${url}: ${e.message}` }));
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
|
$("urlform").addEventListener("submit", (e) => {
|
||||||
|
e.preventDefault();
|
||||||
|
const url = $("url").value.trim();
|
||||||
|
if (url) loadUrl(url);
|
||||||
|
});
|
||||||
$("picker").addEventListener("change", (e) => e.target.files[0] && load(e.target.files[0]));
|
$("picker").addEventListener("change", (e) => e.target.files[0] && load(e.target.files[0]));
|
||||||
document.addEventListener("dragover", (e) => { e.preventDefault(); document.body.classList.add("dragging"); });
|
document.addEventListener("dragover", (e) => { e.preventDefault(); document.body.classList.add("dragging"); });
|
||||||
document.addEventListener("dragleave", () => document.body.classList.remove("dragging"));
|
document.addEventListener("dragleave", () => document.body.classList.remove("dragging"));
|
||||||
|
|||||||
@@ -3,10 +3,13 @@
|
|||||||
#
|
#
|
||||||
# browser.sh FIXTURE_DIR
|
# browser.sh FIXTURE_DIR
|
||||||
#
|
#
|
||||||
# FIXTURE_DIR holds fixture.h5 from make_fixture.py; ../pkg must be built.
|
# FIXTURE_DIR holds fixture.h5 from make_fixture.py (and big.h5 when it was
|
||||||
# The page is opened with ?file=fixture.h5&path=<object>, which fetches the
|
# run with WASM_BIG_MB); ../pkg must be built. test/serve.py serves the
|
||||||
# file, builds the tree down to <object> and shows it; the rendered DOM is
|
# page and, under /fix, the fixtures, with HTTP range requests (no copies,
|
||||||
# dumped and checked for the values libhdf5 reads.
|
# no links). The page is opened with ?file=fix/<name>&path=<object>, which
|
||||||
|
# opens the file with openUrl (range requests), builds the tree down to
|
||||||
|
# <object> and shows it; the rendered DOM is dumped and checked for the
|
||||||
|
# values libhdf5 reads and for the request counter.
|
||||||
#
|
#
|
||||||
# Browser: $CHROME, else chromium/google-chrome on PATH, else a Playwright
|
# Browser: $CHROME, else chromium/google-chrome on PATH, else a Playwright
|
||||||
# download under ~/.cache/ms-playwright. Exit 3 when none is found.
|
# download under ~/.cache/ms-playwright. Exit 3 when none is found.
|
||||||
@@ -31,63 +34,90 @@ if [ -z "$chrome" ] || [ ! -x "$chrome" ]; then
|
|||||||
exit 3
|
exit 3
|
||||||
fi
|
fi
|
||||||
|
|
||||||
root="$(mktemp -d)"
|
profiles="$(mktemp -d)"
|
||||||
server=""
|
server=""
|
||||||
cleanup() {
|
cleanup() {
|
||||||
[ -n "$server" ] && kill "$server" 2>/dev/null || true
|
[ -n "$server" ] && kill "$server" 2>/dev/null || true
|
||||||
rm -rf "$root"
|
rm -rf "$profiles"
|
||||||
}
|
}
|
||||||
trap cleanup EXIT
|
trap cleanup EXIT
|
||||||
ln -s "$HERE/../index.html" "$HERE/../viewer-lib.js" "$HERE/../pkg" "$FIX/fixture.h5" "$root/"
|
|
||||||
|
|
||||||
port=$("$PY" -c 'import socket; s = socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1])')
|
"$PY" "$HERE/serve.py" --root "fix=$FIX" --root "=$HERE/.." > "$profiles/port" &
|
||||||
"$PY" -m http.server --bind 127.0.0.1 --directory "$root" "$port" >/dev/null 2>&1 &
|
|
||||||
server=$!
|
server=$!
|
||||||
for _ in $(seq 50); do
|
for _ in $(seq 50); do
|
||||||
"$PY" -c "import urllib.request; urllib.request.urlopen('http://127.0.0.1:$port/index.html')" 2>/dev/null && break
|
[ -s "$profiles/port" ] && break
|
||||||
sleep 0.1
|
sleep 0.1
|
||||||
done
|
done
|
||||||
|
port="$(head -1 "$profiles/port")"
|
||||||
|
|
||||||
fails=0
|
fails=0
|
||||||
# A fresh profile per page: a second instance on the same profile fails.
|
# A fresh profile per page: a second instance on the same profile fails.
|
||||||
render() {
|
render() {
|
||||||
local profile
|
local profile
|
||||||
profile="$(mktemp -d "$root/profile.XXXXXX")"
|
profile="$(mktemp -d "$profiles/profile.XXXXXX")"
|
||||||
"$chrome" --headless --no-sandbox --disable-gpu --user-data-dir="$profile" \
|
"$chrome" --headless --no-sandbox --disable-gpu --user-data-dir="$profile" \
|
||||||
--virtual-time-budget=20000 \
|
--virtual-time-budget=20000 \
|
||||||
--dump-dom "http://127.0.0.1:$port/index.html?file=fixture.h5&path=$1" 2>/dev/null
|
--dump-dom "http://127.0.0.1:$port/index.html?file=$1&path=$2" 2>/dev/null
|
||||||
}
|
}
|
||||||
# expect PATH TEXT...: every TEXT appears in the page rendered for PATH.
|
# expect FILE PATH TEXT...: every TEXT appears in the page rendered for PATH
|
||||||
|
# of FILE (a URL relative to the page); "re:TEXT" is an extended regex.
|
||||||
expect() {
|
expect() {
|
||||||
local path="$1" dom
|
local file="$1" path="$2" dom
|
||||||
shift
|
shift 2
|
||||||
dom="$(render "$path")"
|
dom="$(render "$file" "$path")"
|
||||||
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$path" | tr / _).html"
|
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$file$path" | tr / _).html"
|
||||||
for text in "$@"; do
|
for text in "$@"; do
|
||||||
if ! grep -qF -- "$text" <<<"$dom"; then
|
local flag=-qF
|
||||||
echo "FAIL: page for $path lacks: $text" >&2
|
if [ "${text#re:}" != "$text" ]; then flag=-qE; text="${text#re:}"; fi
|
||||||
|
if ! grep $flag -- "$text" <<<"$dom"; then
|
||||||
|
echo "FAIL: page for $file $path lacks: $text" >&2
|
||||||
fails=$((fails + 1))
|
fails=$((fails + 1))
|
||||||
fi
|
fi
|
||||||
done
|
done
|
||||||
echo "rendered $path"
|
echo "rendered $file $path"
|
||||||
}
|
}
|
||||||
|
|
||||||
|
H5=fix/fixture.h5
|
||||||
# Tree (root expanded; the group row carries its path) and root attributes.
|
# Tree (root expanded; the group row carries its path) and root attributes.
|
||||||
expect "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
|
expect $H5 "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
|
||||||
'<td>big</td><td>9223372036854775813</td>' '(compound{x: f64, n: i32})'
|
'<td>big</td><td>9223372036854775813</td>' '(compound{x: f64, n: i32})'
|
||||||
# A chunked, deflated 2-D dataset: type, shape, the first window of values.
|
# A chunked, deflated 2-D dataset: type, shape, the first window of values;
|
||||||
expect "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
|
# the request counter.
|
||||||
'showing rows 0–5, columns 0–9 of 60 values'
|
expect $H5 "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
|
||||||
|
'showing rows 0–5, columns 0–9 of 60 values' 'request' 'fetched (100.0%)'
|
||||||
# Nested path revealed through the tree; big-endian float32.
|
# Nested path revealed through the tree; big-endian float32.
|
||||||
expect "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
|
expect $H5 "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
|
||||||
# 64-bit integers stay exact; strings; array datatype cells.
|
# 64-bit integers stay exact; strings; array datatype cells.
|
||||||
expect "/u64" '<td>18446744073709551615</td>'
|
expect $H5 "/u64" '<td>18446744073709551615</td>'
|
||||||
expect "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
|
expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
|
||||||
expect "/pairs" '<td>[2, 3]</td>' '<dd>array[2]<i32></dd>'
|
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]<i32></dd>'
|
||||||
# 3-D: leading dimension held at 0, window over the last two.
|
# 3-D: leading dimension held at 0, window over the last two.
|
||||||
expect "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
|
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
|
||||||
# Unsupported type: an error, not values.
|
# Unsupported type: an error, not values.
|
||||||
expect "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
|
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
|
||||||
|
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
|
||||||
|
# exposes Content-Range, and CORS that exposes nothing (/unexposed/: the
|
||||||
|
# browser hides Content-Range and ETag, so the page learns the length with
|
||||||
|
# a HEAD request, which the server's log must show).
|
||||||
|
expect "http://localhost:$port/$H5" "/grid" '<td>14.75</td>' 'request'
|
||||||
|
get() { "$PY" -c 'import sys, urllib.request; print(urllib.request.urlopen(sys.argv[1]).read().decode())' "$1"; }
|
||||||
|
get "http://127.0.0.1:$port/__reset" >/dev/null
|
||||||
|
expect "http://localhost:$port/unexposed/$H5" "/grid" '<td>0.25</td>' '<td>14.75</td>' 'request'
|
||||||
|
heads="$(get "http://127.0.0.1:$port/__stats" | "$PY" -c \
|
||||||
|
'import json, sys; print(sum(1 for l in json.load(sys.stdin)["log"] if l[4] == "HEAD" and l[0].endswith("fixture.h5")))')"
|
||||||
|
if [ "$heads" -lt 1 ]; then
|
||||||
|
echo "FAIL: no HEAD request for the length with unexposed headers" >&2
|
||||||
|
fails=$((fails + 1))
|
||||||
|
fi
|
||||||
|
# A server without range support: the file is downloaded whole, and says so.
|
||||||
|
expect norange/$H5 "/grid" '<td>14.75</td>' 'downloaded whole'
|
||||||
|
# A large file: a small dataset, and a window of the large one, fetch a few
|
||||||
|
# blocks (the counter shows a small percentage: "(1.5%)", not "(100.0%)").
|
||||||
|
if [ -f "$FIX/big.h5" ]; then
|
||||||
|
expect fix/big.h5 "/small" '<td>1.5</td>' '<td>3.25</td>' 're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
|
||||||
|
expect fix/big.h5 "/big" '<dd>(25000000)</dd>' '<td>0</td>' '<td>24.5</td>' \
|
||||||
|
're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
|
||||||
|
fi
|
||||||
|
|
||||||
if [ "$fails" -gt 0 ]; then
|
if [ "$fails" -gt 0 ]; then
|
||||||
echo "browser: $fails checks failed" >&2
|
echo "browser: $fails checks failed" >&2
|
||||||
|
|||||||
@@ -3,7 +3,9 @@ reads back from them, for the clawhdf5-wasm tests.
|
|||||||
|
|
||||||
python make_fixture.py OUT_DIR
|
python make_fixture.py OUT_DIR
|
||||||
|
|
||||||
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json.
|
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json,
|
||||||
|
and the limit-test files OUT_DIR/limits.h5 and OUT_DIR/hostile_vl.h5 (see
|
||||||
|
write_limits).
|
||||||
Both the Rust test (crates/clawhdf5-wasm/tests/h5py_interop.rs, native) and
|
Both the Rust test (crates/clawhdf5-wasm/tests/h5py_interop.rs, native) and
|
||||||
the Node test (test.mjs, the built wasm package) compare against the same
|
the Node test (test.mjs, the built wasm package) compare against the same
|
||||||
expected.json, so the two check the same values.
|
expected.json, so the two check the same values.
|
||||||
@@ -14,6 +16,7 @@ encoded as strings so JSON.parse keeps 64-bit values exact.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
import json
|
import json
|
||||||
|
import os
|
||||||
import sys
|
import sys
|
||||||
import warnings
|
import warnings
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
@@ -201,3 +204,105 @@ def slab_for(obj):
|
|||||||
|
|
||||||
json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)},
|
json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)},
|
||||||
open(out / "expected.json", "w"), indent=1, ensure_ascii=False)
|
open(out / "expected.json", "w"), indent=1, ensure_ascii=False)
|
||||||
|
|
||||||
|
|
||||||
|
HUGE_U8 = 2**28 + 1024
|
||||||
|
# The collection size hostile_vl.h5 claims: past 2 GiB, which a wasm32
|
||||||
|
# buffer cannot hold.
|
||||||
|
HOSTILE_GCOL_SIZE = 2**31 + 4096
|
||||||
|
# The file length a server claims for hostile_vl.h5 (the tests' mock fetch
|
||||||
|
# answers every range with zeros past the real bytes): 3 GiB, within what
|
||||||
|
# wasm32 opens, and room for the collection.
|
||||||
|
HOSTILE_LENGTH = 3 << 30
|
||||||
|
|
||||||
|
|
||||||
|
def write_limits(out):
|
||||||
|
"""Files for the size limits (the tests must get errors, not aborts):
|
||||||
|
|
||||||
|
- limits.h5: /huge_u8, 2^28 + 1024 bytes of u8 in compressed chunks
|
||||||
|
(a small file): read whole it would take over 2 GiB while decoding;
|
||||||
|
its last value is 7.
|
||||||
|
- hostile_vl.h5: a variable-length string dataset /a whose global heap
|
||||||
|
collection claims HOSTILE_GCOL_SIZE bytes, with the superblock's end
|
||||||
|
of file set to HOSTILE_LENGTH (libhdf5 cannot read it; it is only
|
||||||
|
served by a mock that claims that length).
|
||||||
|
- far.h5 and far.json: /x, 16 float64 values, whose contiguous data
|
||||||
|
address is moved FAR_SHIFT bytes on (past 2 GiB, the sign bit of a
|
||||||
|
wasm32 isize) in a file whose end of file is moved as far; the tests'
|
||||||
|
mock serves the data there, to show offsets up to 4 GiB work on
|
||||||
|
wasm32.
|
||||||
|
"""
|
||||||
|
with h5py.File(out / "limits.h5", "w") as f:
|
||||||
|
d = f.create_dataset("huge_u8", shape=(HUGE_U8,), dtype="u1",
|
||||||
|
chunks=(1 << 20,), compression="gzip")
|
||||||
|
d[-1] = 7
|
||||||
|
path = out / "hostile_vl.h5"
|
||||||
|
with h5py.File(path, "w", libver="earliest") as f:
|
||||||
|
f.create_dataset("a", data=["x", "yy"], dtype=h5py.string_dtype())
|
||||||
|
b = bytearray(path.read_bytes())
|
||||||
|
assert b[8] == 0, "a version 0 superblock"
|
||||||
|
b[40:48] = HOSTILE_LENGTH.to_bytes(8, "little") # end of file address
|
||||||
|
at = b.index(b"GCOL")
|
||||||
|
b[at + 8:at + 16] = HOSTILE_GCOL_SIZE.to_bytes(8, "little")
|
||||||
|
path.write_bytes(bytes(b))
|
||||||
|
|
||||||
|
path = out / "far.h5"
|
||||||
|
values = np.arange(16, dtype="<f8") * 1.5
|
||||||
|
with h5py.File(path, "w", libver="earliest") as f:
|
||||||
|
f.create_dataset("x", data=values)
|
||||||
|
data_at = f["x"].id.get_offset()
|
||||||
|
b = bytearray(path.read_bytes())
|
||||||
|
# The layout message: the data's address, then its size.
|
||||||
|
old = data_at.to_bytes(8, "little") + (values.nbytes).to_bytes(8, "little")
|
||||||
|
assert b.count(old) == 1
|
||||||
|
at = b.index(old)
|
||||||
|
b[at:at + 8] = (data_at + FAR_SHIFT).to_bytes(8, "little")
|
||||||
|
length = len(b) + FAR_SHIFT
|
||||||
|
b[40:48] = length.to_bytes(8, "little")
|
||||||
|
path.write_bytes(bytes(b))
|
||||||
|
json.dump({"data_at": data_at, "far_at": data_at + FAR_SHIFT,
|
||||||
|
"nbytes": values.nbytes, "length": length,
|
||||||
|
"values": [float(x) for x in values]},
|
||||||
|
open(out / "far.json", "w"))
|
||||||
|
|
||||||
|
|
||||||
|
# far.h5's data moves this far: past 2 GiB, below 4 GiB.
|
||||||
|
FAR_SHIFT = 3 << 30
|
||||||
|
|
||||||
|
write_limits(out)
|
||||||
|
|
||||||
|
|
||||||
|
def write_big(path, megabytes):
|
||||||
|
"""A large file for the range-request tests (`openUrl`): `/big`, about
|
||||||
|
`megabytes` MB of float64 in 1 MiB chunks, written after a small
|
||||||
|
dataset and a group, so listing and reading `/small` touch a few blocks
|
||||||
|
of the file and a window of `/big` one chunk. Returns what h5py reads
|
||||||
|
back."""
|
||||||
|
n = megabytes * 1_000_000 // 8
|
||||||
|
chunk = 1 << 17
|
||||||
|
with h5py.File(path, "w") as f:
|
||||||
|
f.attrs["note"] = "large file for range reads"
|
||||||
|
f.create_dataset("small", data=np.array([1.5, -2.0, 3.25]))
|
||||||
|
g = f.create_group("meta")
|
||||||
|
g.attrs["units"] = "m"
|
||||||
|
g.create_dataset("ids", data=np.arange(10, dtype="<i4"))
|
||||||
|
big = f.create_dataset("big", shape=(n,), dtype="<f8", chunks=(chunk,))
|
||||||
|
for s in range(0, n, 1 << 22):
|
||||||
|
e = min(n, s + (1 << 22))
|
||||||
|
big[s:e] = np.arange(s, e, dtype="<f8") * 0.5
|
||||||
|
start = n // 2 + 12_345
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
return {
|
||||||
|
"size": path.stat().st_size,
|
||||||
|
"list": {"groups": ["meta"], "datasets": ["big", "small"]},
|
||||||
|
"small": [float(x) for x in f["small"][()]],
|
||||||
|
"ids": [str(int(x)) for x in f["meta/ids"][()]],
|
||||||
|
"big_shape": list(f["big"].shape),
|
||||||
|
"window": {"start": start, "count": 10,
|
||||||
|
"values": [float(x) for x in f["big"][start:start + 10]]},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
big_mb = int(os.environ.get("WASM_BIG_MB", "0"))
|
||||||
|
if big_mb > 0:
|
||||||
|
json.dump(write_big(out / "big.h5", big_mb), open(out / "big.json", "w"), indent=1)
|
||||||
|
|||||||
@@ -1,11 +1,17 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
# Build the wasm package (../build.sh), test it under Node against files
|
# Build the wasm package (../build.sh), test it under Node against files
|
||||||
# written by h5py and netCDF4 (make_fixture.py), then load the viewer page
|
# written by h5py and netCDF4 (make_fixture.py) — opened from bytes, and
|
||||||
# in headless Chromium if one is found (browser.sh).
|
# opened by URL from a local range-capable HTTP server (serve.py) — then
|
||||||
|
# load the viewer page in headless Chromium if one is found (browser.sh).
|
||||||
#
|
#
|
||||||
# Needs node, the wasm-bindgen CLI (see ../build.sh) and a Python with h5py,
|
# Needs node, the wasm-bindgen CLI (see ../build.sh) and a Python with h5py,
|
||||||
# netCDF4 and numpy: CLAWHDF5_PYTHON names it (default python3). Without that
|
# netCDF4 and numpy: CLAWHDF5_PYTHON names it (default python3). Without that
|
||||||
# Python the test is skipped, unless CLAWHDF5_REQUIRE_INTEROP=1.
|
# Python the test is skipped, unless CLAWHDF5_REQUIRE_INTEROP=1.
|
||||||
|
#
|
||||||
|
# WASM_BIG_MB (default 200) sizes the large file of the range-request
|
||||||
|
# budget test (0 leaves it out); it is written under TMPDIR.
|
||||||
|
# CLAWHDF5_WASM_CORPUS=DIR also compares every HDF5 file under DIR (up to
|
||||||
|
# 16 MiB) read over HTTP with the same file read from bytes.
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||||
@@ -23,9 +29,24 @@ fi
|
|||||||
|
|
||||||
bash "$HERE/../build.sh"
|
bash "$HERE/../build.sh"
|
||||||
fix="$(mktemp -d)"
|
fix="$(mktemp -d)"
|
||||||
trap 'rm -rf "$fix"' EXIT
|
server=""
|
||||||
"$PY" "$HERE/make_fixture.py" "$fix"
|
cleanup() {
|
||||||
node "$HERE/test.mjs" "$HERE/../pkg" "$fix"
|
[ -n "$server" ] && kill "$server" 2>/dev/null || true
|
||||||
|
rm -rf "$fix"
|
||||||
|
}
|
||||||
|
trap cleanup EXIT
|
||||||
|
WASM_BIG_MB="${WASM_BIG_MB:-200}" "$PY" "$HERE/make_fixture.py" "$fix"
|
||||||
|
|
||||||
|
# The fixtures (and the corpus) over HTTP with range support.
|
||||||
|
roots=(--root "fix=$fix")
|
||||||
|
[ -n "${CLAWHDF5_WASM_CORPUS:-}" ] && roots+=(--root "corpus=$CLAWHDF5_WASM_CORPUS")
|
||||||
|
"$PY" "$HERE/serve.py" "${roots[@]}" > "$fix/port" &
|
||||||
|
server=$!
|
||||||
|
for _ in $(seq 50); do
|
||||||
|
[ -s "$fix/port" ] && break
|
||||||
|
sleep 0.1
|
||||||
|
done
|
||||||
|
node "$HERE/test.mjs" "$HERE/../pkg" "$fix" "http://127.0.0.1:$(head -1 "$fix/port")"
|
||||||
|
|
||||||
# The page itself, in headless Chromium when one is available.
|
# The page itself, in headless Chromium when one is available.
|
||||||
status=0
|
status=0
|
||||||
|
|||||||
@@ -0,0 +1,223 @@
|
|||||||
|
"""A static HTTP server for the wasm tests, with HTTP Range support and
|
||||||
|
request counting.
|
||||||
|
|
||||||
|
python serve.py [--root PREFIX=DIR ...]
|
||||||
|
|
||||||
|
Serves each DIR under URL PREFIX (the first match wins; PREFIX "" is the
|
||||||
|
site root), prints the port on its first line of stdout, and runs until
|
||||||
|
killed. No symlinks or copies are made: files are read where they are.
|
||||||
|
|
||||||
|
- `Range: bytes=a-b`, `bytes=a-` and `bytes=-n` get 206 with Content-Range,
|
||||||
|
an unsatisfiable range 416; every file answer carries an ETag, and CORS
|
||||||
|
headers exposing Content-Range, so a page on another origin can use it.
|
||||||
|
- Under `/norange/...` the same files are served but Range is ignored
|
||||||
|
(200 with the whole file), as by a server without range support.
|
||||||
|
- Under `/noexpose/...` ranges are served without Content-Range, ETag,
|
||||||
|
Last-Modified or Accept-Ranges, and CORS exposes none of them: what a
|
||||||
|
page sees of a cross-origin server that does not list them in
|
||||||
|
Access-Control-Expose-Headers. The length comes from Content-Length of
|
||||||
|
a HEAD request (always readable).
|
||||||
|
- Under `/unexposed/...` ranges are served with all those headers, but
|
||||||
|
CORS exposes none of them: a browser page on another origin cannot read
|
||||||
|
them (test/browser.sh loads the page from 127.0.0.1 and the file from
|
||||||
|
localhost), so it has to take the same path.
|
||||||
|
- `GET /__stats` returns `{"requests": n, "bytes": n, "log": [...]}` for
|
||||||
|
file requests since the last `GET /__reset`, which zeroes them.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import posixpath
|
||||||
|
import sys
|
||||||
|
import threading
|
||||||
|
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||||
|
from urllib.parse import unquote, urlsplit
|
||||||
|
|
||||||
|
TYPES = {
|
||||||
|
".html": "text/html; charset=utf-8",
|
||||||
|
".js": "text/javascript; charset=utf-8",
|
||||||
|
".mjs": "text/javascript; charset=utf-8",
|
||||||
|
".wasm": "application/wasm",
|
||||||
|
".json": "application/json",
|
||||||
|
".ts": "text/plain; charset=utf-8",
|
||||||
|
}
|
||||||
|
|
||||||
|
lock = threading.Lock()
|
||||||
|
stats = {"requests": 0, "bytes": 0, "log": []}
|
||||||
|
|
||||||
|
|
||||||
|
def resolve(roots, path):
|
||||||
|
"""The file for URL `path`, or None. `..` never leaves a root."""
|
||||||
|
parts = [p for p in posixpath.normpath(unquote(path)).split("/") if p]
|
||||||
|
if any(p in (".", "..") for p in parts):
|
||||||
|
return None
|
||||||
|
for prefix, root in roots:
|
||||||
|
pre = [p for p in prefix.split("/") if p]
|
||||||
|
if parts[: len(pre)] == pre:
|
||||||
|
rest = parts[len(pre):] or ["index.html"]
|
||||||
|
f = os.path.join(root, *rest)
|
||||||
|
if os.path.isfile(f):
|
||||||
|
return f
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def parse_range(header, size):
|
||||||
|
"""(start, end exclusive) for a single `bytes=` range, "bad" when
|
||||||
|
unsatisfiable, None when absent or unparsable (served whole)."""
|
||||||
|
if not header or not header.startswith("bytes=") or "," in header:
|
||||||
|
return None
|
||||||
|
a, _, b = header[len("bytes="):].strip().partition("-")
|
||||||
|
try:
|
||||||
|
if a == "":
|
||||||
|
n = int(b)
|
||||||
|
return (max(0, size - n), size) if n > 0 and size > 0 else "bad"
|
||||||
|
start = int(a)
|
||||||
|
end = int(b) + 1 if b else size
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
if start >= size or end <= start:
|
||||||
|
return "bad"
|
||||||
|
return start, min(end, size)
|
||||||
|
|
||||||
|
|
||||||
|
def make_handler(roots):
|
||||||
|
class Handler(BaseHTTPRequestHandler):
|
||||||
|
protocol_version = "HTTP/1.1"
|
||||||
|
|
||||||
|
def log_message(self, *args):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def cors(self, expose=True):
|
||||||
|
self.send_header("Access-Control-Allow-Origin", "*")
|
||||||
|
if expose:
|
||||||
|
self.send_header("Access-Control-Expose-Headers",
|
||||||
|
"Content-Range, Content-Length, ETag, Accept-Ranges")
|
||||||
|
|
||||||
|
def do_OPTIONS(self):
|
||||||
|
self.send_response(204)
|
||||||
|
self.cors()
|
||||||
|
self.send_header("Access-Control-Allow-Headers", "Range")
|
||||||
|
self.send_header("Content-Length", "0")
|
||||||
|
self.end_headers()
|
||||||
|
|
||||||
|
def do_HEAD(self):
|
||||||
|
self.serve(head=True)
|
||||||
|
|
||||||
|
def do_GET(self):
|
||||||
|
self.serve(head=False)
|
||||||
|
|
||||||
|
def json(self, obj):
|
||||||
|
body = json.dumps(obj).encode()
|
||||||
|
self.send_response(200)
|
||||||
|
self.cors()
|
||||||
|
self.send_header("Content-Type", "application/json")
|
||||||
|
self.send_header("Content-Length", str(len(body)))
|
||||||
|
self.send_header("Cache-Control", "no-store")
|
||||||
|
self.end_headers()
|
||||||
|
self.wfile.write(body)
|
||||||
|
|
||||||
|
def serve(self, head):
|
||||||
|
path = urlsplit(self.path).path
|
||||||
|
if path == "/__stats":
|
||||||
|
with lock:
|
||||||
|
return self.json(stats)
|
||||||
|
if path == "/__reset":
|
||||||
|
with lock:
|
||||||
|
stats.update(requests=0, bytes=0, log=[])
|
||||||
|
return self.json({})
|
||||||
|
# ranges: honour Range; send_all: send Content-Range, ETag and
|
||||||
|
# Accept-Ranges; expose: list them for CORS.
|
||||||
|
ranges = send_all = expose = True
|
||||||
|
if path.startswith("/norange/"):
|
||||||
|
ranges = False
|
||||||
|
path = path[len("/norange"):]
|
||||||
|
elif path.startswith("/noexpose/"):
|
||||||
|
expose = send_all = False
|
||||||
|
path = path[len("/noexpose"):]
|
||||||
|
elif path.startswith("/unexposed/"):
|
||||||
|
expose = False
|
||||||
|
path = path[len("/unexposed"):]
|
||||||
|
f = resolve(roots, path)
|
||||||
|
if f is None:
|
||||||
|
self.send_response(404)
|
||||||
|
self.cors()
|
||||||
|
self.send_header("Content-Length", "0")
|
||||||
|
self.end_headers()
|
||||||
|
return
|
||||||
|
size = os.path.getsize(f)
|
||||||
|
st = os.stat(f)
|
||||||
|
etag = '"%s"' % hashlib.sha1(
|
||||||
|
f"{f}:{size}:{st.st_mtime_ns}".encode()).hexdigest()[:16]
|
||||||
|
r = parse_range(self.headers.get("Range"), size) if ranges else None
|
||||||
|
if r == "bad":
|
||||||
|
self.send_response(416)
|
||||||
|
self.cors()
|
||||||
|
self.send_header("Content-Range", f"bytes */{size}")
|
||||||
|
self.send_header("Content-Length", "0")
|
||||||
|
self.end_headers()
|
||||||
|
return
|
||||||
|
start, end = r if r else (0, size)
|
||||||
|
self.send_response(206 if r else 200)
|
||||||
|
self.cors(expose)
|
||||||
|
ext = os.path.splitext(f)[1]
|
||||||
|
self.send_header("Content-Type", TYPES.get(ext, "application/octet-stream"))
|
||||||
|
self.send_header("Content-Length", str(end - start))
|
||||||
|
self.send_header("Cache-Control", "no-store")
|
||||||
|
if send_all:
|
||||||
|
self.send_header("ETag", etag)
|
||||||
|
if ranges:
|
||||||
|
self.send_header("Accept-Ranges", "bytes")
|
||||||
|
if r:
|
||||||
|
self.send_header("Content-Range", f"bytes {start}-{end - 1}/{size}")
|
||||||
|
self.end_headers()
|
||||||
|
with lock:
|
||||||
|
# A HEAD is a request too (openUrl makes one when it cannot
|
||||||
|
# see Content-Range); it sends no bytes.
|
||||||
|
stats["requests"] += 1
|
||||||
|
sent = 0 if head else end - start
|
||||||
|
stats["bytes"] += sent
|
||||||
|
stats["log"].append([path, start, end, 206 if r else 200, "HEAD" if head else "GET"])
|
||||||
|
if not head:
|
||||||
|
with open(f, "rb") as fh:
|
||||||
|
fh.seek(start)
|
||||||
|
left = end - start
|
||||||
|
try:
|
||||||
|
while left:
|
||||||
|
buf = fh.read(min(left, 1 << 20))
|
||||||
|
if not buf:
|
||||||
|
break
|
||||||
|
self.wfile.write(buf)
|
||||||
|
left -= len(buf)
|
||||||
|
except (BrokenPipeError, ConnectionResetError):
|
||||||
|
pass
|
||||||
|
|
||||||
|
return Handler
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument("--root", action="append", default=[],
|
||||||
|
help="PREFIX=DIR: serve DIR under URL PREFIX")
|
||||||
|
args = ap.parse_args()
|
||||||
|
roots = []
|
||||||
|
for spec in args.root:
|
||||||
|
prefix, _, d = spec.partition("=")
|
||||||
|
roots.append((prefix, os.path.abspath(d)))
|
||||||
|
class Server(ThreadingHTTPServer):
|
||||||
|
def handle_error(self, request, client_address):
|
||||||
|
# A client that drops a connection (a cancelled download) is
|
||||||
|
# not an error of the server.
|
||||||
|
if not isinstance(sys.exc_info()[1], (ConnectionError, TimeoutError)):
|
||||||
|
super().handle_error(request, client_address)
|
||||||
|
|
||||||
|
httpd = Server(("127.0.0.1", 0), make_handler(roots))
|
||||||
|
httpd.daemon_threads = True
|
||||||
|
print(httpd.server_address[1], flush=True)
|
||||||
|
sys.stdout.close()
|
||||||
|
httpd.serve_forever()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -1,20 +1,31 @@
|
|||||||
// Node test of the built wasm package (the exact pkg/ the viewer page loads)
|
// Node test of the built wasm package (the exact pkg/ the viewer page loads)
|
||||||
// and the viewer's DOM-free helpers. Run by test/run.sh:
|
// and the viewer's DOM-free helpers. Run by test/run.sh:
|
||||||
// node test.mjs PKG_DIR FIXTURE_DIR
|
// node test.mjs PKG_DIR FIXTURE_DIR [SERVER_URL]
|
||||||
// FIXTURE_DIR holds fixture.h5, fixture.nc and expected.json from
|
// FIXTURE_DIR holds fixture.h5, fixture.nc and expected.json from
|
||||||
// make_fixture.py (values as libhdf5 reads them back).
|
// make_fixture.py (values as libhdf5 reads them back), and big.h5/big.json
|
||||||
|
// when it was run with WASM_BIG_MB. SERVER_URL is test/serve.py serving
|
||||||
|
// FIXTURE_DIR under /fix (and $CLAWHDF5_WASM_CORPUS under /corpus): with
|
||||||
|
// it, every check is repeated on files opened with openUrl (HTTP range
|
||||||
|
// requests), and the request budget, the full-download fallback and the
|
||||||
|
// error paths of openUrl are tested.
|
||||||
import assert from "node:assert/strict";
|
import assert from "node:assert/strict";
|
||||||
import { readFileSync } from "node:fs";
|
import { existsSync, readFileSync, readdirSync, statSync } from "node:fs";
|
||||||
import { join } from "node:path";
|
import { join, relative } from "node:path";
|
||||||
import { pathToFileURL } from "node:url";
|
import { pathToFileURL } from "node:url";
|
||||||
|
|
||||||
const [pkgDir, fixDir] = process.argv.slice(2);
|
const [pkgDir, fixDir, base] = process.argv.slice(2);
|
||||||
const pkg = await import(pathToFileURL(join(pkgDir, "clawhdf5_wasm.js")));
|
const pkg = await import(pathToFileURL(join(pkgDir, "clawhdf5_wasm.js")));
|
||||||
pkg.initSync({ module: readFileSync(join(pkgDir, "clawhdf5_wasm_bg.wasm")) });
|
pkg.initSync({ module: readFileSync(join(pkgDir, "clawhdf5_wasm_bg.wasm")) });
|
||||||
const lib = await import(pathToFileURL(join(import.meta.dirname, "..", "viewer-lib.js")));
|
const lib = await import(pathToFileURL(join(import.meta.dirname, "..", "viewer-lib.js")));
|
||||||
|
|
||||||
let checks = 0;
|
let checks = 0;
|
||||||
const eq = (a, b, msg) => { assert.deepEqual(a, b, msg); checks++; };
|
const eq = (a, b, msg) => { assert.deepEqual(a, b, msg); checks++; };
|
||||||
|
// A call that must fail: a thrown Error (open) or a rejected promise
|
||||||
|
// (openUrl), whose message matches `re`.
|
||||||
|
const fails = async (fn, re, msg) => {
|
||||||
|
await assert.rejects(async () => fn(), (e) => e instanceof Error && re.test(e.message), msg);
|
||||||
|
checks++;
|
||||||
|
};
|
||||||
|
|
||||||
const ARRAY_TYPES = {
|
const ARRAY_TYPES = {
|
||||||
f32: Float32Array, f64: Float64Array, i8: Int8Array, i16: Int16Array, i32: Int32Array,
|
f32: Float32Array, f64: Float64Array, i8: Int8Array, i16: Int16Array, i32: Int32Array,
|
||||||
@@ -54,13 +65,13 @@ function checkAttr(ctx, a, want) {
|
|||||||
assert.fail(`${ctx}: unknown expectation ${JSON.stringify(want)}`);
|
assert.fail(`${ctx}: unknown expectation ${JSON.stringify(want)}`);
|
||||||
}
|
}
|
||||||
|
|
||||||
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
|
// Every listing, dataset, error and attribute of `file` against what
|
||||||
for (const [name, exp] of Object.entries(expected)) {
|
// libhdf5 reads (expected.json). `file` is an H5File (synchronous methods)
|
||||||
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
|
// or a RemoteFile (promises): every call is awaited.
|
||||||
|
async function checkFile(name, exp, file) {
|
||||||
for (const [path, want] of Object.entries(exp.lists)) {
|
for (const [path, want] of Object.entries(exp.lists)) {
|
||||||
eq(file.kind(path), "group", `${name}:${path} kind`);
|
eq(await file.kind(path), "group", `${name}:${path} kind`);
|
||||||
const list = file.list(path);
|
const list = await file.list(path);
|
||||||
for (const [kind, key] of [["group", "groups"], ["dataset", "datasets"]]) {
|
for (const [kind, key] of [["group", "groups"], ["dataset", "datasets"]]) {
|
||||||
eq(list.filter((c) => c.kind === kind).map((c) => c.name).sort(), want[key], `${name}:${path} ${key}`);
|
eq(list.filter((c) => c.kind === kind).map((c) => c.name).sort(), want[key], `${name}:${path} ${key}`);
|
||||||
}
|
}
|
||||||
@@ -68,58 +79,67 @@ for (const [name, exp] of Object.entries(expected)) {
|
|||||||
|
|
||||||
for (const [path, want] of Object.entries(exp.datasets)) {
|
for (const [path, want] of Object.entries(exp.datasets)) {
|
||||||
const ctx = `${name}:${path}`;
|
const ctx = `${name}:${path}`;
|
||||||
eq(file.kind(path), "dataset", `${ctx} kind`);
|
eq(await file.kind(path), "dataset", `${ctx} kind`);
|
||||||
if (want.unavailable) {
|
if (want.unavailable) {
|
||||||
// The wasm build has no zstd (it links C): a clear error, no data.
|
// The wasm build has no zstd (it links C): a clear error, no data.
|
||||||
assert.throws(() => file.read(path), (e) => e.message.includes(want.unavailable), ctx);
|
await fails(() => file.read(path), new RegExp(want.unavailable), ctx);
|
||||||
checks++;
|
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
const info = file.info(path);
|
const info = await file.info(path);
|
||||||
eq([...info.shape, ...info.elementShape], want.shape, `${ctx} info shape`);
|
eq([...info.shape, ...info.elementShape], want.shape, `${ctx} info shape`);
|
||||||
const r = file.read(path);
|
const r = await file.read(path);
|
||||||
eq(r.shape, want.shape, `${ctx} shape`);
|
eq(r.shape, want.shape, `${ctx} shape`);
|
||||||
eq(r.dtype, info.dtype, `${ctx} dtype`);
|
eq(r.dtype, info.dtype, `${ctx} dtype`);
|
||||||
assert.ok(r.data instanceof ARRAY_TYPES[want.kind], `${ctx}: ${r.data.constructor.name} for ${want.kind}`);
|
assert.ok(r.data instanceof ARRAY_TYPES[want.kind], `${ctx}: ${r.data.constructor.name} for ${want.kind}`);
|
||||||
eq(values(want.kind, r.data), want.values, ctx);
|
eq(values(want.kind, r.data), want.values, ctx);
|
||||||
if (want.slab) {
|
if (want.slab) {
|
||||||
const s = want.slab;
|
const s = want.slab;
|
||||||
const part = file.readHyperslab(path, s.start, s.count, s.stride);
|
const part = await file.readHyperslab(path, s.start, s.count, s.stride);
|
||||||
eq(part.shape, s.shape, `${ctx} slab shape`);
|
eq(part.shape, s.shape, `${ctx} slab shape`);
|
||||||
eq(values(want.kind, part.data), s.values, `${ctx} slab`);
|
eq(values(want.kind, part.data), s.values, `${ctx} slab`);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
for (const [path, what] of Object.entries(exp.errors)) {
|
for (const [path, what] of Object.entries(exp.errors)) {
|
||||||
assert.throws(() => file.read(path), (e) => e instanceof Error && e.message.includes(what), `${name}:${path}`);
|
await fails(() => file.read(path), new RegExp(what), `${name}:${path}`);
|
||||||
checks++;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
for (const [path, want] of Object.entries(exp.attrs)) {
|
for (const [path, want] of Object.entries(exp.attrs)) {
|
||||||
const attrs = file.attrs(path);
|
const attrs = await file.attrs(path);
|
||||||
eq(file.attrErrors(path), [], `${name}:${path} attr errors`);
|
eq(await file.attrErrors(path), [], `${name}:${path} attr errors`);
|
||||||
const seen = attrs.filter((a) => !a.name.startsWith("_") && !exp.skip_attrs.includes(a.name));
|
const seen = attrs.filter((a) => !a.name.startsWith("_") && !exp.skip_attrs.includes(a.name));
|
||||||
eq(seen.map((a) => a.name).sort(), Object.keys(want).sort(), `${name}:${path} attr names`);
|
eq(seen.map((a) => a.name).sort(), Object.keys(want).sort(), `${name}:${path} attr names`);
|
||||||
for (const a of seen) checkAttr(`${name}:${path}@${a.name}`, a, want[a.name]);
|
for (const a of seen) checkAttr(`${name}:${path}@${a.name}`, a, want[a.name]);
|
||||||
}
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The error paths of the reader, for an H5File or a RemoteFile of
|
||||||
|
// fixture.h5.
|
||||||
|
async function checkErrors(h5) {
|
||||||
|
await fails(() => h5.read("/nope"), /./);
|
||||||
|
await fails(() => h5.list("/grid"), /not a group/);
|
||||||
|
await fails(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
|
||||||
|
await fails(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
|
||||||
|
await fails(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
|
||||||
|
await fails(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
|
||||||
|
// Big integers stay exact.
|
||||||
|
eq((await h5.read("/u64")).data[0], 18446744073709551615n, "u64 max");
|
||||||
|
}
|
||||||
|
|
||||||
|
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
|
||||||
|
for (const [name, exp] of Object.entries(expected)) {
|
||||||
|
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
|
||||||
|
await checkFile(name, exp, file);
|
||||||
file.free();
|
file.free();
|
||||||
}
|
}
|
||||||
|
|
||||||
// Errors reach JavaScript as thrown Errors, never as data.
|
// Errors reach JavaScript as thrown Errors, never as data.
|
||||||
const h5 = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.h5"))));
|
const h5 = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.h5"))));
|
||||||
const throwsMsg = (fn, re) => { assert.throws(fn, (e) => e instanceof Error && re.test(e.message)); checks++; };
|
await fails(() => pkg.open(new Uint8Array(64)), /./);
|
||||||
throwsMsg(() => pkg.open(new Uint8Array(64)), /./);
|
await checkErrors(h5);
|
||||||
throwsMsg(() => h5.read("/nope"), /./);
|
|
||||||
throwsMsg(() => h5.list("/grid"), /not a group/);
|
|
||||||
throwsMsg(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
|
|
||||||
throwsMsg(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
|
|
||||||
throwsMsg(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
|
|
||||||
throwsMsg(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
|
|
||||||
// Info for a dataset with an unlimited dimension (netCDF "time").
|
// Info for a dataset with an unlimited dimension (netCDF "time").
|
||||||
const nc = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.nc"))));
|
const nc = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.nc"))));
|
||||||
eq(nc.info("/time").maxshape, [null], "unlimited dimension is null");
|
eq(nc.info("/time").maxshape, [null], "unlimited dimension is null");
|
||||||
// Big integers stay exact.
|
|
||||||
eq(h5.read("/u64").data[0], 18446744073709551615n, "u64 max");
|
|
||||||
eq(typeof pkg.version(), "string", "version");
|
eq(typeof pkg.version(), "string", "version");
|
||||||
|
|
||||||
// Viewer helpers.
|
// Viewer helpers.
|
||||||
@@ -137,7 +157,486 @@ eq(lib.toRows(pairs.data, 2, 1, lib.perElement([2])), [["[2, 3]"], ["[4, 5]"]],
|
|||||||
eq(lib.formatValue(0.1 + 0.2), "0.3", "float formatting");
|
eq(lib.formatValue(0.1 + 0.2), "0.3", "float formatting");
|
||||||
eq(lib.formatValue(2n ** 64n - 1n), "18446744073709551615", "bigint formatting");
|
eq(lib.formatValue(2n ** 64n - 1n), "18446744073709551615", "bigint formatting");
|
||||||
eq(lib.formatValue("x"), '"x"', "string formatting");
|
eq(lib.formatValue("x"), '"x"', "string formatting");
|
||||||
|
eq(lib.formatBytes(0), "0 B", "bytes");
|
||||||
|
eq(lib.formatBytes(1536), "1.5 KiB", "KiB");
|
||||||
|
eq(lib.formatBytes(200 * 1024 * 1024), "200 MiB", "MiB");
|
||||||
|
eq(lib.formatStats({ lazy: true, requests: 3, bytesFetched: 2 << 20, size: 200 << 20 }),
|
||||||
|
"3 requests, 2 MiB of 200 MiB fetched (1.0%)", "stats line");
|
||||||
|
eq(lib.formatStats({ lazy: false, requests: 1, bytesFetched: 1024, size: 1024 }),
|
||||||
|
"downloaded whole (1 KiB): the server does not support range requests", "stats line, no ranges");
|
||||||
h5.free();
|
h5.free();
|
||||||
nc.free();
|
nc.free();
|
||||||
|
|
||||||
console.log(`wasm package: ${checks} checks passed`);
|
console.log(`wasm package: ${checks} checks passed`);
|
||||||
|
|
||||||
|
if (base) await remoteTests();
|
||||||
|
|
||||||
|
async function serverStats() {
|
||||||
|
return (await fetch(`${base}/__stats`)).json();
|
||||||
|
}
|
||||||
|
|
||||||
|
async function remoteTests() {
|
||||||
|
const before = checks;
|
||||||
|
await fetch(`${base}/__reset`);
|
||||||
|
|
||||||
|
// The same checks over HTTP range requests, at the default block size
|
||||||
|
// and at 512-byte blocks with a 4 KiB cache (almost every structure read
|
||||||
|
// a miss, evictions between calls).
|
||||||
|
for (const opts of [undefined, { blockSize: 512, cacheSize: 4096 }]) {
|
||||||
|
for (const [name, exp] of Object.entries(expected)) {
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/${name}`, opts);
|
||||||
|
await checkFile(`${name} (openUrl ${JSON.stringify(opts ?? {})})`, exp, f);
|
||||||
|
const st = f.stats();
|
||||||
|
eq(st.lazy, true, "read by ranges");
|
||||||
|
eq(st.size, statSync(join(fixDir, name)).size, "size");
|
||||||
|
f.free();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const remote = await pkg.openUrl(`${base}/fix/fixture.h5`);
|
||||||
|
await checkErrors(remote);
|
||||||
|
eq((await (await pkg.openUrl(`${base}/fix/fixture.nc`)).info("/time")).maxshape, [null], "remote unlimited");
|
||||||
|
|
||||||
|
// What the page counts is what the server served.
|
||||||
|
await fetch(`${base}/__reset`);
|
||||||
|
const counted = await pkg.openUrl(`${base}/fix/fixture.nc`, { blockSize: 1024 });
|
||||||
|
await counted.read("/temp");
|
||||||
|
const server = await serverStats();
|
||||||
|
eq(counted.stats().requests, server.requests, "requests counted");
|
||||||
|
eq(counted.stats().bytesFetched, server.bytes, "bytes counted");
|
||||||
|
|
||||||
|
// A custom fetch is used for every request, with the caller's headers,
|
||||||
|
// given in any form fetch takes (a caller's Range is not sent).
|
||||||
|
for (const headers of [{ "X-Test": "1" }, new Headers({ "X-Test": "1", Range: "bytes=0-0" }), [["X-Test", "1"]]]) {
|
||||||
|
let calls = 0;
|
||||||
|
const seen = new Set();
|
||||||
|
const viaCustom = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||||
|
blockSize: 4096,
|
||||||
|
headers,
|
||||||
|
fetch: (url, init) => {
|
||||||
|
calls++;
|
||||||
|
const h = new Headers(init.headers);
|
||||||
|
seen.add(`${h.get("X-Test")} ${h.get("Range").startsWith("bytes=0-0") ? "caller's range" : "ours"}`);
|
||||||
|
return fetch(url, init);
|
||||||
|
},
|
||||||
|
});
|
||||||
|
const kind = headers.constructor.name;
|
||||||
|
eq((await viaCustom.read("/sensors/temp")).data[0], 21.5, `custom fetch values (${kind})`);
|
||||||
|
eq(calls, viaCustom.stats().requests, `custom fetch calls (${kind})`);
|
||||||
|
eq([...seen], ["1 ours"], `headers passed (${kind})`);
|
||||||
|
}
|
||||||
|
|
||||||
|
// parallel must be a positive integer.
|
||||||
|
for (const parallel of [0, -1, 1.5, "4", NaN, 5000]) {
|
||||||
|
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { parallel }), /parallel/, `parallel ${String(parallel)}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
// When one range request fails, the others in flight are aborted and no
|
||||||
|
// more start (fetchRanges, called directly: remote.js is in the package).
|
||||||
|
{
|
||||||
|
const snippets = join(pkgDir, "snippets");
|
||||||
|
const dir = readdirSync(snippets).find((d) => existsSync(join(snippets, d, "js", "remote.js")));
|
||||||
|
const remote = await import(pathToFileURL(join(snippets, dir, "js", "remote.js")));
|
||||||
|
let started = 0;
|
||||||
|
const aborted = [];
|
||||||
|
const slowFetch = (url, init) => {
|
||||||
|
const k = started++;
|
||||||
|
if (k === 2) return Promise.resolve(new Response(null, { status: 500 }));
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const t = setTimeout(() => resolve(new Response(new Uint8Array(10), { status: 206 })), 200);
|
||||||
|
init.signal?.addEventListener("abort", () => {
|
||||||
|
clearTimeout(t);
|
||||||
|
aborted.push(k);
|
||||||
|
reject(new DOMException("aborted", "AbortError"));
|
||||||
|
});
|
||||||
|
});
|
||||||
|
};
|
||||||
|
const ranges = Array.from({ length: 20 }, (_, i) => [i * 10, i * 10 + 10]).flat();
|
||||||
|
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: 3 }, null, 200),
|
||||||
|
/HTTP 500/, "a failed range request is the error");
|
||||||
|
await new Promise((r) => setTimeout(r, 300));
|
||||||
|
eq(started, 3, "no request starts after a failure");
|
||||||
|
eq(aborted.sort(), [0, 1], "requests in flight are aborted");
|
||||||
|
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: "x" }, null, 200),
|
||||||
|
/parallel must be a positive integer/, "fetchRanges checks parallel");
|
||||||
|
}
|
||||||
|
|
||||||
|
// Calls in flight at once share the cache (a 1 KiB budget: nothing is
|
||||||
|
// evicted while any of them runs) and each gets its own answer.
|
||||||
|
const both = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, cacheSize: 1024 });
|
||||||
|
const [g, t, l, a] = await Promise.all([
|
||||||
|
both.read("/grid"), both.read("/sensors/temp"), both.list("/sensors"), both.attrs("/"),
|
||||||
|
]);
|
||||||
|
const exp = expected["fixture.h5"];
|
||||||
|
eq(Array.from(g.data), exp.datasets["/grid"].values, "concurrent /grid");
|
||||||
|
eq(Array.from(t.data), exp.datasets["/sensors/temp"].values, "concurrent /sensors/temp");
|
||||||
|
eq(l.map((c) => c.name).sort(), [...exp.lists["/sensors"].groups, ...exp.lists["/sensors"].datasets].sort(), "concurrent list");
|
||||||
|
eq(a.length > 0, true, "concurrent attrs");
|
||||||
|
assert.ok(both.stats().cachedBytes <= 1024, "trimmed to the budget when idle");
|
||||||
|
|
||||||
|
// Listing and reading small things of a large file fetches a few blocks,
|
||||||
|
// not the file.
|
||||||
|
if (existsSync(join(fixDir, "big.json"))) {
|
||||||
|
const big = JSON.parse(readFileSync(join(fixDir, "big.json"), "utf8"));
|
||||||
|
await fetch(`${base}/__reset`);
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/big.h5`);
|
||||||
|
const list = await f.list("/");
|
||||||
|
eq(list.filter((c) => c.kind === "group").map((c) => c.name), big.list.groups, "big: groups");
|
||||||
|
eq(list.filter((c) => c.kind === "dataset").map((c) => c.name).sort(), big.list.datasets, "big: datasets");
|
||||||
|
eq(Array.from((await f.read("/small")).data), big.small, "big: /small");
|
||||||
|
eq(Array.from((await f.read("/meta/ids")).data, String), big.ids, "big: /meta/ids");
|
||||||
|
eq((await f.info("/big")).shape, big.big_shape, "big: shape");
|
||||||
|
eq((await f.attrs("/meta"))[0].value, "m", "big: group attribute");
|
||||||
|
const win = await f.readHyperslab("/big", [big.window.start], [big.window.count]);
|
||||||
|
eq(Array.from(win.data), big.window.values, "big: window");
|
||||||
|
const st = f.stats();
|
||||||
|
const server = await serverStats();
|
||||||
|
eq(st.requests, server.requests, "big: requests counted");
|
||||||
|
eq(st.bytesFetched, server.bytes, "big: bytes counted");
|
||||||
|
console.log(`big.h5 (${big.size} bytes): listed, 2 small reads, attributes, info and a window in ` +
|
||||||
|
`${server.requests} requests, ${server.bytes} bytes (${(100 * server.bytes / big.size).toFixed(2)}%)`);
|
||||||
|
assert.ok(server.requests <= 8, `big: ${server.requests} requests`);
|
||||||
|
assert.ok(server.bytes * 20 < big.size, `big: ${server.bytes} bytes fetched`);
|
||||||
|
checks += 2;
|
||||||
|
}
|
||||||
|
|
||||||
|
// A cross-origin server that does not expose Content-Range, ETag or
|
||||||
|
// Last-Modified (serve.py's /noexpose/): the length comes from a HEAD
|
||||||
|
// request and answers are checked by their length alone. Every fixture
|
||||||
|
// check, calls in flight at once, and what the page counts.
|
||||||
|
for (const opts of [undefined, { blockSize: 512, cacheSize: 1024 }]) {
|
||||||
|
for (const [name, exp] of Object.entries(expected)) {
|
||||||
|
await fetch(`${base}/__reset`);
|
||||||
|
const f = await pkg.openUrl(`${base}/noexpose/fix/${name}`, opts);
|
||||||
|
await checkFile(`${name} (no exposed headers, ${JSON.stringify(opts ?? {})})`, exp, f);
|
||||||
|
const st = f.stats();
|
||||||
|
eq(st.lazy, true, "no exposed headers: read by ranges");
|
||||||
|
eq(st.size, statSync(join(fixDir, name)).size, "no exposed headers: size from HEAD");
|
||||||
|
const server = await serverStats();
|
||||||
|
eq(server.log.filter((l) => l[4] === "HEAD").length, 1, "no exposed headers: one HEAD");
|
||||||
|
eq(st.requests, server.requests, "no exposed headers: requests counted");
|
||||||
|
eq(st.bytesFetched, server.bytes, "no exposed headers: bytes counted");
|
||||||
|
f.free();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
{
|
||||||
|
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, { blockSize: 512, cacheSize: 0, parallel: 3 });
|
||||||
|
const exp = expected["fixture.h5"];
|
||||||
|
const paths = ["/grid", "/sensors/temp", "/vlen_str", "/cube"];
|
||||||
|
const got = await Promise.all([...paths, ...paths].map((p) => f.read(p)));
|
||||||
|
got.forEach((r, i) => eq(values(exp.datasets[paths[i % 4]].kind, r.data), exp.datasets[paths[i % 4]].values,
|
||||||
|
`no exposed headers: concurrent ${paths[i % 4]}`));
|
||||||
|
await checkErrors(f);
|
||||||
|
}
|
||||||
|
// Without a validator a changed file cannot be told apart; a short
|
||||||
|
// answer still can.
|
||||||
|
await fails(async () => {
|
||||||
|
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
|
||||||
|
blockSize: 512,
|
||||||
|
fetch: async (url, init) => {
|
||||||
|
const r = await fetch(url, init);
|
||||||
|
if (init.method === "HEAD" || init.headers.Range === "bytes=0-511") return r;
|
||||||
|
return new Response((await r.arrayBuffer()).slice(1), { status: 206 });
|
||||||
|
},
|
||||||
|
});
|
||||||
|
await f.read("/grid");
|
||||||
|
}, /got \d+/, "no exposed headers: short answer");
|
||||||
|
// No HEAD length either: a clear error.
|
||||||
|
await fails(() => pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
|
||||||
|
fetch: async (url, init) => (init.method === "HEAD" ? new Response(null, { status: 405 }) : fetch(url, init)),
|
||||||
|
}), /cannot learn the file's size/, "no exposed headers, no HEAD");
|
||||||
|
|
||||||
|
// A server without range support: downloaded whole (the default), or
|
||||||
|
// refused.
|
||||||
|
const whole = await pkg.openUrl(`${base}/norange/fix/fixture.h5`);
|
||||||
|
await checkFile("fixture.h5 (no range support)", expected["fixture.h5"], whole);
|
||||||
|
eq(whole.stats().lazy, false, "downloaded whole");
|
||||||
|
eq(whole.stats().requests, 1, "one request");
|
||||||
|
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fallback: "error" }),
|
||||||
|
/does not support HTTP range requests/, "fallback: error");
|
||||||
|
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000 }),
|
||||||
|
/more than maxDownload/, "maxDownload");
|
||||||
|
// Without a Content-Length the body is streamed, and stopped at the limit.
|
||||||
|
const undeclared = async (url, init) => {
|
||||||
|
const r = await fetch(url, init);
|
||||||
|
return new Response(r.body, { status: r.status });
|
||||||
|
};
|
||||||
|
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000, fetch: undeclared }),
|
||||||
|
/more than maxDownload/, "maxDownload, streamed");
|
||||||
|
const streamed = await pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fetch: undeclared });
|
||||||
|
eq(Array.from((await streamed.read("/sensors/temp")).data), [21.5, 22, 22.25], "streamed download");
|
||||||
|
|
||||||
|
// Errors: HTTP status, not HDF5, bad options, a file that changes, a
|
||||||
|
// server that answers with the wrong bytes.
|
||||||
|
await fails(() => pkg.openUrl(`${base}/fix/missing.h5`), /HTTP 404/, "404");
|
||||||
|
await fails(() => pkg.openUrl(`${base}/fix/expected.json`), /./, "not HDF5");
|
||||||
|
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 100 }), /blockSize/, "blockSize");
|
||||||
|
const tamper = (edit) => async (url, init) => {
|
||||||
|
const r = await fetch(url, init);
|
||||||
|
return init.headers.Range === "bytes=0-511" ? r : edit(r);
|
||||||
|
};
|
||||||
|
const withHeaders = async (r, headers) => {
|
||||||
|
const h = new Headers(r.headers);
|
||||||
|
for (const [k, v] of Object.entries(headers)) h.set(k, v);
|
||||||
|
return new Response(await r.arrayBuffer(), { status: r.status, headers: h });
|
||||||
|
};
|
||||||
|
await fails(async () => {
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, fetch: tamper((r) => withHeaders(r, { ETag: '"other"' })) });
|
||||||
|
await f.read("/grid");
|
||||||
|
}, /changed on the server/, "changed file");
|
||||||
|
await fails(async () => {
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||||
|
blockSize: 512,
|
||||||
|
fetch: tamper(async (r) => new Response((await r.arrayBuffer()).slice(1), { status: 206 })),
|
||||||
|
});
|
||||||
|
await f.read("/grid");
|
||||||
|
}, /got \d+/, "short answer");
|
||||||
|
await fails(async () => {
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||||
|
blockSize: 512,
|
||||||
|
fetch: tamper(async (r) => new Response(await r.arrayBuffer(), { status: 200 })),
|
||||||
|
});
|
||||||
|
await f.read("/grid");
|
||||||
|
}, /stopped honouring range requests/, "200 mid-file");
|
||||||
|
await fails(async () => {
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||||
|
blockSize: 512,
|
||||||
|
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
|
||||||
|
});
|
||||||
|
await f.read("/grid");
|
||||||
|
}, /the server sent 0-511/, "wrong range");
|
||||||
|
await fails(async () => {
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||||
|
blockSize: 512,
|
||||||
|
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes */25752" })),
|
||||||
|
});
|
||||||
|
await f.read("/grid");
|
||||||
|
}, /unusable Content-Range/, "unusable Content-Range");
|
||||||
|
|
||||||
|
await limitTests();
|
||||||
|
await floodTests();
|
||||||
|
|
||||||
|
// Corpus files: what the viewer can show of each is the same read by
|
||||||
|
// ranges as in memory (an error wherever it gives one).
|
||||||
|
const corpus = process.env.CLAWHDF5_WASM_CORPUS;
|
||||||
|
if (corpus) await corpusTests(corpus);
|
||||||
|
console.log(`openUrl: ${checks - before} checks passed`);
|
||||||
|
}
|
||||||
|
|
||||||
|
// A fetch that serves `buf` as a file of `total` bytes (zeros past the end
|
||||||
|
// of `buf`, and `extra` = [[offset, bytes], ...] laid over them), counting
|
||||||
|
// its calls. `hide` leaves out Content-Range and the validators, as a
|
||||||
|
// cross-origin server that exposes neither does; HEAD then gives the length.
|
||||||
|
function mockFetch(buf, { total = buf.length, extra = [], hide = false } = {}) {
|
||||||
|
const f = async (url, init) => {
|
||||||
|
f.calls++;
|
||||||
|
if (init.method === "HEAD") {
|
||||||
|
return new Response(null, { status: 200, headers: { "Content-Length": String(total) } });
|
||||||
|
}
|
||||||
|
const m = /^bytes=(\d+)-(\d+)$/.exec(new Headers(init.headers).get("Range"));
|
||||||
|
const a = Number(m[1]);
|
||||||
|
const b = Math.min(Number(m[2]) + 1, total);
|
||||||
|
const out = new Uint8Array(b - a);
|
||||||
|
for (const [at, bytes] of [[0, buf], ...extra]) {
|
||||||
|
const from = Math.max(a, at);
|
||||||
|
const to = Math.min(b, at + bytes.length);
|
||||||
|
if (from < to) out.set(bytes.subarray(from - at, to - at), from - a);
|
||||||
|
}
|
||||||
|
const headers = { "Content-Length": String(b - a) };
|
||||||
|
if (!hide) headers["Content-Range"] = `bytes ${a}-${b - 1}/${total}`;
|
||||||
|
return new Response(out, { status: 206, headers });
|
||||||
|
};
|
||||||
|
f.calls = 0;
|
||||||
|
return f;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Sizes a hostile server or a large dataset can name are errors, never an
|
||||||
|
// allocation that aborts the module (which would take every open file on
|
||||||
|
// the page with it); see write_limits in make_fixture.py.
|
||||||
|
async function limitTests() {
|
||||||
|
// Read whole, /huge_u8 (2^28 + 1024 bytes) would take over 2 GiB while
|
||||||
|
// decoding: refused before its chunks are fetched; a window reads.
|
||||||
|
const n = 2 ** 28 + 1024;
|
||||||
|
const limits = readFileSync(join(fixDir, "limits.h5"));
|
||||||
|
const local = pkg.open(new Uint8Array(limits));
|
||||||
|
await fails(() => local.read("/huge_u8"), /readHyperslab/, "huge read, in memory");
|
||||||
|
const remote = await pkg.openUrl(`${base}/fix/limits.h5`);
|
||||||
|
const before = remote.stats().requests;
|
||||||
|
await fails(() => remote.read("/huge_u8"), /readHyperslab/, "huge read, openUrl");
|
||||||
|
eq(remote.stats().requests, before, "huge read fetched nothing");
|
||||||
|
eq(Array.from((await remote.readHyperslab("/huge_u8", [n - 4], [4])).data), [0, 0, 0, 7], "huge window");
|
||||||
|
eq(Array.from(local.readHyperslab("/huge_u8", [n - 4], [4]).data), [0, 0, 0, 7], "huge window, in memory");
|
||||||
|
local.free();
|
||||||
|
|
||||||
|
// A server claiming 3 GiB, a heap collection claiming 2 GiB + 4 KiB: an
|
||||||
|
// error after a few requests (it used to fetch 2 GiB, then abort).
|
||||||
|
const hostile = mockFetch(new Uint8Array(readFileSync(join(fixDir, "hostile_vl.h5"))), { total: 3 * 2 ** 30 });
|
||||||
|
const h = await pkg.openUrl("http://hostile.invalid/h.h5", { fetch: hostile });
|
||||||
|
await fails(() => h.read("/a"), /maxFetch/, "hostile collection size");
|
||||||
|
assert.ok(hostile.calls <= 4, `hostile: ${hostile.calls} requests`);
|
||||||
|
// The module survived: files open and read.
|
||||||
|
eq((await (await pkg.openUrl(`${base}/fix/fixture.h5`)).read("/sensors/temp")).data[0], 21.5, "alive after hostile");
|
||||||
|
|
||||||
|
// A call that would fetch more than maxFetch fails before fetching it.
|
||||||
|
await fails(async () => {
|
||||||
|
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, maxFetch: 1024 });
|
||||||
|
await f.read("/grid");
|
||||||
|
}, /maxFetch/, "maxFetch");
|
||||||
|
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxFetch: 2 ** 31 }), /maxFetch/, "maxFetch range");
|
||||||
|
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxDownload: 2 ** 31 }), /maxDownload/, "maxDownload range");
|
||||||
|
|
||||||
|
// Offsets past 2 GiB work on wasm32: far.h5's data sits at 3 GiB.
|
||||||
|
const far = JSON.parse(readFileSync(join(fixDir, "far.json"), "utf8"));
|
||||||
|
const farBytes = new Uint8Array(readFileSync(join(fixDir, "far.h5")));
|
||||||
|
const farData = farBytes.subarray(far.data_at, far.data_at + far.nbytes);
|
||||||
|
const ff = await pkg.openUrl("http://far.invalid/far.h5", {
|
||||||
|
fetch: mockFetch(farBytes, { total: far.length, extra: [[far.far_at, farData]] }),
|
||||||
|
});
|
||||||
|
eq(Array.from((await ff.read("/x")).data), far.values, "data past 2 GiB");
|
||||||
|
eq(ff.stats().size, far.length, "size past 2 GiB");
|
||||||
|
|
||||||
|
// wasm32's reader holds file offsets in 32 bits: a file of 4 GiB or more
|
||||||
|
// is refused at open, with the limit in the message (past 2^53 - 1 bytes
|
||||||
|
// a JavaScript number is not even exact).
|
||||||
|
for (const total of [2 ** 32, 2 ** 40, 2 ** 53 + 2]) {
|
||||||
|
await fails(() => pkg.openUrl("http://huge.invalid/x.h5", { fetch: mockFetch(farBytes, { total }) }),
|
||||||
|
total > 2 ** 53 ? /2\^53 - 1/ : /4 GiB/, `length ${total}`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
|
||||||
|
// tells how many were.
|
||||||
|
function flood(total) {
|
||||||
|
let pulled = 0;
|
||||||
|
const body = new ReadableStream({
|
||||||
|
pull(c) {
|
||||||
|
if (pulled >= total) return c.close();
|
||||||
|
const n = Math.min(65536, total - pulled);
|
||||||
|
pulled += n;
|
||||||
|
c.enqueue(new Uint8Array(n));
|
||||||
|
},
|
||||||
|
}, { highWaterMark: 0 });
|
||||||
|
return { body, pulled: () => pulled };
|
||||||
|
}
|
||||||
|
|
||||||
|
// A 206 whose body runs past the range asked for is cut off as it arrives:
|
||||||
|
// the page never reads (or holds) more than it asked for, whatever the
|
||||||
|
// server sends. 64 MiB stands for "gigabytes".
|
||||||
|
async function floodTests() {
|
||||||
|
const FLOOD = 64 << 20;
|
||||||
|
// The probe: asked for the first 1 MiB.
|
||||||
|
let f = flood(FLOOD);
|
||||||
|
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
|
||||||
|
fetch: async () => new Response(f.body, { status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000" } }),
|
||||||
|
}), /the server sent more/, "flooded probe");
|
||||||
|
assert.ok(f.pulled() <= (1 << 20) + 65536, `probe: read ${f.pulled()} bytes`);
|
||||||
|
// A range request after a correct probe.
|
||||||
|
f = flood(FLOOD);
|
||||||
|
const file = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||||
|
blockSize: 512,
|
||||||
|
fetch: async (url, init) => {
|
||||||
|
const range = new Headers(init.headers).get("Range");
|
||||||
|
if (range === "bytes=0-511") return fetch(url, init);
|
||||||
|
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
|
||||||
|
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
|
||||||
|
},
|
||||||
|
}).catch((e) => e);
|
||||||
|
// The open itself may need a second range: flooded either way.
|
||||||
|
const read = file instanceof Error ? Promise.reject(file) : file.read("/grid");
|
||||||
|
await fails(() => read, /the server sent more/, "flooded range");
|
||||||
|
assert.ok(f.pulled() <= 8 * 512 + 65536, `range: read ${f.pulled()} bytes`);
|
||||||
|
// A declared Content-Length past the range is refused before reading.
|
||||||
|
f = flood(FLOOD);
|
||||||
|
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
|
||||||
|
fetch: async () => new Response(f.body, {
|
||||||
|
status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000", "Content-Length": String(FLOOD) },
|
||||||
|
}),
|
||||||
|
}), /the server sent more/, "declared too long");
|
||||||
|
assert.ok(f.pulled() <= 65536, `declared: read ${f.pulled()} bytes`);
|
||||||
|
}
|
||||||
|
|
||||||
|
function hdf5Files(dir, out) {
|
||||||
|
for (const e of readdirSync(dir, { withFileTypes: true })) {
|
||||||
|
const p = join(dir, e.name);
|
||||||
|
let st;
|
||||||
|
try {
|
||||||
|
st = statSync(p);
|
||||||
|
} catch {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (st.isDirectory()) hdf5Files(p, out);
|
||||||
|
else if (st.size <= 16 << 20) {
|
||||||
|
const b = readFileSync(p);
|
||||||
|
const sig = [0x89, 0x48, 0x44, 0x46, 0x0d, 0x0a, 0x1a, 0x0a];
|
||||||
|
for (let at = 0; at + 8 <= b.length; at = at === 0 ? 512 : at * 2) {
|
||||||
|
if (sig.every((x, i) => b[at + i] === x)) { out.push(p); break; }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
function show(v) {
|
||||||
|
return JSON.stringify(v, (_, x) => {
|
||||||
|
if (typeof x === "bigint") return `${x}n`;
|
||||||
|
if (ArrayBuffer.isView(x)) return Array.from(x, (y) => (typeof y === "bigint" ? `${y}n` : Number.isNaN(y) ? "NaN" : y));
|
||||||
|
return x;
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async function transcript(f) {
|
||||||
|
const out = [];
|
||||||
|
const call = async (what, fn) => {
|
||||||
|
try {
|
||||||
|
out.push(`${what}: ${show(await fn())}`);
|
||||||
|
} catch {
|
||||||
|
out.push(`${what}: Err`);
|
||||||
|
}
|
||||||
|
};
|
||||||
|
const todo = [["/", 0]];
|
||||||
|
while (todo.length && out.length < 3000) {
|
||||||
|
const [path, depth] = todo.pop();
|
||||||
|
let kind = null;
|
||||||
|
await call(`${path} kind`, async () => (kind = await f.kind(path)));
|
||||||
|
await call(`${path} attrs`, () => f.attrs(path));
|
||||||
|
if (kind === "group") {
|
||||||
|
let list = [];
|
||||||
|
await call(`${path} list`, async () => (list = await f.list(path)));
|
||||||
|
if (depth < 12) for (const c of list.reverse()) todo.push([lib.joinPath(path, c.name), depth + 1]);
|
||||||
|
} else if (kind === "dataset") {
|
||||||
|
let info = null;
|
||||||
|
await call(`${path} info`, async () => (info = await f.info(path)));
|
||||||
|
if (info && [...info.shape, ...info.elementShape].reduce((a, b) => a * b, 1) <= 1 << 20) {
|
||||||
|
await call(`${path} read`, () => f.read(path));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function corpusTests(corpus) {
|
||||||
|
const files = hdf5Files(corpus, []).sort();
|
||||||
|
let opened = 0;
|
||||||
|
let bytes = 0;
|
||||||
|
let size = 0;
|
||||||
|
for (const p of files) {
|
||||||
|
const rel = relative(corpus, p).split("/").map(encodeURIComponent).join("/");
|
||||||
|
let local;
|
||||||
|
try {
|
||||||
|
local = pkg.open(new Uint8Array(readFileSync(p)));
|
||||||
|
} catch {
|
||||||
|
await fails(() => pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 }), /./, `${rel}: opens in neither`);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
const remote = await pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 });
|
||||||
|
const want = await transcript(local);
|
||||||
|
const got = await transcript(remote);
|
||||||
|
// Errors are compared as errors: a malformed file can fail at another
|
||||||
|
// check, with another message, when read by ranges.
|
||||||
|
eq(got, want, `${rel}: transcript`);
|
||||||
|
opened++;
|
||||||
|
bytes += remote.stats().bytesFetched;
|
||||||
|
size += remote.stats().size;
|
||||||
|
local.free();
|
||||||
|
remote.free();
|
||||||
|
}
|
||||||
|
console.log(`corpus: ${opened} of ${files.length} files agree over HTTP (${bytes} of ${size} bytes fetched)`);
|
||||||
|
}
|
||||||
|
|||||||
@@ -75,3 +75,24 @@ export function toRows(data, rows, cols, per = 1) {
|
|||||||
export function perElement(elementShape) {
|
export function perElement(elementShape) {
|
||||||
return elementShape.reduce((a, b) => a * b, 1);
|
return elementShape.reduce((a, b) => a * b, 1);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** A byte count for people: `1.5 KiB`, `200 MiB`. */
|
||||||
|
export function formatBytes(n) {
|
||||||
|
const units = ["B", "KiB", "MiB", "GiB", "TiB"];
|
||||||
|
let i = 0;
|
||||||
|
let v = n;
|
||||||
|
while (v >= 1024 && i < units.length - 1) {
|
||||||
|
v /= 1024;
|
||||||
|
i++;
|
||||||
|
}
|
||||||
|
const s = i === 0 ? String(v) : v < 10 ? String(Number(v.toFixed(1))) : String(Math.round(v));
|
||||||
|
return `${s} ${units[i]}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** What reading a remote file has cost, from `RemoteFile.stats()`. */
|
||||||
|
export function formatStats({ lazy, requests, bytesFetched, size }) {
|
||||||
|
if (!lazy) return `downloaded whole (${formatBytes(size)}): the server does not support range requests`;
|
||||||
|
const pct = size > 0 ? (100 * bytesFetched) / size : 0;
|
||||||
|
return `${requests} request${requests === 1 ? "" : "s"}, ${formatBytes(bytesFetched)} of ` +
|
||||||
|
`${formatBytes(size)} fetched (${pct.toFixed(1)}%)`;
|
||||||
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user