Merge branch 'feat/p3-m4-wasm-lazy' into feat/p3-wasm-swmr-python

# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
This commit is contained in:
osobh
2026-09-27 07:59:14 -05:00
24 changed files with 3752 additions and 293 deletions
+81
View File
@@ -214,6 +214,87 @@ Design: `docs/design/swmr.md`.
that hangs up, 16 threads on one remote file, and a thread that keeps
running while a read waits on 0.2 s requests.
### Range reads, milestone M4: remote files in the browser (2026-09-27)
- **`openUrl(url, opts)`** in `clawhdf5-wasm` opens an HDF5/NetCDF-4 file
on a web server without downloading it: the returned `RemoteFile` has
`H5File`'s methods (`kind`, `list`, `info`, `attrs`, `attrErrors`,
`read`, `readHyperslab`), each returning a promise, and `stats()`
(requests, bytes fetched, file size). Only the byte ranges a call needs
are fetched, with `fetch` and `Range` headers, in 1 MiB blocks
(`blockSize`) kept in a 64 MiB cache (`cacheSize`). `open(bytes)` is
unchanged.
- **How, on the browser's main thread:** the design's restartable
"NeedBytes" mode (`clawhdf5_wasm::lazy::LazyStorage`). A call runs as a
pass over the blocks fetched so far; a read that misses records the
missing blocks and fails, the pass's result is dropped (even if a parser
caught the error and carried on), the blocks are fetched and the pass is
re-run. No Web Worker and no synchronous XHR (what h5wasm's lazy files
need). No block is evicted while a call is in flight, so every call ends;
the cache is trimmed between calls, raw-data blocks before metadata.
`clawhdf5-remote`'s `BlockCache` is not reused: it fetches by blocking,
and it evicts, and does not keep large reads, during a read, where a
restartable pass needs every block it read to stay until it finishes.
- **Every answer is checked** (`js/remote.js`): a `206` with exactly the
bytes asked for (`Content-Range`, when the page can see it, and the body
length), and the same `ETag`/`Last-Modified` and length as at open, or
the call fails: never data from another file or offset. A server that
answers `200` to a `Range` request is downloaded whole (up to
`maxDownload`, 512 MiB) unless `fallback: "error"`. Other options:
`headers`, `credentials`, `parallel` (6 requests at a time), `fetch`.
- **The viewer** (`examples/wasm-viewer`) has a URL box, opens
`?file=<url>` lazily, and shows the requests and bytes fetched.
- Counted on tank, 2026-09-27 (`WASM_BIG_MB=200 bash
examples/wasm-viewer/test/run.sh`, requests as `test/serve.py` counted
them): in a 200 MB h5py file, listing the root, reading two small
datasets, a group's attributes, the large dataset's shape and a
10-value window of it (25 million values) took 5 requests and 6 MiB.
With
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, the 622 corpus files
up to 16 MiB that open read over HTTP as from bytes: the same listings,
attributes and values (datasets up to 2^20 values), and an error
wherever bytes give one. Natively, `cargo test -p clawhdf5-wasm --test
lazy` with the same variable compares 656 files (up to 64 MiB, error
messages included, against the facade's range-storage path).
- `Reader::open_storage` (the wasm crate's core over any `Storage`), and
variable-length strings resolve through the file's storage rather than
`File::as_bytes`.
- The package is larger: the facade's `Storage` read path is now
reachable from JavaScript (it was compiled out before), and the promise
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
build machine was shared); the viewer README's size table predates M4.
- **Hardened after review (2026-09-27):**
- Sizes a server or a dataset names are errors, never an abort of the
wasm module (which took every open file on the page with it): a read
past 2 GiB aborted in the lazy cache, reachable by a hostile server
claiming a large file and a 2 GiB heap collection, and `read()` of a
256 MiB `u8` dataset aborted widening it to 64 bits. New option
`maxFetch` (512 MiB, at most 1 GiB): what one call may fetch, and the
longest single read, refused before fetching. `read()` refuses a
dataset that would take more than 1 GiB to decode, naming
`readHyperslab`. A file of 4 GiB or more is refused at open (wasm32
reads offsets as 32-bit); `maxDownload` is at most 1 GiB.
- Response bodies are read as they arrive and cut off at the length
asked for (`maxDownload` for a `200`): a `206` with a gigabyte body
was buffered whole before its length was checked.
- Listing a group reads every child's header, and every node of each
level of the group's index, in one pass: 3000 datasets (h5py, 198 MB)
listed in 6 passes and 73 requests at 1 MiB blocks instead of 185
passes and 184 serial requests (`libver="latest"`: 9 passes instead of
189). In `clawhdf5-format`, the B-tree v1/v2 collectors, the symbol
table node loop and the dense-link loop read (without using) the
siblings after the first that fails, then return that error: same
results and errors, more reads only on failure (free in memory). The
lazy cache no longer re-fetches a cached block to merge two requests.
- `headers` may be a `Headers` instance or `[name, value]` pairs (a
`Headers` was silently dropped); when one range request fails the
others in flight are aborted; `parallel` must be an integer from 1 to
1024.
- Tests: `test/serve.py` serves ranges without exposed `Content-Range`/
`ETag` (`/noexpose/`, and `/unexposed/` for a real cross-origin page
in Chromium), so the HEAD-length path runs end to end; hostile and
oversized files (`make_fixture.py`'s `write_limits`), flooding
bodies, aborted siblings.
### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
+13 -8
View File
@@ -181,14 +181,19 @@ Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal F
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
file over HTTP with `File::open`.
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only, file held in memory;
no Zstd/SZIP since they link C) and the `examples/wasm-viewer/` page.
`examples/wasm-viewer/test/run.sh` builds the package (needs the
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node
and headless Chromium (a Playwright download in `~/.cache/ms-playwright`
on tank); the CI container has neither, so CI runs the native
`clawhdf5-wasm` `h5py_interop` test on the same fixture. Size numbers are
in the example's README.
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
call is re-run after each wave of misses; no block evicted while a call
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
version) and tests it under Node and headless Chromium (a Playwright
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
(range server with request counts, 200 MB budget file); the CI container
has neither, so CI runs the native `h5py_interop` and `lazy` tests
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
numbers in the example's README predate `openUrl`.
- Python and Node.js bindings for cross-language use
- NetCDF-4 compatibility for scientific data interop
+1 -1
View File
@@ -827,7 +827,7 @@ clawhdf5 workspace (19 crates, ~86K lines of Rust in src/, ~104K with tests
├── Bindings
│ ├── clawhdf5-py — Python (PyO3)
│ ├── clawhdf5-napi — Node.js (napi-rs)
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only)
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only; remote files by HTTP range requests)
│
└── Tooling
├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check
+18 -5
View File
@@ -199,19 +199,32 @@ fn collect_symbol_table_nodes_inner<S: Storage + ?Sized>(
// Leaf: children are SNOD addresses
Ok(node.children)
} else {
// Internal: recurse into children
// Internal: recurse into children. After the first child that
// fails, the others are only read (as `storage::touch` does), not
// descended into; that error is returned.
let mut result = Vec::new();
let mut failed = None;
for &child_addr in &node.children {
let child_snods = collect_symbol_table_nodes_inner(
if failed.is_some() {
// Parsing reads the node's header, then its body.
let _ = BTreeV1Node::parse_in(file, child_addr, offset_size, length_size);
continue;
}
match collect_symbol_table_nodes_inner(
file,
child_addr,
offset_size,
length_size,
depth + 1,
)?;
result.extend(child_snods);
) {
Ok(child_snods) => result.extend(child_snods),
Err(e) => failed = Some(e),
}
}
match failed {
Some(e) => Err(e),
None => Ok(result),
}
Ok(result)
}
}
+51 -33
View File
@@ -504,45 +504,63 @@ fn collect_internal_records<S: Storage + ?Sized>(
// Interleave: child[0], record[0], child[1], record[1], ..., child[nr]
// We collect child[0] records, then record[0], then child[1], etc.
// After the first child that fails, the others are only touched (see
// `storage::touch`); that error is returned.
let mut failed = None;
for (i, &(child_addr, child_nrec)) in node.children.iter().enumerate() {
if child_depth == 0 {
// Before parsing, so a refused tree is not also a large allocation.
spend(budget, usize::from(child_nrec))?;
let leaf_recs = parse_leaf_records(
file,
to_usize(child_addr)?,
child_nrec,
record_size,
node_size,
)?;
out.extend(leaf_recs);
} else {
collect_internal_records(
file,
to_usize(child_addr)?,
child_nrec,
child_depth,
record_size,
node_size,
offset_size,
length_size,
max_leaf_nrec,
budget,
out,
)?;
if failed.is_some() {
let len = usize::try_from(node_size)
.unwrap_or(usize::MAX)
.min(1 << 16);
crate::storage::touch(file, child_addr, len);
continue;
}
if let Err(e) = (|| -> Result<(), FormatError> {
if child_depth == 0 {
// Before parsing, so a refused tree is not also a large allocation.
spend(budget, usize::from(child_nrec))?;
let leaf_recs = parse_leaf_records(
file,
to_usize(child_addr)?,
child_nrec,
record_size,
node_size,
)?;
out.extend(leaf_recs);
} else {
collect_internal_records(
file,
to_usize(child_addr)?,
child_nrec,
child_depth,
record_size,
node_size,
offset_size,
length_size,
max_leaf_nrec,
budget,
out,
)?;
}
// Add record[i] (except after the last child)
if i < nr {
let data = node.record(i, rs)?;
spend(budget, 1)?;
out.push(BTreeV2Record {
data: data.to_vec(),
});
// Add record[i] (except after the last child)
if i < nr {
let data = node.record(i, rs)?;
spend(budget, 1)?;
out.push(BTreeV2Record {
data: data.to_vec(),
});
}
Ok(())
})() {
failed = Some(e);
}
}
Ok(())
match failed {
Some(e) => Err(e),
None => Ok(()),
}
}
/// The records of a B-tree v2 that fall in one key range, found by
+38 -14
View File
@@ -79,27 +79,51 @@ pub(crate) fn v1_group_entries<S: Storage + ?Sized>(
length_size,
)?;
// The names are read one by one from the heap's data segment; read
// (up to 1 MiB of) it first, so a storage that records what it lacks
// asks for it at once (see `storage::touch`).
if !snod_addrs.is_empty() {
let len = usize::try_from(heap.data_segment_size).map_or(1 << 20, |n| n.min(1 << 20));
crate::storage::touch(file_data, heap.data_segment_address, len);
}
let mut entries = Vec::new();
let mut heap_checked = false;
// After the first node that fails, the others are only read (as
// `storage::touch` does); that error is returned.
let mut failed = None;
for snod_addr in snod_addrs {
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
for entry in &snod.entries {
// Like libhdf5, look at the heap's free list only once a name is
// needed: an empty group with a damaged heap still lists.
if !heap_checked {
heap.validate_free_list_in(file_data, length_size)?;
heap_checked = true;
if failed.is_some() {
let _ = SymbolTableNode::parse_in(file_data, snod_addr, offset_size);
continue;
}
let mut node = || -> Result<(), FormatError> {
let snod = SymbolTableNode::parse_in(file_data, checked_addr(snod_addr)?, offset_size)?;
for entry in &snod.entries {
// Like libhdf5, look at the heap's free list only once a name
// is needed: an empty group with a damaged heap still lists.
if !heap_checked {
heap.validate_free_list_in(file_data, length_size)?;
heap_checked = true;
}
let name = heap.read_string_in(file_data, entry.link_name_offset)?;
entries.push(GroupEntry {
name,
object_header_address: entry.object_header_address,
cache_type: entry.cache_type,
});
}
let name = heap.read_string_in(file_data, entry.link_name_offset)?;
entries.push(GroupEntry {
name,
object_header_address: entry.object_header_address,
cache_type: entry.cache_type,
});
Ok(())
};
if let Err(e) = node() {
failed = Some(e);
}
}
Ok(entries)
match failed {
Some(e) => Err(e),
None => Ok(entries),
}
}
/// Symbol table cache type for a soft link: the scratch pad's first four bytes
+15 -4
View File
@@ -126,6 +126,9 @@ fn for_each_dense_link<S: Storage + ?Sized>(
)?;
let records = collect_btree_v2_records_in(file_data, &btree_hdr, offset_size, length_size)?;
// After the first link that fails, the others are only read, not
// visited (a touch, see `storage::touch`); that error is returned.
let mut failed = None;
for record in &records {
// For type 5 (name index): hash(4) + heap_id(heap_id_length)
// For type 6 (creation order): creation_order(8) + heap_id(heap_id_length)
@@ -141,12 +144,20 @@ fn for_each_dense_link<S: Storage + ?Sized>(
let id_bytes = &record.data[id_offset..id_offset + fh.heap_id_length as usize];
// Read managed object from fractal heap
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size)?;
if let Some(link) = parse_link(&link_data, offset_size)? {
visit(link);
let link_data = fh.read_managed_object_in(file_data, id_bytes, offset_size);
if failed.is_some() {
continue;
}
match link_data.and_then(|d| parse_link(&d, offset_size)) {
Ok(Some(link)) => visit(link),
Ok(None) => {}
Err(e) => failed = Some(e),
}
}
Ok(())
match failed {
Some(e) => Err(e),
None => Ok(()),
}
}
/// Resolve entries from dense storage (fractal heap + B-tree v2).
+13
View File
@@ -202,6 +202,19 @@ pub(crate) fn len_usize<S: Storage + ?Sized>(file: &S) -> usize {
usize::try_from(file.len()).unwrap_or(usize::MAX)
}
/// Read `len` bytes at `offset` and drop them, ignoring any error.
///
/// For a traversal that has failed on one sibling (a B-tree child, a
/// symbol table node, a heap object) and would stop there: it first
/// touches the siblings it did not get to, so a storage that records what
/// it lacks — the browser's restartable reader, which fetches over the
/// network between attempts — learns about all of them in one attempt
/// instead of one per attempt. Results and errors are unchanged (the first
/// error is still the one returned); an in-memory read is free.
pub fn touch<S: Storage + ?Sized>(file: &S, offset: u64, len: usize) {
let _ = file.read_at(offset, len);
}
/// Bytes `[offset, offset + len)`, all of them.
///
/// A range that runs past the end of the storage is
+2
View File
@@ -24,6 +24,8 @@ clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" }
# Must match the wasm-bindgen CLI exactly; build.sh checks.
wasm-bindgen = "0.2.129"
js-sys = "0.3.106"
# Promises for openUrl and RemoteFile (pure Rust over js-sys).
wasm-bindgen-futures = "0.4.79"
[dev-dependencies]
serde_json = "1"
+235
View File
@@ -0,0 +1,235 @@
// HTTP for clawhdf5-wasm's openUrl (see src/lib.rs and src/lazy.rs).
//
// The Rust side decides which byte ranges a read needs; this file fetches
// them with `fetch` and `Range` headers and checks every answer, so a server
// that ignores the range, answers with other bytes, or serves a file that
// changed since it was opened is an error, never data. wasm-bindgen copies
// it into the package (pkg/snippets/...).
const DEFAULT_MAX_DOWNLOAD = 512 * 1024 * 1024;
const DEFAULT_PARALLEL = 6;
function fetcher(opts) {
const f = opts?.fetch ?? globalThis.fetch;
if (typeof f !== "function") {
throw new Error("openUrl: no fetch() in this environment (pass opts.fetch)");
}
return f;
}
// The caller's headers (`opts.headers`: a Headers, [name, value] pairs or a
// plain object, as fetch takes them; names come out in lower case) with
// `extra` over them. A Range of the caller's is dropped: this file asks for
// the ranges.
function init(opts, extra, method = "GET", signal = undefined) {
const headers = {};
if (opts?.headers != null) {
for (const [name, value] of new Headers(opts.headers)) {
if (name !== "range") headers[name] = value;
}
}
Object.assign(headers, extra);
return { method, headers, credentials: opts?.credentials, signal };
}
// `opts.parallel`: range requests in flight at once.
function parallelism(opts) {
const p = opts?.parallel ?? DEFAULT_PARALLEL;
if (!Number.isSafeInteger(p) || p < 1) {
throw new Error(`openUrl: parallel must be a positive integer, got ${String(p)}`);
}
return p;
}
// "bytes a-b/total" -> { start, end (exclusive), total | null }; null when
// the page cannot see the header (cross-origin, not exposed).
function contentRange(resp, url) {
const v = resp.headers.get("Content-Range");
if (v === null) return null;
const m = /^bytes (\d+)-(\d+)\/(\d+|\*)$/.exec(v.trim());
if (!m) throw new Error(`${url}: the server sent an unusable Content-Range: ${v}`);
return { start: Number(m[1]), end: Number(m[2]) + 1, total: m[3] === "*" ? null : Number(m[3]) };
}
// What pins the file: its ETag, else its Last-Modified (null if neither is
// visible to this page).
function validatorOf(resp) {
return resp.headers.get("ETag") ?? resp.headers.get("Last-Modified");
}
async function discard(resp) {
try {
await resp.body?.cancel();
} catch {
// Nothing to release.
}
}
// The body, refusing more than `limit` bytes as they arrive: it is piped
// through a TransformStream that errors the moment the count passes the
// limit, which cancels the body and so aborts the request. Whatever the
// server declares or sends, the page never holds more than `limit` bytes of
// it. (`tooBig(n)` makes the error; n is the count so far.) A reader loop
// would do the same, but it stalls on small bodies in headless Chromium
// under --virtual-time-budget, which the page test uses; a pipe does not.
async function readCapped(resp, limit, tooBig) {
const declared = resp.headers.get("Content-Length");
if (declared !== null && Number(declared) > limit) {
await discard(resp);
throw tooBig(declared);
}
if (!resp.body) {
// No stream to read from (some fetch implementations): all at once.
const all = new Uint8Array(await resp.arrayBuffer());
if (all.length > limit) throw tooBig(all.length);
return all;
}
let n = 0;
let over = null;
const capped = resp.body.pipeThrough(new TransformStream({
transform(chunk, ctl) {
n += chunk.length;
if (n > limit) {
over = tooBig(`over ${limit}`);
ctl.error(over);
return;
}
ctl.enqueue(chunk);
},
}));
try {
return new Uint8Array(await new Response(capped).arrayBuffer());
} catch (e) {
throw over ?? e;
}
}
// The whole body of a 200 answer (a server without range support), at
// most `limit` (maxDownload) bytes.
function readAll(resp, limit, url) {
return readCapped(resp, limit, (n) =>
new Error(`${url} is ${n} bytes, more than maxDownload (${limit}); ` +
"the server does not support range requests, so the whole file would have to be downloaded"));
}
// The body of a 206 answer, which may not be longer than the `limit` bytes
// asked for at `start` (the caller checks the exact length).
function readLimited(resp, limit, url, start) {
return readCapped(resp, limit, () =>
new Error(`${url}: asked for ${limit} bytes at offset ${start}, the server sent more`));
}
/**
* Ask for the file's first `firstLen` bytes. A server that honours the
* range (206) gives `{ length, first, validator, requests }`; one that
* answers 200 sends the whole file, which is kept (`{ whole, requests }`)
* when `opts.fallback` is "download" (the default) and the file is at most
* `opts.maxDownload` bytes, and is an error otherwise.
*/
export async function probe(url, firstLen, opts) {
const f = fetcher(opts);
parallelism(opts);
const resp = await f(url, init(opts, { Range: `bytes=0-${firstLen - 1}` }));
if (resp.status === 206) {
const cr = contentRange(resp, url);
if (cr && cr.start !== 0) {
await discard(resp);
throw new Error(`${url}: asked for bytes from 0, the server sent bytes from ${cr.start}`);
}
const first = await readLimited(resp, firstLen, url, 0);
let length = cr?.total ?? null;
let requests = 1;
if (length === null) {
// Content-Range is not readable here: a cross-origin server that does
// not list it in Access-Control-Expose-Headers. Content-Length of a
// HEAD request is always readable.
const head = await f(url, init(opts, {}, "HEAD"));
requests++;
const cl = head.headers.get("Content-Length");
if (!head.ok || cl === null) {
throw new Error(`${url}: cannot learn the file's size (a cross-origin server must send ` +
"Access-Control-Expose-Headers: Content-Range, or answer HEAD with Content-Length)");
}
length = Number(cl);
}
if (!Number.isSafeInteger(length) || length < 0) {
throw new Error(`${url}: the server gave a file size of ${length} bytes; openUrl reads files ` +
"of up to 2^53 - 1 bytes (the largest offset a JavaScript number holds exactly)");
}
if (first.length !== Math.min(firstLen, length)) {
throw new Error(`${url}: asked for the first ${firstLen} bytes of ${length}, got ${first.length}`);
}
return { length, first, validator: validatorOf(resp), requests };
}
if (resp.status === 200) {
if ((opts?.fallback ?? "download") !== "download") {
await discard(resp);
throw new Error(`${url}: the server does not support HTTP range requests (it answered 200 ` +
"to a Range request); open it with { fallback: \"download\" } to download the whole file");
}
const whole = await readAll(resp, opts?.maxDownload ?? DEFAULT_MAX_DOWNLOAD, url);
return { whole, requests: 1 };
}
await discard(resp);
throw new Error(`${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
}
/**
* Fetch `ranges` ([start0, end0, start1, end1, ...], ends exclusive) of a
* file opened by `probe`, at most `opts.parallel` (default 6) at a time.
* Every answer must be a 206 with exactly the bytes asked for, from the same
* file (validator and length). When one request fails, the others in
* flight are aborted and no more are made; that failure is the error.
*/
export async function fetchRanges(url, ranges, opts, validator, length) {
const f = fetcher(opts);
const parallel = parallelism(opts);
const n = ranges.length / 2;
const out = new Array(n);
const abort = new AbortController();
let next = 0;
async function one(i) {
const start = ranges[2 * i];
const end = ranges[2 * i + 1];
const resp = await f(url, init(opts, { Range: `bytes=${start}-${end - 1}` }, "GET", abort.signal));
if (resp.status !== 206) {
await discard(resp);
throw new Error(resp.status === 200
? `${url}: the server stopped honouring range requests`
: `${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim());
}
const cr = contentRange(resp, url);
const v = validatorOf(resp);
if ((validator != null && v !== null && v !== validator) ||
(cr?.total != null && cr.total !== length)) {
await discard(resp);
throw new Error(`${url} changed on the server since it was opened`);
}
if (cr && (cr.start !== start || cr.end !== end)) {
await discard(resp);
throw new Error(`${url}: asked for bytes ${start}-${end - 1}, the server sent ${cr.start}-${cr.end - 1}`);
}
const body = await readLimited(resp, end - start, url, start);
if (body.length !== end - start) {
throw new Error(`${url}: asked for ${end - start} bytes at offset ${start}, got ${body.length}`);
}
out[i] = body;
}
async function worker() {
while (next < n && !abort.signal.aborted) {
try {
await one(next++);
} catch (e) {
// The first failure stops the rest: requests in flight are aborted
// (their AbortErrors are not reported) and no new ones start.
if (!abort.signal.aborted) {
abort.abort();
throw e;
}
return;
}
}
}
await Promise.all(Array.from({ length: Math.min(parallel, n) }, worker));
return out;
}
+113 -28
View File
@@ -6,11 +6,25 @@
//! with no typed-array mapping (compound, reference, opaque, ...) is refused
//! with a message naming it, never returned as reinterpreted bytes.
use std::sync::Arc;
use clawhdf5::{AttrValue, File, Selection};
use clawhdf5_format::data_read;
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::storage::Storage;
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
/// The most memory one read may use while it decodes: the stored bytes,
/// the values at 64 bits (integers are widened first) and the values
/// returned. A larger read fails with an error naming `readHyperslab`,
/// before anything is read: on wasm32 a buffer past 2 GiB cannot be
/// allocated at all, and failing to allocate aborts the module (every open
/// file on the page with it). 1 GiB leaves room in wasm32's 4 GiB for the
/// file's cached blocks and the JavaScript copy of the result.
pub const MAX_READ_BYTES: u64 = 1 << 30;
/// Errors are reported to JavaScript as messages.
pub type Result<T> = std::result::Result<T, String>;
@@ -119,7 +133,9 @@ pub struct Attr {
pub value: AttrValue,
}
/// An open file, held in memory.
/// An open file: held in memory ([`Reader::open`]) or read through a
/// [`Storage`] ([`Reader::open_storage`], such as a
/// [`LazyStorage`](crate::lazy::LazyStorage)).
pub struct Reader {
file: File,
}
@@ -132,6 +148,14 @@ impl Reader {
})
}
/// Open a file read through `storage` (the file's bytes from offset 0,
/// user block included, as [`File::open_storage`] takes them).
pub fn open_storage(storage: Arc<dyn Storage + Send + Sync>) -> Result<Self> {
Ok(Self {
file: File::open_storage(storage).map_err(err)?,
})
}
/// Whether `path` names a group or a dataset.
pub fn kind(&self, path: &str) -> Result<Kind> {
match self.file.dataset(path) {
@@ -144,31 +168,54 @@ impl Reader {
/// The groups, then the datasets, in the group at `path` (`/` is the
/// root). Soft links are listed as their targets; external and dangling
/// links, and named datatypes, are left out.
///
/// What [`Group::groups`](clawhdf5::Group::groups) and `datasets` list,
/// but every child's object header is read before an error ends the
/// listing (the first error, in listing order, is the one returned, as
/// there). Over a [`LazyStorage`](crate::lazy::LazyStorage) that makes
/// one pass ask for all the headers it is missing at once, instead of
/// one pass, and one round trip, per header.
pub fn list(&self, path: &str) -> Result<Vec<Child>> {
if self.kind(path)? != Kind::Group {
return Err(format!("not a group: {path}"));
}
let group = self.file.group(path).map_err(err)?;
let mut out: Vec<Child> = group
.groups()
.map_err(err)?
.into_iter()
.map(|name| Child {
name,
kind: Kind::Group,
})
.collect();
out.extend(
group
.datasets()
.map_err(err)?
.into_iter()
.map(|name| Child {
name,
kind: Kind::Dataset,
}),
);
Ok(out)
let entries = group.entries().map_err(err)?;
let sb = self.file.superblock();
let storage = self.file.storage();
let mut groups = Vec::new();
let mut datasets = Vec::new();
let mut first_error = None;
for (name, address) in entries {
match ObjectHeader::parse_in(storage, address, sb.offset_size, sb.length_size) {
Ok(header) => {
let has = |t: MessageType| header.messages.iter().any(|m| m.msg_type == t);
if has(MessageType::LinkInfo)
|| has(MessageType::Link)
|| has(MessageType::SymbolTable)
{
groups.push(Child {
name: name.clone(),
kind: Kind::Group,
});
}
if has(MessageType::DataLayout) {
datasets.push(Child {
name,
kind: Kind::Dataset,
});
}
}
Err(e) => {
first_error.get_or_insert(e);
}
}
}
if let Some(e) = first_error {
return Err(err(clawhdf5::Error::from(e)));
}
groups.extend(datasets);
Ok(groups)
}
/// Shape, max shape and datatype of the dataset at `path`.
@@ -228,13 +275,27 @@ impl Reader {
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
}
let raw = ds.read_selection(&selection).map_err(err)?;
let data = self.decode(&raw, &dt)?;
out_shape.extend(element_shape(&dt));
let expected = out_shape
.iter()
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
.ok_or("selection size overflows")?;
let cost = expected.saturating_mul(bytes_per_value(&dt));
if cost > MAX_READ_BYTES {
return Err(format!(
"reading {path}{} would take about {} MiB of memory, more than the {} MiB \
one read may use; read it in parts (readHyperslab)",
if slab.is_some() {
" (this selection)"
} else {
" whole"
},
cost >> 20,
MAX_READ_BYTES >> 20
));
}
let raw = ds.read_selection(&selection).map_err(err)?;
let data = self.decode(&raw, &dt)?;
if data.len() as u64 != expected {
return Err(format!(
"read {} values for shape {out_shape:?} ({expected} expected)",
@@ -281,11 +342,14 @@ impl Reader {
// string ends at its first NUL and a heap object of the
// wrong size is an error, as in libhdf5 and h5py.
let sb = self.file.superblock();
Data::Strings(
VlResolver::new(self.file.as_bytes(), sb.offset_size, sb.length_size)
.strings(raw)
.map_err(err)?,
)
let strings = match self.file.contiguous_bytes() {
Some(bytes) => {
VlResolver::new(bytes, sb.offset_size, sb.length_size).strings(raw)
}
None => VlResolver::new_in(self.file.storage(), sb.offset_size, sb.length_size)
.strings(raw),
};
Data::Strings(strings.map_err(err)?)
}
Datatype::Enumeration { .. } if !is_array => {
Data::Strings(data_read::read_enum_names(raw, dt).map_err(err)?)
@@ -300,6 +364,27 @@ impl Reader {
}
}
/// Memory one value of type `dt` takes while [`Reader::read`] decodes it
/// (an array type's elements count as values): its stored bytes, plus what
/// [`Reader::decode`] builds from them. A string counts its `String` (24
/// bytes on 64-bit targets, less on wasm32) and, for a fixed-length one,
/// its text; a variable-length string's text lives in the heap and is
/// bounded by the storage's own read limit.
fn bytes_per_value(dt: &Datatype) -> u64 {
let base = array_base(dt);
let stored = u64::from(base.type_size());
stored
+ match base {
Datatype::FloatingPoint { size, .. } if *size <= 4 => 4,
Datatype::FloatingPoint { .. } => 8,
// Widened to 64 bits, then narrowed to a new vector.
Datatype::FixedPoint { .. } => 8 + stored,
Datatype::String { .. } => 24 + stored,
Datatype::VariableLength { .. } | Datatype::Enumeration { .. } => 24,
_ => 0,
}
}
/// Narrow integers read at 64 bits to the dataset's own width. The source is
/// that width, so this cannot fail on correct input; it is checked anyway.
fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T>> {
+815
View File
@@ -0,0 +1,815 @@
//! Reading a file that is not all here, when no read may wait for the
//! network: the restartable "NeedBytes" mode of `docs/design/range-reads.md`
//! (milestone M4).
//!
//! A browser's main thread cannot block on `fetch`, and the parsers are
//! synchronous. So an operation (open, list a group, read a dataset) runs
//! as a *pass* over a [`LazyStorage`] that holds the blocks fetched so far:
//!
//! 1. [`LazyStorage::attempt`] runs the operation. A read whose blocks are
//! all present is served; a read that misses records the missing blocks
//! and fails with a storage error.
//! 2. If the pass missed anything, its result is thrown away — whatever it
//! is, since a parser may have caught the error and carried on (a
//! listing skips a link it cannot resolve) — and the caller gets the
//! byte ranges to fetch ([`Step::Need`]).
//! 3. The caller fetches them (asynchronously, with HTTP `Range` requests),
//! hands them over with [`LazyStorage::supply`] and runs the operation
//! again.
//!
//! A pass is pure over the storage: the facade only caches what completed
//! reads decoded (its chunk cache), so re-running it is safe. Every pass
//! that does not finish asks for at least one block not yet present, and no
//! block is evicted while an operation is in flight
//! ([`LazyStorage::operation`]), so an operation finishes after at most one
//! pass per block it needs. In practice it is one pass per *wave* of
//! misses: a chunked read asks for all the chunks of a batch at once.
//!
//! Blocks are kept in an LRU cache with a byte budget, trimmed only when no
//! operation is in flight. Blocks fetched for bulk reads (raw data: a
//! `read_ranges` call, or a read longer than a block) go first, so reading
//! a large dataset does not evict the metadata.
use std::borrow::Cow;
use std::collections::{BTreeSet, HashMap};
use std::ops::Range;
use std::sync::{Arc, Mutex, MutexGuard};
use clawhdf5_format::error::FormatError;
use clawhdf5_format::storage::Storage;
/// Default block size: 1 MiB, as `clawhdf5-remote`'s block cache (the size
/// `docs/design/range-reads.md` §2 measured).
pub const DEFAULT_BLOCK_SIZE: u64 = 1 << 20;
/// Default of [`LazyConfig::max_fetch`]: 512 MiB, the same as `openUrl`'s
/// `maxDownload` for a server without range support.
pub const DEFAULT_MAX_FETCH: u64 = 512 << 20;
/// The message of the error a read that misses returns. It never reaches
/// the caller of [`LazyStorage::attempt`]: a pass that missed is re-run.
pub const NEED_BYTES: &str = "bytes not fetched yet (restartable read)";
/// Settings of a [`LazyStorage`].
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct LazyConfig {
/// Size of a block in bytes (at least 512); fetches are whole, aligned
/// blocks (the file's last block is shorter).
pub block_size: u64,
/// Byte budget of cached blocks between operations. An operation keeps
/// every block it needs until it finishes, whatever the budget.
pub capacity: u64,
/// Largest single range asked for, in bytes (whole blocks, at least
/// one); longer runs are split so they can be fetched in parallel.
pub max_request: u64,
/// Most bytes one operation may fetch (at least one block), and so the
/// longest single read: a read longer than this fails at once, before
/// anything is fetched, and so does an operation whose passes would
/// fetch more. The file's length comes from the server, so without
/// this a hostile file (a heap "collection" claiming 2 GiB) makes the
/// reader fetch and hold whatever it names; on wasm32 a buffer past
/// 2 GiB cannot even be allocated.
pub max_fetch: u64,
}
impl Default for LazyConfig {
fn default() -> Self {
LazyConfig {
block_size: DEFAULT_BLOCK_SIZE,
capacity: 64 << 20,
max_request: 8 << 20,
max_fetch: DEFAULT_MAX_FETCH,
}
}
}
/// What a [`LazyStorage`] has done so far.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub struct LazyStats {
/// Passes run by [`LazyStorage::attempt`].
pub passes: u64,
/// Ranges handed to [`LazyStorage::supply`]: one HTTP request each.
pub requests: u64,
/// Bytes handed to [`LazyStorage::supply`].
pub bytes_fetched: u64,
/// Blocks evicted to stay within the budget.
pub evictions: u64,
/// Bytes cached now.
pub cached_bytes: u64,
}
/// The outcome of one pass.
#[derive(Debug)]
pub enum Step<T> {
/// The pass read only bytes that were present: its result stands.
Done(T),
/// The pass missed: fetch these byte ranges (sorted, disjoint, block
/// aligned), [`supply`](LazyStorage::supply) them and run it again.
Need(Vec<Range<u64>>),
}
struct Block {
data: Arc<[u8]>,
/// Eviction order: bulk blocks (`false`) before metadata (`true`),
/// then least recently used first.
key: (bool, u64),
}
#[derive(Default)]
struct State {
blocks: HashMap<u64, Block>,
/// `(metadata?, tick, block index)`, in eviction order.
order: BTreeSet<(bool, u64, u64)>,
tick: u64,
bytes: u64,
/// Blocks the current pass missed, and whether a small read wanted
/// them (metadata).
missing: HashMap<u64, bool>,
/// Blocks a bulk read missed that have not been supplied yet: kept
/// as bulk when they arrive.
bulk_pending: BTreeSet<u64>,
/// Operations in flight: no eviction while any is.
active: u32,
stats: LazyStats,
}
/// A [`Storage`] over the blocks of a file fetched so far; a read of
/// anything else fails and is recorded, so the pass can be re-run once the
/// bytes arrive. See the [module documentation](self).
pub struct LazyStorage {
len: u64,
config: LazyConfig,
state: Mutex<State>,
}
fn lock(m: &Mutex<State>) -> MutexGuard<'_, State> {
m.lock().unwrap_or_else(std::sync::PoisonError::into_inner)
}
/// Keeps an operation's blocks cached until it is dropped; see
/// [`LazyStorage::operation`].
pub struct Operation<'a> {
storage: &'a LazyStorage,
/// Bytes fetched for this operation so far.
fetched: std::cell::Cell<u64>,
}
impl Operation<'_> {
/// Count `ranges` against the operation's budget
/// ([`LazyConfig::max_fetch`]) before they are fetched: an error, and
/// nothing counted, if they would take it past the budget.
pub fn charge(&self, ranges: &[Range<u64>]) -> Result<(), String> {
let max = self.storage.config.max_fetch;
let total = ranges.iter().fold(self.fetched.get(), |n, r| {
n.saturating_add(r.end.saturating_sub(r.start))
});
if total > max {
return Err(format!(
"this call would fetch more than {max} bytes of the file (the maxFetch limit); \
read less at a time (readHyperslab) or raise maxFetch"
));
}
self.fetched.set(total);
Ok(())
}
}
impl Drop for Operation<'_> {
fn drop(&mut self) {
let mut st = lock(&self.storage.state);
st.active = st.active.saturating_sub(1);
if st.active == 0 {
self.storage.evict(&mut st);
}
}
}
impl LazyStorage {
/// An empty cache for a file of `len` bytes.
pub fn new(len: u64, mut config: LazyConfig) -> Self {
config.block_size = config.block_size.max(512);
config.max_request = (config.max_request / config.block_size).max(1) * config.block_size;
config.max_fetch = config.max_fetch.max(config.block_size);
LazyStorage {
len,
config,
state: Mutex::new(State::default()),
}
}
/// The settings in use (after rounding).
pub fn config(&self) -> &LazyConfig {
&self.config
}
/// Counters since the storage was made.
pub fn stats(&self) -> LazyStats {
let st = lock(&self.state);
LazyStats {
cached_bytes: st.bytes,
..st.stats
}
}
/// Mark an operation in flight until the guard is dropped: no block is
/// evicted meanwhile, so re-running its passes always makes progress.
/// Hold it across every pass of one operation.
pub fn operation(&self) -> Operation<'_> {
lock(&self.state).active += 1;
Operation {
storage: self,
fetched: std::cell::Cell::new(0),
}
}
/// Run one pass of `f` over this storage. `Done` when `f` read nothing
/// that is missing; otherwise `Need` with the ranges to fetch, and `f`'s
/// result is dropped (it may be an error caused by the miss, or a
/// result built around one).
pub fn attempt<T>(&self, f: impl FnOnce() -> T) -> Step<T> {
{
let mut st = lock(&self.state);
st.missing.clear();
st.stats.passes += 1;
}
let out = f();
let missing = std::mem::take(&mut lock(&self.state).missing);
if missing.is_empty() {
return Step::Done(out);
}
drop(out);
Step::Need(self.runs(missing))
}
/// The bytes of the file at `offset`, fetched for a range a pass asked
/// for. `offset` must be block aligned and the bytes whole blocks (the
/// last block of the file may be short) inside the file, or this is an
/// error and nothing is kept. Blocks already present are left alone.
pub fn supply(&self, offset: u64, bytes: &[u8]) -> Result<(), String> {
let bs = self.config.block_size;
let end = offset
.checked_add(bytes.len() as u64)
.filter(|&e| e <= self.len)
.ok_or_else(|| {
format!(
"{} bytes at offset {offset} run past the end of the {}-byte file",
bytes.len(),
self.len
)
})?;
if !offset.is_multiple_of(bs) || (!end.is_multiple_of(bs) && end != self.len) {
return Err(format!(
"{} bytes at offset {offset} are not whole {bs}-byte blocks",
bytes.len()
));
}
let mut st = lock(&self.state);
st.stats.requests += 1;
st.stats.bytes_fetched += bytes.len() as u64;
let mut start = offset;
while start < end {
let i = start / bs;
let stop = (start + bs).min(end);
if !st.blocks.contains_key(&i) {
let rel = (start - offset) as usize..(stop - offset) as usize;
let metadata = !st.bulk_pending.remove(&i);
self.keep(&mut st, i, Arc::from(&bytes[rel]), metadata);
}
start = stop;
}
if st.active == 0 {
self.evict(&mut st);
}
Ok(())
}
/// [`supply`](Self::supply) the bytes fetched for `range`, one of the
/// ranges a [`Step::Need`] asked for: anything but exactly its length
/// (a server that answered with more or less) is an error.
pub fn supply_range(&self, range: &Range<u64>, bytes: &[u8]) -> Result<(), String> {
let want = range.end.saturating_sub(range.start);
if bytes.len() as u64 != want {
return Err(format!(
"asked for {want} bytes at offset {}, got {}",
range.start,
bytes.len()
));
}
self.supply(range.start, bytes)
}
/// Run `f` to completion, fetching what its passes miss with `fetch`
/// (a byte range to its bytes). The blocking driver, for native code
/// and tests; the browser's is the same loop with an `await` between
/// passes.
pub fn run_blocking<T>(
&self,
mut f: impl FnMut() -> T,
mut fetch: impl FnMut(Range<u64>) -> Result<Vec<u8>, String>,
) -> Result<T, String> {
let op = self.operation();
loop {
match self.attempt(&mut f) {
Step::Done(v) => return Ok(v),
Step::Need(ranges) => {
op.charge(&ranges)?;
for r in ranges {
let bytes = fetch(r.clone())?;
self.supply_range(&r, &bytes)?;
}
}
}
}
}
/// Cache block `i`.
fn keep(&self, st: &mut State, i: u64, data: Arc<[u8]>, metadata: bool) {
st.tick += 1;
let key = (metadata, st.tick);
st.bytes += <[u8]>::len(&data) as u64;
st.order.insert((key.0, key.1, i));
if let Some(old) = st.blocks.insert(i, Block { data, key }) {
st.order.remove(&(old.key.0, old.key.1, i));
st.bytes -= <[u8]>::len(&old.data) as u64;
}
}
fn evict(&self, st: &mut State) {
while st.bytes > self.config.capacity {
let Some((_, _, i)) = st.order.pop_first() else {
break;
};
if let Some(b) = st.blocks.remove(&i) {
st.bytes -= <[u8]>::len(&b.data) as u64;
st.stats.evictions += 1;
}
}
}
/// Byte ranges covering the missing blocks: runs of consecutive
/// blocks, a one-block hole between two runs filled so they merge
/// (unless the hole is cached: it would be fetched again), each at
/// most `max_request` long.
fn runs(&self, missing: HashMap<u64, bool>) -> Vec<Range<u64>> {
let bs = self.config.block_size;
let mut wanted: Vec<u64> = missing.keys().copied().collect();
wanted.sort_unstable();
let mut st = lock(&self.state);
// Remember which blocks only bulk reads asked for: they are kept
// as bulk once supplied.
for (&i, &metadata) in &missing {
if metadata {
st.bulk_pending.remove(&i);
} else {
st.bulk_pending.insert(i);
}
}
let per_request = self.config.max_request / bs;
let mut runs: Vec<(u64, u64)> = Vec::new();
for i in wanted {
match runs.last_mut() {
Some((first, last))
if (i == *last + 1
|| (i == *last + 2 && !st.blocks.contains_key(&(i - 1))))
&& i - *first < per_request =>
{
*last = i
}
_ => runs.push((i, i)),
}
}
drop(st);
runs.into_iter()
.map(|(a, b)| a * bs..((b + 1) * bs).min(self.len))
.collect()
}
/// Block indices covering `[offset, offset + len)`, clamped to the file.
fn span(&self, offset: u64, len: u64) -> Option<Range<u64>> {
let end = offset.saturating_add(len).min(self.len);
if offset >= end {
return None;
}
let bs = self.config.block_size;
Some(offset / bs..(end - 1) / bs + 1)
}
/// The blocks of `spans` if all are present (touching them), else
/// record the missing ones and fail.
fn blocks(
&self,
spans: &[Range<u64>],
metadata: bool,
) -> Result<HashMap<u64, Arc<[u8]>>, FormatError> {
let mut st = lock(&self.state);
let mut have = HashMap::new();
let mut missed = false;
for span in spans {
for i in span.clone() {
if have.contains_key(&i) {
continue;
}
match st.blocks.get(&i) {
Some(b) => {
have.insert(i, b.data.clone());
}
None => {
missed = true;
let m = st.missing.entry(i).or_insert(metadata);
*m |= metadata;
}
}
}
}
if missed {
return Err(FormatError::Storage(NEED_BYTES.into()));
}
// Touch: most recently used last; a small read promotes a bulk
// block to metadata.
for &i in have.keys() {
st.tick += 1;
let tick = st.tick;
let Some(b) = st.blocks.get_mut(&i) else {
continue;
};
let old = b.key;
b.key = (old.0 || metadata, tick);
let new = b.key;
st.order.remove(&(old.0, old.1, i));
st.order.insert((new.0, new.1, i));
}
Ok(have)
}
/// Refuse a read of `n` bytes longer than an operation may fetch
/// ([`LazyConfig::max_fetch`]), before its blocks are asked for.
fn check_len(&self, n: u64) -> Result<(), FormatError> {
let max = self.config.max_fetch;
if n > max {
return Err(FormatError::Storage(format!(
"a read of {n} bytes is more than one call may fetch ({max} bytes, the maxFetch limit)"
)));
}
Ok(())
}
/// The bytes `offset..end` from `blocks`, which hold every block of
/// that span. The buffer is reserved fallibly: a length the address
/// space cannot hold (past `isize::MAX` on wasm32) is an error, never
/// an abort.
fn assemble(
&self,
offset: u64,
end: u64,
blocks: &HashMap<u64, Arc<[u8]>>,
) -> Result<Vec<u8>, FormatError> {
let bs = self.config.block_size;
let n = end - offset;
let too_long =
|| FormatError::Storage(format!("cannot hold a read of {n} bytes in memory"));
let mut out = Vec::new();
out.try_reserve_exact(usize::try_from(n).map_err(|_| too_long())?)
.map_err(|_| too_long())?;
let mut pos = offset;
while pos < end {
let i = pos / bs;
let block = &blocks[&i];
let from = (pos - i * bs) as usize;
let to = ((end - i * bs) as usize).min(<[u8]>::len(block));
out.extend_from_slice(&block[from..to]);
pos = i * bs + to as u64;
}
Ok(out)
}
}
impl Storage for LazyStorage {
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
let Some(span) = self.span(offset, len as u64) else {
return Ok(Cow::Owned(Vec::new()));
};
let end = offset.saturating_add(len as u64).min(self.len);
self.check_len(end - offset)?;
let metadata = len as u64 <= self.config.block_size;
let blocks = self.blocks(std::slice::from_ref(&span), metadata)?;
Ok(Cow::Owned(self.assemble(offset, end, &blocks)?))
}
fn len(&self) -> u64 {
self.len
}
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
let mut spans = Vec::with_capacity(ranges.len());
let mut total = 0u64;
for r in ranges {
if r.end < r.start {
return Err(FormatError::Storage(
"read range ends before it starts".into(),
));
}
total = total.saturating_add(r.end.min(self.len).saturating_sub(r.start));
spans.extend(self.span(r.start, r.end - r.start));
}
self.check_len(total)?;
let blocks = self.blocks(&spans, false)?;
ranges
.iter()
.map(|r| {
let end = r.end.min(self.len);
if r.start >= end {
Ok(Cow::Owned(Vec::new()))
} else {
self.assemble(r.start, end, &blocks).map(Cow::Owned)
}
})
.collect()
}
}
#[cfg(test)]
#[allow(clippy::single_range_in_vec_init)]
mod tests {
use super::*;
fn file(n: usize) -> Vec<u8> {
(0..n).map(|i| (i * 7 + i / 251) as u8).collect()
}
fn config(block: u64, capacity: u64) -> LazyConfig {
LazyConfig {
block_size: block,
capacity,
max_request: 4 * block,
max_fetch: DEFAULT_MAX_FETCH,
}
}
/// Supply every range of `need` from `data`.
fn serve(s: &LazyStorage, data: &[u8], need: &[Range<u64>]) {
for r in need {
s.supply_range(r, &data[r.start as usize..r.end as usize])
.unwrap();
}
}
fn owned(r: Result<Cow<'_, [u8]>, FormatError>) -> Result<Vec<u8>, FormatError> {
r.map(Cow::into_owned)
}
#[test]
fn a_miss_asks_for_whole_blocks_then_the_rerun_reads_them() {
let data = file(10_000);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let Step::Need(need) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
panic!("nothing is cached yet");
};
assert_eq!(need, vec![1024..3072]);
serve(&s, &data, &need);
let Step::Done(got) = s.attempt(|| owned(s.read_at(1500, 1000))) else {
panic!("the blocks were supplied");
};
assert_eq!(got.unwrap(), &data[1500..2500]);
// Past the end: short, then empty, as for a slice.
serve(&s, &data, &[9216..10_000]);
let Step::Done(tail) = s.attempt(|| owned(s.read_at(9_990, 100))) else {
panic!("the last block was supplied");
};
assert_eq!(tail.unwrap(), &data[9_990..]);
assert!(matches!(
s.attempt(|| s.read_at(20_000, 10).map(|c| c.len())),
Step::Done(Ok(0))
));
let st = s.stats();
assert_eq!(
(st.passes, st.requests, st.bytes_fetched),
(4, 2, 2048 + 784)
);
}
#[test]
fn a_pass_that_swallowed_the_miss_is_still_rerun() {
// A parser that catches the error and returns something anyway
// (a listing skipping a link it cannot resolve) must not have its
// result used.
let data = file(4096);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let step = s.attempt(|| s.read_at(0, 4).map(|b| b.len()).unwrap_or(0));
assert!(
matches!(step, Step::Need(ref n) if n == &vec![0..1024]),
"{step:?}"
);
}
#[test]
fn read_ranges_asks_for_every_missing_block_in_one_pass() {
let data = file(64 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let ranges = [
100..200,
5000..5100,
5200..5300,
30_000..33_000,
60_000..60_010,
];
let read = || {
s.read_ranges(&ranges)
.map(|v| v.into_iter().map(Cow::into_owned).collect::<Vec<_>>())
};
let Step::Need(need) = s.attempt(read) else {
panic!("nothing is cached yet");
};
// 5000..5300 is blocks 4 and 5; 30_000..33_000 is blocks 29..=32.
assert_eq!(
need,
vec![
0..1024,
4096..6144,
29 * 1024..33 * 1024,
58 * 1024..59 * 1024
]
);
serve(&s, &data, &need);
let Step::Done(got) = s.attempt(read) else {
panic!("one pass fetched everything");
};
for (r, g) in ranges.iter().zip(&got.unwrap()) {
assert_eq!(g, &data[r.start as usize..r.end as usize]);
}
}
#[test]
fn runs_merge_one_block_holes_and_split_long_runs() {
let data = file(32 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
// Blocks 0 and 2 (a hole of one: merged), 5..=14 (split in fours).
let Step::Need(need) = s.attempt(|| {
let _ = s.read_at(0, 10);
let _ = s.read_at(2048, 10);
s.read_ranges(&[5120..15 * 1024]).map(|_| ())
}) else {
panic!("nothing is cached yet");
};
assert_eq!(
need,
vec![0..3072, 5120..9216, 9216..13_312, 13_312..15_360]
);
}
#[test]
fn a_cached_hole_is_not_fetched_again() {
let data = file(8 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
serve(&s, &data, &[1024..2048]);
// Blocks 0 and 2 missing, 1 cached: two requests, not 0..3072.
let Step::Need(need) = s.attempt(|| {
let _ = s.read_at(0, 10);
s.read_at(2048, 10).map(|_| ())
}) else {
panic!("blocks 0 and 2 are missing");
};
assert_eq!(need, vec![0..1024, 2048..3072]);
}
#[test]
fn supply_refuses_what_was_not_asked_for() {
let data = file(10_000);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
assert!(s.supply(1, &data[1..1025]).unwrap_err().contains("whole"));
assert!(s.supply(0, &data[..1000]).unwrap_err().contains("whole"));
assert!(
s.supply(9216, &[0u8; 1024])
.unwrap_err()
.contains("past the end")
);
for wrong in [&data[..1023], &data[..2048]] {
assert!(
s.supply_range(&(0..1024), wrong)
.unwrap_err()
.contains("asked for 1024 bytes")
);
}
assert_eq!(s.stats().cached_bytes, 0);
// The file's last, short block is whole.
s.supply(9216, &data[9216..]).unwrap();
assert_eq!(s.stats().cached_bytes, 784);
}
#[test]
fn an_operation_keeps_its_blocks_whatever_the_budget() {
// A budget of one block, an operation that needs eight: without
// the operation guard every pass would evict what the last one
// fetched and never finish.
let data = file(8 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 1024));
let got = s
.run_blocking(
|| {
(0..8)
.map(|i| owned(s.read_at(i * 1024 + 3, 10)))
.collect::<Result<Vec<_>, _>>()
},
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap()
.unwrap();
for (i, g) in got.iter().enumerate() {
assert_eq!(g, &data[i * 1024 + 3..i * 1024 + 13]);
}
// Trimmed to the budget once the operation is over.
let st = s.stats();
assert_eq!(st.cached_bytes, 1024);
assert_eq!(st.evictions, 7);
assert_eq!(st.passes, 9, "one pass per block, then the one that ends");
}
#[test]
fn bulk_blocks_are_evicted_before_metadata() {
let data = file(16 * 1024);
let s = LazyStorage::new(data.len() as u64, config(1024, 4 * 1024));
let fetch = |r: Range<u64>| Ok(data[r.start as usize..r.end as usize].to_vec());
// Metadata: a small read of block 0.
s.run_blocking(|| s.read_at(0, 16).map(|_| ()), fetch)
.unwrap()
.unwrap();
// Bulk: raw data over blocks 4..12, more than the budget.
s.run_blocking(|| s.read_ranges(&[4096..12 * 1024]).map(|_| ()), fetch)
.unwrap()
.unwrap();
// The metadata block survived: reading it again is a hit.
let before = s.stats();
assert!(before.cached_bytes <= 4 * 1024);
assert!(matches!(
s.attempt(|| s.read_at(0, 16).map(|_| ())),
Step::Done(Ok(()))
));
assert_eq!(s.stats().requests, before.requests);
}
#[test]
fn a_read_longer_than_max_fetch_fails_without_fetching() {
// A hostile file names a 2 GiB heap collection in a "file" the
// server claims is 1 TiB: the read is refused before any block is
// asked for (on wasm32 its buffer could not even be allocated).
let s = LazyStorage::new(1 << 40, LazyConfig::default());
let step = s.attempt(|| s.read_at(4096, (1usize << 31) + 4096).map(|b| b.len()));
match step {
Step::Done(Err(e)) => assert!(e.to_string().contains("maxFetch"), "{e}"),
other => panic!("expected a refusal, got {other:?}"),
}
let step = s.attempt(|| s.read_ranges(&[0..(600 << 20)]).map(|v| v.len()));
assert!(matches!(step, Step::Done(Err(_))), "{step:?}");
assert_eq!(s.stats().requests, 0);
// At the limit it is an ordinary miss.
let s = LazyStorage::new(1 << 40, config(1024, 1 << 20));
let step = s.attempt(|| s.read_at(0, DEFAULT_MAX_FETCH as usize).map(|b| b.len()));
assert!(matches!(step, Step::Need(_)), "{step:?}");
}
#[test]
fn an_operation_stops_at_its_fetch_budget() {
// Many small reads, none over the limit, that together would fetch
// more than the budget: the operation fails before fetching past it.
let data = file(64 * 1024);
let mut c = config(1024, 1 << 20);
c.max_fetch = 8 * 1024;
let s = LazyStorage::new(data.len() as u64, c);
let e = s
.run_blocking(
|| {
(0..64)
.map(|i| owned(s.read_at(i * 1024, 8)))
.collect::<Result<Vec<_>, _>>()
},
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap_err();
assert!(e.contains("maxFetch"), "{e}");
assert!(s.stats().bytes_fetched <= 8 * 1024, "{:?}", s.stats());
// Within the budget it completes, and the budget is per operation.
for _ in 0..3 {
s.run_blocking(
|| owned(s.read_at(10 * 1024, 3000)),
|r| Ok(data[r.start as usize..r.end as usize].to_vec()),
)
.unwrap()
.unwrap();
}
}
#[test]
fn a_failed_fetch_is_an_error_not_data() {
let data = file(4096);
let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20));
let e = s
.run_blocking(|| owned(s.read_at(0, 8)), |_| Err("HTTP 500".into()))
.unwrap_err();
assert_eq!(e, "HTTP 500");
let e = s
.run_blocking(|| owned(s.read_at(0, 8)), |_| Ok(vec![0; 10]))
.unwrap_err();
assert!(e.contains("asked for 1024 bytes"), "{e}");
assert_eq!(s.stats().cached_bytes, 0);
drop(data);
}
}
+476 -62
View File
@@ -1,7 +1,7 @@
//! clawhdf5's HDF5 reader for JavaScript, via `wasm-bindgen`.
//!
//! ```js
//! import init, { open } from "./pkg/clawhdf5_wasm.js";
//! import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
//! await init();
//! const file = open(new Uint8Array(await blob.arrayBuffer()));
//! file.list("/"); // [{ name, kind: "group" | "dataset" }]
@@ -10,6 +10,12 @@
//! file.read("/x"); // { shape, dtype, data: Float64Array | ... | string[] }
//! file.readHyperslab("/x", [0, 0], [10, 10]); // stride, block optional
//! file.free();
//!
//! // A file on a web server, read by HTTP range requests as needed: the
//! // same methods, returning promises.
//! const remote = await openUrl("https://example.org/data.h5");
//! await remote.list("/");
//! remote.stats(); // { requests, bytesFetched, size, ... }
//! ```
//!
//! Numeric data comes back in the typed array of the stored width
@@ -18,15 +24,23 @@
//! Anything else is a thrown `Error` naming the datatype. Only the reader is
//! exposed: nothing here writes files.
//!
//! The logic lives in [`core`], which is plain Rust and tested natively.
//! The logic lives in [`core`] and [`lazy`], which are plain Rust and tested
//! natively; `js/remote.js` does the HTTP.
pub mod core;
pub mod lazy;
use std::ops::Range;
use std::rc::Rc;
use std::sync::Arc;
use clawhdf5::AttrValue;
use js_sys::{Array, Object, Reflect};
use js_sys::{Array, Object, Promise, Reflect, Uint8Array};
use wasm_bindgen::prelude::*;
use wasm_bindgen_futures::future_to_promise;
use crate::core::{Data, Hyperslab, Reader};
use crate::core::{Attr, Child, Data, DatasetInfo, Hyperslab, Reader};
use crate::lazy::{LazyConfig, LazyStorage, Step};
/// JavaScript numbers are exact up to 2^53.
const MAX_SAFE_INTEGER: f64 = 9_007_199_254_740_991.0;
@@ -35,11 +49,30 @@ fn js_err(msg: String) -> JsError {
JsError::new(&msg)
}
/// A JavaScript exception (from `fetch`, or `js/remote.js`) as a `JsError`
/// with its message.
fn js_exception(e: JsValue) -> JsError {
let msg = e
.dyn_ref::<js_sys::Error>()
.map(|e| String::from(e.message()))
.or_else(|| e.as_string())
.unwrap_or_else(|| format!("{e:?}"));
JsError::new(&msg)
}
fn set(obj: &Object, key: &str, value: impl Into<JsValue>) {
// Defining a property on a fresh plain object cannot fail.
Reflect::set(obj, &JsValue::from_str(key), &value.into()).unwrap_throw();
}
fn get(obj: &JsValue, key: &str) -> JsValue {
if obj.is_object() {
Reflect::get(obj, &JsValue::from_str(key)).unwrap_or(JsValue::UNDEFINED)
} else {
JsValue::UNDEFINED
}
}
fn shape_to_js(shape: &[u64]) -> Array {
shape.iter().map(|&d| JsValue::from_f64(d as f64)).collect()
}
@@ -58,6 +91,20 @@ fn indices_from_js(what: &str, v: &[f64]) -> Result<Vec<u64>, JsError> {
.collect()
}
fn slab_from_js(
start: &[f64],
count: &[f64],
stride: Option<Vec<f64>>,
block: Option<Vec<f64>>,
) -> Result<Hyperslab, JsError> {
Ok(Hyperslab {
start: indices_from_js("start", start)?,
count: indices_from_js("count", count)?,
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
})
}
fn data_to_js(data: Data) -> JsValue {
match data {
Data::F32(v) => js_sys::Float32Array::from(&v[..]).into(),
@@ -110,6 +157,72 @@ fn attr_to_js(value: AttrValue) -> (JsValue, Option<String>) {
}
}
fn list_to_js(children: Vec<Child>) -> Array {
children
.into_iter()
.map(|c| {
let o = Object::new();
set(&o, "name", c.name);
set(&o, "kind", c.kind.as_str());
JsValue::from(o)
})
.collect()
}
fn info_to_js(i: DatasetInfo) -> Object {
let o = Object::new();
set(&o, "shape", shape_to_js(&i.shape));
let max: JsValue = match i.maxshape {
None => JsValue::NULL,
Some(dims) => dims
.into_iter()
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
.collect::<Array>()
.into(),
};
set(&o, "maxshape", max);
set(&o, "dtype", i.dtype);
set(&o, "elementShape", shape_to_js(&i.element_shape));
o
}
fn attrs_to_js(attrs: Vec<Attr>) -> Array {
attrs
.into_iter()
.map(|a| {
let o = Object::new();
set(&o, "name", a.name);
let (value, dtype) = attr_to_js(a.value);
set(&o, "value", value);
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
JsValue::from(o)
})
.collect()
}
fn errors_to_js(errors: Vec<String>) -> Array {
errors.into_iter().map(JsValue::from).collect()
}
/// A read's values with the dataset's datatype.
fn values_to_js((dtype, a): (String, core::Array)) -> Object {
let o = Object::new();
set(&o, "shape", shape_to_js(&a.shape));
set(&o, "dtype", dtype);
set(&o, "data", data_to_js(a.data));
o
}
/// Read the dataset at `path` (whole, or `slab`) with its datatype.
fn read_values(
r: &Reader,
path: &str,
slab: Option<&Hyperslab>,
) -> core::Result<(String, core::Array)> {
let dtype = r.info(path)?.dtype;
Ok((dtype, r.read(path, slab)?))
}
/// An open HDF5 (or NetCDF-4) file.
#[wasm_bindgen]
pub struct H5File {
@@ -145,38 +258,13 @@ impl H5File {
/// The group's members: `[{ name, kind }]`, groups first.
pub fn list(&self, path: &str) -> Result<Array, JsError> {
Ok(self
.inner
.list(path)
.map_err(js_err)?
.into_iter()
.map(|c| {
let o = Object::new();
set(&o, "name", c.name);
set(&o, "kind", c.kind.as_str());
JsValue::from(o)
})
.collect())
Ok(list_to_js(self.inner.list(path).map_err(js_err)?))
}
/// `{ shape, maxshape, dtype, elementShape }`. `maxshape` is `null`
/// when not recorded, with `null` for each unlimited dimension.
pub fn info(&self, path: &str) -> Result<Object, JsError> {
let i = self.inner.info(path).map_err(js_err)?;
let o = Object::new();
set(&o, "shape", shape_to_js(&i.shape));
let max: JsValue = match i.maxshape {
None => JsValue::NULL,
Some(dims) => dims
.into_iter()
.map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64)))
.collect::<Array>()
.into(),
};
set(&o, "maxshape", max);
set(&o, "dtype", i.dtype);
set(&o, "elementShape", shape_to_js(&i.element_shape));
Ok(o)
Ok(info_to_js(self.inner.info(path).map_err(js_err)?))
}
/// `[{ name, value, dtype }]`, sorted by name. Scalars are `number`
@@ -185,31 +273,21 @@ impl H5File {
/// and its `dtype`; one that could not be read at all is reported by
/// [`attrErrors`](Self::attr_errors).
pub fn attrs(&self, path: &str) -> Result<Array, JsError> {
let (attrs, _) = self.inner.attrs(path).map_err(js_err)?;
Ok(attrs
.into_iter()
.map(|a| {
let o = Object::new();
set(&o, "name", a.name);
let (value, dtype) = attr_to_js(a.value);
set(&o, "value", value);
set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from));
JsValue::from(o)
})
.collect())
Ok(attrs_to_js(self.inner.attrs(path).map_err(js_err)?.0))
}
/// Messages for attributes that could not be read.
#[wasm_bindgen(js_name = attrErrors)]
pub fn attr_errors(&self, path: &str) -> Result<Array, JsError> {
let (_, errors) = self.inner.attrs(path).map_err(js_err)?;
Ok(errors.into_iter().map(JsValue::from).collect())
Ok(errors_to_js(self.inner.attrs(path).map_err(js_err)?.1))
}
/// The whole dataset: `{ shape, dtype, data }`, `data` in row-major
/// order.
pub fn read(&self, path: &str) -> Result<Object, JsError> {
self.read_impl(path, None)
Ok(values_to_js(
read_values(&self.inner, path, None).map_err(js_err)?,
))
}
/// A regular hyperslab (`H5Sselect_hyperslab`): `stride` and `block`
@@ -223,22 +301,358 @@ impl H5File {
stride: Option<Vec<f64>>,
block: Option<Vec<f64>>,
) -> Result<Object, JsError> {
let slab = Hyperslab {
start: indices_from_js("start", &start)?,
count: indices_from_js("count", &count)?,
stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?,
block: block.map(|b| indices_from_js("block", &b)).transpose()?,
};
self.read_impl(path, Some(&slab))
}
fn read_impl(&self, path: &str, slab: Option<&Hyperslab>) -> Result<Object, JsError> {
let dtype = self.inner.info(path).map_err(js_err)?.dtype;
let a = self.inner.read(path, slab).map_err(js_err)?;
let slab = slab_from_js(&start, &count, stride, block)?;
Ok(values_to_js(
read_values(&self.inner, path, Some(&slab)).map_err(js_err)?,
))
}
}
// ---------------------------------------------------------------------------
// Remote files: openUrl.
#[wasm_bindgen(module = "/js/remote.js")]
extern "C" {
#[wasm_bindgen(catch)]
async fn probe(url: &str, first_len: f64, opts: &JsValue) -> Result<JsValue, JsValue>;
#[wasm_bindgen(catch, js_name = fetchRanges)]
async fn fetch_ranges(
url: &str,
ranges: Vec<f64>,
opts: &JsValue,
validator: &JsValue,
length: f64,
) -> Result<JsValue, JsValue>;
}
/// Where a remote file's bytes come from.
struct Http {
url: String,
opts: JsValue,
/// ETag or Last-Modified at open (`null` if the server sent neither).
validator: JsValue,
length: u64,
/// Requests the probe made (1, or 2 with a HEAD for the length).
probe_requests: u64,
}
enum Source {
/// Read by range requests through a restartable cache.
Lazy {
http: Http,
storage: Arc<LazyStorage>,
reader: Reader,
},
/// The server ignored `Range`: the whole file, downloaded at open.
Whole {
reader: Reader,
size: u64,
requests: u64,
},
}
impl Http {
/// Fetch `ranges` and hand them to `storage`.
async fn fetch(&self, storage: &LazyStorage, ranges: &[Range<u64>]) -> Result<(), JsError> {
let flat: Vec<f64> = ranges
.iter()
.flat_map(|r| [r.start as f64, r.end as f64])
.collect();
let got = fetch_ranges(
&self.url,
flat,
&self.opts,
&self.validator,
self.length as f64,
)
.await
.map_err(js_exception)?;
let got = Array::from(&got);
if got.length() as usize != ranges.len() {
return Err(js_err(format!(
"fetchRanges returned {} ranges for {}",
got.length(),
ranges.len()
)));
}
for (r, bytes) in ranges.iter().zip(got.iter()) {
let bytes = Uint8Array::new(&bytes).to_vec();
storage.supply_range(r, &bytes).map_err(js_err)?;
}
Ok(())
}
/// Run `f` over `storage` until it has every byte it reads.
async fn drive<T>(
&self,
storage: &LazyStorage,
mut f: impl FnMut() -> T,
) -> Result<T, JsError> {
let op = storage.operation();
loop {
match storage.attempt(&mut f) {
Step::Done(v) => return Ok(v),
Step::Need(ranges) => {
op.charge(&ranges).map_err(js_err)?;
self.fetch(storage, &ranges).await?
}
}
}
}
}
impl Source {
async fn run<T>(&self, op: impl Fn(&Reader) -> core::Result<T>) -> Result<T, JsError> {
match self {
Source::Whole { reader, .. } => op(reader).map_err(js_err),
Source::Lazy {
http,
storage,
reader,
} => http.drive(storage, || op(reader)).await?.map_err(js_err),
}
}
}
/// A non-negative integer option, or `None` when not given.
fn int_opt(opts: &JsValue, key: &str, min: f64, max: f64) -> Result<Option<u64>, JsError> {
let v = get(opts, key);
if v.is_undefined() || v.is_null() {
return Ok(None);
}
match v.as_f64() {
Some(x) if x.fract() == 0.0 && (min..=max).contains(&x) => Ok(Some(x as u64)),
_ => Err(js_err(format!(
"openUrl: {key} must be an integer from {min} to {max}"
))),
}
}
/// Most bytes `maxFetch` and `maxDownload` may allow: 1 GiB. wasm32 has
/// 4 GiB of memory and no buffer past 2 GiB, and what is fetched is held
/// while it is decoded.
const MAX_FETCH_LIMIT: u64 = 1 << 30;
/// The largest file `openUrl` reads by ranges: on wasm32, 4 GiB - 1 bytes.
/// The format code turns file offsets into `usize` to use them (with a
/// clean error past it, see scripts/check-32bit-casts.sh), so on a 32-bit
/// target nothing at 4 GiB or beyond can be read; a larger file is refused
/// at open rather than failing on whichever read reaches past 4 GiB. On
/// 64-bit targets it is 2^53 - 1, the largest offset a JavaScript number
/// holds exactly.
const MAX_REMOTE_LENGTH: u64 = if (usize::MAX as u64) < MAX_SAFE_INTEGER as u64 {
usize::MAX as u64
} else {
MAX_SAFE_INTEGER as u64
};
fn config_from(opts: &JsValue) -> Result<LazyConfig, JsError> {
let mut c = LazyConfig::default();
if let Some(b) = int_opt(opts, "blockSize", 512.0, (64u64 << 20) as f64)? {
c.block_size = b;
}
if let Some(n) = int_opt(opts, "cacheSize", 0.0, MAX_SAFE_INTEGER)? {
c.capacity = n;
}
if let Some(n) = int_opt(opts, "maxFetch", 512.0, MAX_FETCH_LIMIT as f64)? {
c.max_fetch = n;
}
// Read by remote.js; checked here so a value wasm32 cannot hold is an
// option error rather than a download that cannot be kept.
int_opt(opts, "maxDownload", 0.0, MAX_FETCH_LIMIT as f64)?;
// Also read by remote.js (which checks it too, for direct callers).
int_opt(opts, "parallel", 1.0, 1024.0)?;
Ok(c)
}
/// Open the HDF5 file at `url` without downloading it: its bytes are
/// fetched with HTTP `Range` requests as the methods of the returned
/// [`RemoteFile`] need them, through a block cache.
///
/// `opts` (all optional):
/// - `blockSize` — bytes per request block, 512 to 64 MiB (default 1 MiB);
/// - `cacheSize` — bytes of blocks kept between calls (default 64 MiB);
/// - `maxFetch` — most bytes one call may fetch, and so the longest single
/// read, up to 1 GiB (default 512 MiB): a call that would fetch more
/// fails before fetching it;
/// - `fallback` — `"download"` (default) reads the whole file when the
/// server ignores `Range` (answers 200), up to `maxDownload` bytes
/// (default 512 MiB, at most 1 GiB); `"error"` refuses such a server;
/// - `headers`, `credentials` — passed to every `fetch` (`headers` as
/// `fetch` takes them: a `Headers`, `[name, value]` pairs or an object);
/// - `parallel` — range requests in flight at once, 1 to 1024 (default
/// 6); when one fails the others are aborted;
/// - `fetch` — a `fetch`-compatible function to use instead of the global.
///
/// Cross-origin servers must allow CORS and expose `Content-Range` (or
/// answer `HEAD` with `Content-Length`). A file may be up to 4 GiB - 1
/// bytes long (wasm32 offsets); a longer one is refused at open. A whole-dataset `read` that would use more than 1 GiB of
/// memory ([`core::MAX_READ_BYTES`]) is refused: read it in parts with
/// `readHyperslab`.
#[wasm_bindgen(js_name = openUrl)]
pub async fn open_url(url: String, opts: JsValue) -> Result<RemoteFile, JsError> {
let config = config_from(&opts)?;
let p = probe(&url, config.block_size as f64, &opts)
.await
.map_err(js_exception)?;
let requests = get(&p, "requests").as_f64().unwrap_or(1.0) as u64;
let whole = get(&p, "whole");
if !whole.is_undefined() {
let bytes = Uint8Array::new(&whole).to_vec();
let size = bytes.len() as u64;
let reader = Reader::open(bytes).map_err(js_err)?;
return Ok(RemoteFile {
inner: Rc::new(Source::Whole {
reader,
size,
requests,
}),
});
}
let length = get(&p, "length")
.as_f64()
.filter(|x| x.fract() == 0.0 && (0.0..=MAX_SAFE_INTEGER).contains(x))
.ok_or_else(|| js_err(format!("{url}: the server gave no usable file size")))?;
if length as u64 > MAX_REMOTE_LENGTH {
return Err(js_err(format!(
"{url} is {length} bytes; openUrl reads files of up to {MAX_REMOTE_LENGTH} bytes \
(4 GiB - 1: the WebAssembly reader addresses a file with 32-bit offsets)"
)));
}
let http = Http {
url,
opts,
validator: get(&p, "validator"),
length: length as u64,
probe_requests: requests,
};
let storage = Arc::new(LazyStorage::new(http.length, config));
let first = Uint8Array::new(&get(&p, "first")).to_vec();
storage.supply(0, &first).map_err(js_err)?;
let s = storage.clone();
let reader = http
.drive(&storage, || Reader::open_storage(s.clone()))
.await?
.map_err(js_err)?;
Ok(RemoteFile {
inner: Rc::new(Source::Lazy {
http,
storage,
reader,
}),
})
}
/// A file opened with [`openUrl`](open_url): the methods of [`H5File`],
/// each returning a `Promise` (it may have to fetch bytes first).
#[wasm_bindgen]
pub struct RemoteFile {
inner: Rc<Source>,
}
impl RemoteFile {
/// Run `op` (fetching what it needs) and convert its result.
fn call<T: 'static>(
&self,
op: impl Fn(&Reader) -> core::Result<T> + 'static,
to_js: impl FnOnce(T) -> JsValue + 'static,
) -> Promise {
let inner = self.inner.clone();
future_to_promise(async move {
let v = inner.run(op).await.map_err(JsValue::from)?;
Ok(to_js(v))
})
}
}
#[wasm_bindgen]
impl RemoteFile {
/// `"group"` or `"dataset"`.
#[wasm_bindgen(unchecked_return_type = "Promise<string>")]
pub fn kind(&self, path: String) -> Promise {
self.call(move |r| r.kind(&path), |k| k.as_str().into())
}
/// The group's members: `[{ name, kind }]`, groups first.
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
pub fn list(&self, path: String) -> Promise {
self.call(move |r| r.list(&path), |c| list_to_js(c).into())
}
/// `{ shape, maxshape, dtype, elementShape }`, as [`H5File::info`].
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
pub fn info(&self, path: String) -> Promise {
self.call(move |r| r.info(&path), |i| info_to_js(i).into())
}
/// `[{ name, value, dtype }]`, as [`H5File::attrs`].
#[wasm_bindgen(unchecked_return_type = "Promise<Array<any>>")]
pub fn attrs(&self, path: String) -> Promise {
self.call(move |r| r.attrs(&path), |a| attrs_to_js(a.0).into())
}
/// Messages for attributes that could not be read.
#[wasm_bindgen(js_name = attrErrors, unchecked_return_type = "Promise<Array<string>>")]
pub fn attr_errors(&self, path: String) -> Promise {
self.call(move |r| r.attrs(&path), |a| errors_to_js(a.1).into())
}
/// The whole dataset: `{ shape, dtype, data }`, as [`H5File::read`].
#[wasm_bindgen(unchecked_return_type = "Promise<any>")]
pub fn read(&self, path: String) -> Promise {
self.call(
move |r| read_values(r, &path, None),
|v| values_to_js(v).into(),
)
}
/// A regular hyperslab, as [`H5File::read_hyperslab`]. Only the chunks
/// (or the contiguous runs) the selection touches are fetched.
#[wasm_bindgen(js_name = readHyperslab, unchecked_return_type = "Promise<any>")]
pub fn read_hyperslab(
&self,
path: String,
start: Vec<f64>,
count: Vec<f64>,
stride: Option<Vec<f64>>,
block: Option<Vec<f64>>,
) -> Result<Promise, JsError> {
let slab = slab_from_js(&start, &count, stride, block)?;
Ok(self.call(
move |r| read_values(r, &path, Some(&slab)),
|v| values_to_js(v).into(),
))
}
/// What reading this file has cost so far: `{ lazy, size, requests,
/// bytesFetched, cachedBytes, passes }`. `lazy` is false when the
/// server ignored `Range` and the file was downloaded whole.
pub fn stats(&self) -> Object {
let o = Object::new();
set(&o, "shape", shape_to_js(&a.shape));
set(&o, "dtype", dtype);
set(&o, "data", data_to_js(a.data));
Ok(o)
match &*self.inner {
Source::Lazy { http, storage, .. } => {
let st = storage.stats();
set(&o, "lazy", true);
set(&o, "size", http.length as f64);
set(
&o,
"requests",
(st.requests.saturating_sub(1) + http.probe_requests) as f64,
);
set(&o, "bytesFetched", st.bytes_fetched as f64);
set(&o, "cachedBytes", st.cached_bytes as f64);
set(&o, "passes", st.passes as f64);
}
Source::Whole { size, requests, .. } => {
set(&o, "lazy", false);
set(&o, "size", *size as f64);
set(&o, "requests", *requests as f64);
set(&o, "bytesFetched", *size as f64);
set(&o, "cachedBytes", *size as f64);
set(&o, "passes", 0.0);
}
}
o
}
}
+578
View File
@@ -0,0 +1,578 @@
//! The restartable ("NeedBytes") reader against the in-memory one: every
//! file must list, describe and read the same through a [`LazyStorage`]
//! that starts empty and is fed only the ranges its passes ask for, as the
//! browser's `openUrl` feeds it from HTTP range requests.
//!
//! - Files written here with `FileBuilder`, at several block sizes (512 B
//! blocks make almost every structure read a miss).
//! - The h5py/netCDF4 fixture of `examples/wasm-viewer/test/make_fixture.py`
//! (skipped without h5py, unless `CLAWHDF5_REQUIRE_INTEROP=1`;
//! `CLAWHDF5_PYTHON` names the interpreter).
//! - `CLAWHDF5_WASM_CORPUS=dir[:dir...]`: every HDF5 file under those
//! directories up to 64 MiB (e.g. `conformance/.cache/corpus`).
//!
//! Also the request budget: listing and reading one small dataset of a large
//! file fetches a few blocks, not the file.
use std::ops::Range;
use std::path::{Path, PathBuf};
use std::process::Command;
use std::sync::Arc;
use clawhdf5::{AttrValue, FileBuilder};
use clawhdf5_format::storage::CountingStorage;
use clawhdf5_wasm::core::{Hyperslab, Kind, Reader};
use clawhdf5_wasm::lazy::{LazyConfig, LazyStorage};
/// The API the JavaScript side calls, one operation at a time.
trait Api {
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T;
}
struct Local(Reader);
impl Api for Local {
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
op(&self.0)
}
}
/// A lazily read file and the "server" it fetches from.
struct Lazy {
data: Arc<Vec<u8>>,
storage: Arc<LazyStorage>,
reader: Reader,
}
fn fetch(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
Ok(data[r.start as usize..r.end as usize].to_vec())
}
impl Lazy {
/// Open as `openUrl` does: the first block comes with the probe that
/// learns the length, then the open is run until it has its bytes.
fn open(data: Vec<u8>, config: LazyConfig) -> Result<Lazy, String> {
let data = Arc::new(data);
let storage = Arc::new(LazyStorage::new(data.len() as u64, config));
let first = (storage.config().block_size as usize).min(data.len());
storage.supply(0, &data[..first])?;
let s = storage.clone();
let reader =
storage.run_blocking(|| Reader::open_storage(s.clone()), |r| fetch(&data, r))??;
Ok(Lazy {
data,
storage,
reader,
})
}
}
impl Api for Lazy {
fn call<T>(&self, op: impl Fn(&Reader) -> T) -> T {
self.storage
.run_blocking(|| op(&self.reader), |r| fetch(&self.data, r))
.expect("serving from memory cannot fail")
}
}
/// Everything the viewer can show of a file, as text: each object's kind,
/// listing, attributes (and attribute errors), dataset info, whole value
/// and a hyperslab — or the error each gives.
fn transcript(api: &impl Api) -> Vec<String> {
let mut out = Vec::new();
let mut todo = vec![("/".to_string(), 0usize)];
while let Some((path, depth)) = todo.pop() {
if out.len() > 4000 {
out.push("... (truncated)".into());
break;
}
let kind = api.call(|r| r.kind(&path));
out.push(format!("{path}: {kind:?}"));
out.push(format!("{path} attrs: {:?}", api.call(|r| r.attrs(&path))));
match kind {
Ok(Kind::Group) => {
let list = api.call(|r| r.list(&path));
out.push(format!("{path} list: {list:?}"));
if let Ok(children) = list
&& depth < 12
{
for c in children.into_iter().rev() {
let child = if path == "/" {
format!("/{}", c.name)
} else {
format!("{path}/{}", c.name)
};
todo.push((child, depth + 1));
}
}
}
Ok(Kind::Dataset) => {
let info = api.call(|r| r.info(&path));
out.push(format!("{path} info: {info:?}"));
let Ok(info) = info else { continue };
let n = info
.shape
.iter()
.chain(&info.element_shape)
.try_fold(1u64, |a, &d| a.checked_mul(d));
if n.is_none_or(|n| n > 4_000_000) {
out.push(format!("{path}: not read ({n:?} values)"));
continue;
}
out.push(format!(
"{path} read: {:?}",
api.call(|r| r.read(&path, None))
));
if !info.shape.is_empty() && info.shape.iter().all(|&d| d > 1) {
let slab = Hyperslab {
start: info.shape.iter().map(|_| 1).collect(),
count: info.shape.iter().map(|&d| d / 2).collect(),
stride: None,
block: None,
};
let part = api.call(|r| r.read(&path, Some(&slab)));
out.push(format!("{path} slab: {part:?}"));
}
}
Err(_) => {}
}
}
out
}
/// The lazy transcript of `data` at `block` bytes per block equals the
/// transcript of the same file through a range storage that has every byte
/// (`CountingStorage`: the facade's `Storage` path, the one the lazy reader
/// takes), and agrees with the in-memory one: the same values, and an error
/// wherever it has one (a malformed file can fail at a different check,
/// with a different message, when read by ranges). Returns what the lazy
/// reader fetched and its transcript.
fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec<String>) {
let ctx = format!("{name} (blocks of {block} B)");
let ranged = Reader::open_storage(Arc::new(CountingStorage::new(data.to_vec())));
let local = Reader::open(data.to_vec());
let lazy = Lazy::open(data.to_vec(), config(block));
let (ranged, local, lazy) = match (ranged, local, lazy) {
(Ok(r), Ok(l), Ok(z)) => (r, l, z),
(Err(r), Err(_), Err(z)) => {
assert_eq!(z, r, "{ctx}: open error");
return (0, 0, Vec::new());
}
(r, l, z) => panic!(
"{ctx}: opens differently: ranged {:?}, in memory {:?}, lazily {:?}",
r.err(),
l.err(),
z.err()
),
};
let got = transcript(&lazy);
let want = transcript(&Local(ranged));
for (i, (w, g)) in want.iter().zip(&got).enumerate() {
assert_eq!(g, w, "{ctx}, line {i}");
}
assert_eq!(got.len(), want.len(), "{ctx}: transcript length");
let local = transcript(&Local(local));
for (i, (l, g)) in local.iter().zip(&got).enumerate() {
let both_errors = match (l.split_once("Err("), g.split_once("Err(")) {
(Some((a, _)), Some((b, _))) => a == b,
_ => false,
};
assert!(
l == g || both_errors,
"{ctx}, line {i}: in memory\n {l}\nlazily\n {g}"
);
}
assert_eq!(got.len(), local.len(), "{ctx}: transcript length");
let st = lazy.storage.stats();
(st.requests, st.bytes_fetched, got)
}
fn config(block: u64) -> LazyConfig {
LazyConfig {
block_size: block,
// A small budget, so eviction between operations is exercised.
capacity: 16 * block,
max_request: 8 * block,
..LazyConfig::default()
}
}
fn builder_file() -> Vec<u8> {
let mut b = FileBuilder::new();
b.create_dataset("grid")
.with_f64_data(&(0..20_000).map(f64::from).collect::<Vec<_>>())
.with_shape(&[100, 200])
.with_chunks(&[10, 25])
.with_deflate(4);
b.create_dataset("contiguous")
.with_i32_data(&(0..50_000).collect::<Vec<_>>());
b.create_dataset("bytes").with_u8_data(&[1, 2, 250]);
let mut g = b.create_group("sensors");
for i in 0..40 {
g.create_dataset(&format!("t{i}"))
.with_f32_data(&[i as f32, 1.5, -2.25]);
}
g.set_attr("location", AttrValue::String("lab".into()));
b.add_group(g.finish());
b.set_attr("version", AttrValue::I64(3));
b.set_attr("scale", AttrValue::F64Array(vec![0.5, 2.0]));
b.finish().unwrap()
}
#[test]
fn builder_files_read_the_same_at_every_block_size() {
let data = builder_file();
for block in [512, 4096, 1 << 20] {
let (requests, _, lines) = check_equal("builder", &data, block);
assert!(requests > 0);
// The transcript covers every object, values included.
assert!(lines.iter().any(|l| l.starts_with("/grid read: Ok")));
assert!(lines.iter().any(|l| l.starts_with("/grid slab: Ok")));
assert!(lines.iter().any(|l| l.starts_with("/sensors/t39 read: Ok")));
}
}
#[test]
fn garbage_fails_to_open_as_in_memory() {
check_equal("zeros", &[0u8; 5000], 512);
check_equal("empty", &[], 512);
let mut cut = builder_file();
cut.truncate(cut.len() / 3);
check_equal("truncated", &cut, 512);
}
/// Listing a large file and reading one small dataset fetches a few blocks,
/// not the file.
#[test]
fn a_small_read_of_a_large_file_fetches_a_few_blocks() {
let mut b = FileBuilder::new();
b.create_dataset("small").with_f64_data(&[1.0, 2.0, 3.0]);
// 48 MB of raw data, written after the small dataset's metadata.
b.create_dataset("big")
.with_f64_data(&(0..6_000_000).map(f64::from).collect::<Vec<_>>());
let mut g = b.create_group("group");
g.create_dataset("inner").with_i32_data(&[7, 8]);
b.add_group(g.finish());
let data = b.finish().unwrap();
let lazy = Lazy::open(data.clone(), LazyConfig::default()).unwrap();
let list = lazy.call(|r| r.list("/")).unwrap();
assert_eq!(list.len(), 3);
assert_eq!(
format!("{:?}", lazy.call(|r| r.read("/small", None)).unwrap().data),
"F64([1.0, 2.0, 3.0])"
);
assert_eq!(
format!(
"{:?}",
lazy.call(|r| r.read("/group/inner", None)).unwrap().data
),
"I32([7, 8])"
);
// A window of the big dataset reads only its block(s).
let slab = Hyperslab {
start: vec![3_000_000],
count: vec![4],
stride: None,
block: None,
};
assert_eq!(
format!(
"{:?}",
lazy.call(|r| r.read("/big", Some(&slab))).unwrap().data
),
"F64([3000000.0, 3000001.0, 3000002.0, 3000003.0])"
);
let st = lazy.storage.stats();
eprintln!("{} bytes: {st:?}", data.len());
assert!(st.requests <= 6, "{st:?}");
assert!(st.bytes_fetched <= 6 << 20, "{st:?}");
assert!(st.bytes_fetched * 8 < data.len() as u64, "{st:?}");
}
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn python_available() -> bool {
Command::new(python())
.args(["-c", "import h5py, netCDF4, numpy"])
.output()
.is_ok_and(|o| o.status.success())
}
#[test]
fn h5py_and_netcdf4_files_read_the_same_lazily() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
python()
);
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
return;
}
let dir = fixture_dir();
for name in ["fixture.h5", "fixture.nc"] {
let data = std::fs::read(dir.path().join(name)).unwrap();
for block in [512, 64 * 1024] {
let (_, _, lines) = check_equal(name, &data, block);
assert!(lines.iter().filter(|l| l.contains(" read: Ok")).count() >= 2);
}
}
}
/// The bytes of `data` at `r`, zero past its end: a server that claims
/// the file is longer than it is.
fn fetch_padded(data: &[u8], r: Range<u64>) -> Result<Vec<u8>, String> {
let mut out = vec![0u8; (r.end - r.start) as usize];
let len = data.len() as u64;
if r.start < len {
let end = r.end.min(len);
out[..(end - r.start) as usize].copy_from_slice(&data[r.start as usize..end as usize]);
}
Ok(out)
}
/// Sizes a hostile server or a large dataset can name are errors, never
/// allocations that abort the wasm module: make_fixture.py's limits.h5 and
/// hostile_vl.h5 (see write_limits there).
#[test]
fn size_limits_are_errors_not_aborts() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
python()
);
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
return;
}
let dir = fixture_dir();
// Read whole, /huge_u8 would widen 2^28 values to 64 bits (2 GiB): an
// error naming readHyperslab, before its chunks are read. A window of
// it reads.
let data = std::fs::read(dir.path().join("limits.h5")).unwrap();
let n = (1u64 << 28) + 1024;
let window = Hyperslab {
start: vec![n - 4],
count: vec![4],
stride: None,
block: None,
};
let local = Reader::open(data.clone()).unwrap();
let lazy = Lazy::open(data, LazyConfig::default()).unwrap();
let before = lazy.storage.stats().requests;
for e in [
local.read("/huge_u8", None).unwrap_err(),
lazy.call(|r| r.read("/huge_u8", None)).unwrap_err(),
] {
assert!(e.contains("readHyperslab"), "{e}");
}
assert_eq!(lazy.storage.stats().requests, before, "nothing fetched");
for part in [
local.read("/huge_u8", Some(&window)).unwrap(),
lazy.call(|r| r.read("/huge_u8", Some(&window))).unwrap(),
] {
assert_eq!(format!("{:?}", part.data), "U8([0, 0, 0, 7])");
}
// A server that claims 3 GiB and a heap collection of 2 GiB + 4 KiB:
// reading the strings fails at once, fetching a few blocks.
let data = std::fs::read(dir.path().join("hostile_vl.h5")).unwrap();
let storage = Arc::new(LazyStorage::new(3 << 30, LazyConfig::default()));
let s = storage.clone();
let reader = storage
.run_blocking(
|| Reader::open_storage(s.clone()),
|r| fetch_padded(&data, r),
)
.unwrap()
.unwrap();
let e = storage
.run_blocking(|| reader.read("/a", None), |r| fetch_padded(&data, r))
.unwrap()
.unwrap_err();
assert!(e.contains("maxFetch"), "{e}");
let st = storage.stats();
assert!(st.requests <= 4 && st.bytes_fetched <= 4 << 20, "{st:?}");
}
/// make_fixture.py's files, written to a temporary directory.
fn fixture_dir() -> tempfile::TempDir {
let dir = tempfile::tempdir().unwrap();
let generator = Path::new(env!("CARGO_MANIFEST_DIR"))
.join("../../examples/wasm-viewer/test/make_fixture.py");
let out = Command::new(python())
.arg(&generator)
.arg(dir.path())
.output()
.unwrap();
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
dir
}
fn hdf5_files(dir: &Path, out: &mut Vec<PathBuf>) {
let Ok(entries) = std::fs::read_dir(dir) else {
return;
};
for e in entries.flatten() {
let p = e.path();
if p.is_dir() {
hdf5_files(&p, out);
} else if std::fs::read(&p)
.ok()
.is_some_and(|b| b.len() <= 64 << 20 && is_hdf5(&b))
{
out.push(p);
}
}
}
/// The HDF5 signature at 0 or a power-of-two user-block offset.
fn is_hdf5(b: &[u8]) -> bool {
const SIG: &[u8] = b"\x89HDF\r\n\x1a\n";
let mut at = 0usize;
loop {
if b.get(at..at + 8) == Some(SIG) {
return true;
}
at = if at == 0 { 512 } else { at * 2 };
if at >= b.len() {
return false;
}
}
}
#[test]
fn corpus_files_read_the_same_lazily() {
let Ok(dirs) = std::env::var("CLAWHDF5_WASM_CORPUS") else {
eprintln!("CLAWHDF5_WASM_CORPUS not set; skipping the corpus");
return;
};
let mut files = Vec::new();
for d in std::env::split_paths(&dirs) {
hdf5_files(&d, &mut files);
}
files.sort();
assert!(!files.is_empty(), "no HDF5 files under {dirs}");
let (mut requests, mut bytes, mut total) = (0u64, 0u64, 0u64);
for f in &files {
let data = std::fs::read(f).unwrap();
total += data.len() as u64;
let (r, b, _) = check_equal(&f.display().to_string(), &data, 64 * 1024);
requests += r;
bytes += b;
}
eprintln!(
"{} files ({total} bytes): {requests} requests, {bytes} bytes fetched",
files.len()
);
}
/// Passes and requests `list(path)` takes on a file opened lazily at
/// `block`-byte blocks (the open not counted), checking the listing against
/// the in-memory one.
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
let lazy = Lazy::open(
data.to_vec(),
LazyConfig {
block_size: block,
..LazyConfig::default()
},
)
.unwrap();
let before = lazy.storage.stats();
assert_eq!(lazy.call(|r| r.list(path)).unwrap(), want);
let after = lazy.storage.stats();
(
after.passes - before.passes,
after.requests - before.requests,
)
}
/// Listing a group reads every child's object header, and its index (B-tree
/// and symbol table nodes, or B-tree v2 and heap blocks) before that. Each
/// pass asks for every node of a level it is missing, not the first one
/// only, so the passes (network round trips) grow with the depth of the
/// index, not with the number of children: 2000 children with headers
/// scattered over 512-byte blocks list in a handful of passes, where each
/// header block used to cost its own.
#[test]
fn listing_a_large_group_takes_a_few_passes() {
let mut b = FileBuilder::new();
let mut g = b.create_group("many");
for i in 0..600 {
g.create_dataset(&format!("d{i}")).with_i32_data(&[i; 64]);
}
b.add_group(g.finish());
let data = b.finish().unwrap();
let (passes, requests) = listing_cost(&data, "/many", 512);
eprintln!("FileBuilder, 600 children: {passes} passes, {requests} requests");
assert!(passes <= 6, "{passes} passes");
if !python_available() {
return;
}
let dir = tempfile::tempdir().unwrap();
for libver in ["earliest", "latest"] {
let path = dir.path().join(format!("{libver}.h5"));
let script = format!(
"import h5py, numpy as np\n\
with h5py.File({:?}, 'w', libver='{libver}') as f:\n\
\x20 for i in range(2000):\n\
\x20 f.create_dataset('d%d' % i, data=np.full(256, i, np.float32))\n",
path.display().to_string()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.unwrap();
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let data = std::fs::read(&path).unwrap();
let (passes, requests) = listing_cost(&data, "/", 512);
eprintln!("h5py libver={libver}, 2000 children: {passes} passes, {requests} requests");
assert!(passes <= 12, "{libver}: {passes} passes");
}
}
/// `CLAWHDF5_WASM_LIST_FILE=file.h5`: what listing the root group of that
/// file costs lazily, at 1 MiB and 64 KiB blocks (a measurement, printed).
#[test]
fn listing_cost_of_a_given_file() {
let Ok(path) = std::env::var("CLAWHDF5_WASM_LIST_FILE") else {
return;
};
let data = std::fs::read(&path).unwrap();
for block in [1 << 20, 64 << 10] {
let lazy = Lazy::open(
data.clone(),
LazyConfig {
block_size: block,
..LazyConfig::default()
},
)
.unwrap();
let open = lazy.storage.stats();
let n = lazy.call(|r| r.list("/")).unwrap().len();
let st = lazy.storage.stats();
eprintln!(
"{path} ({} bytes), {block}-byte blocks: open {} requests / {} passes; list('/') of {n}: {} passes, {} requests, {} bytes",
data.len(),
open.requests,
open.passes,
st.passes - open.passes,
st.requests - open.requests,
st.bytes_fetched - open.bytes_fetched
);
}
}
+74 -3
View File
@@ -20,13 +20,20 @@ change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet.
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
reader, through the restartable `NeedBytes` mode (see the M4 status
below). M5 (SWMR) is not started.
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
converted code without changing the plan: object-header continuation chunks
are followed without recursion (still one bounded `read_at` per chunk), and
implicit chunk indexes are addressed over the maximum chunk grid (in
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
working on the whole file in memory; it is not part of this design. Every count below was
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
next to it. No timing numbers appear here on purpose: the machine was shared
working on the whole file in memory; it is not part of this design. Every
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
every count in a milestone's status on the date it gives, with the
commands given next to it. No timing numbers appear here on purpose: the machine was shared
with other build jobs when this was written.
## The problem
@@ -536,6 +543,70 @@ fast path within benchmark noise.
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
to a whole download when the server does not answer 206.
- `examples/wasm-viewer`: open by URL.
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
with these choices:
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
`Storage` over the blocks fetched so far. A call (open, list, read)
runs as a pass; a read that misses records its blocks and fails with
a storage error. The error itself is not the signal: parsers catch
errors and carry on (a listing leaves out a link it cannot resolve),
so *any* pass that recorded a miss is thrown away, whatever it
returned, and re-run once the missing ranges have been fetched
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
synchronous XHR would have kept the parsers' single pass, but it
needs the page to run the reader off the main thread (a second
module, messages for every call and every typed array copied back),
since synchronous XHR can return binary data only there; the
restartable loop costs a
re-parse per wave of misses instead, which is CPU, not network.
- **Progress:** no block is evicted while a call is in flight
(`LazyStorage::operation`), so each pass that does not finish asks for
at least one new block, and a call ends after at most one pass per
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
evicted before metadata. The price: a call holds everything it reads
until it finishes.
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
backend by blocking, evicts during a read, and does not keep reads
that miss more than half its budget; a restartable pass needs every
block it has read to be there when it is re-run. The block
arithmetic and coalescing (runs of consecutive blocks, a one-block
hole filled, at most 8 MiB per request) follow it; the cache is
about 450 lines (with its documentation) in the wasm crate, with no
in-flight tracking or HTTP. (A hole already cached is not filled:
it would be fetched again.)
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
requests at a time, every answer checked (206, `Content-Range` when
visible, body length, the ETag/Last-Modified and length of the
first answer). A `200` to the first request is kept as the whole
file up to `maxDownload` (512 MiB), or refused with
`fallback: "error"`. The first block comes with the request that
learns the length, as in M3.
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
--test lazy` compares what the viewer shows of each file (kinds,
listings, attributes, info, whole reads, hyperslabs) read lazily at
512 B to 1 MiB blocks with the in-memory reader and the facade's
range-storage path, for built files, the h5py/netCDF4 fixture and,
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
files. The built package, under Node and headless Chromium
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
`test/serve.py`, which counts requests): every fixture check over
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
attributes and a 10-value window of the 25-million-value dataset
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
with `open(bytes)`.
- After review (2026-09-27): a call may fetch at most `maxFetch`
(512 MiB, at most 1 GiB) and a longer single read is refused before
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
more is refused at open on wasm32 (offsets become `usize`); every
response body is cut off at the length asked for. Parsers that walk
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
the wasm listing's child headers) read the siblings after the first
failure before returning it, so a pass asks for a whole level's
missing blocks: listing 3000 datasets went from 185 passes to 6.
The 32-bit risk below is covered by a Node test that reads data at
3 GiB from a mock server and is refused a 4 GiB file.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
+68 -3
View File
@@ -948,6 +948,11 @@ cache, but:
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
(`h5rs` and the Python bindings read through `File::storage`, and take
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
feature; the wasm reader's `openUrl` reads by range requests since
2026-09-27, its `open(bytes)` takes a whole file).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -969,6 +974,10 @@ cache, but:
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
`File::as_bytes`, which a remote file does not have. (The browser can
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.
@@ -1007,10 +1016,66 @@ cache, but:
## `clawhdf5-wasm` (browser) limits
**Status:** open (by design for now; added 2026-09-26).
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
- The whole file is held in memory: `open()` takes its bytes. There are no
HTTP range reads, so a multi-GB file does not fit a browser tab.
- `open()` holds the whole file in memory (it takes its bytes), so a
multi-GB local file does not fit a browser tab. A file on a web server
can be opened with `openUrl()` instead, which fetches only the byte
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
with these limits:
- **Round trips:** a call runs as passes over the blocks fetched so far
and is re-run after each wave of misses, so a call costs one round
trip per wave, not one for everything: a chunk index is walked a
level (or a node) per round trip, while the chunks of a read are
fetched together. Listing a group asks for every child's object
header, and every node of a level of the group's index, in one pass
(since 2026-09-27; it was one round trip per header block): 3000
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
`libver="latest"` file (dense links). Each pass re-parses what the
call reads (CPU, not network). With headers spread through the file
(h5py writes each next to its data) a listing still fetches most of
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
- **Memory:** a call keeps every block it reads until it finishes (the
cache budget applies between calls). It may fetch at most `maxFetch`
bytes (512 MiB by default, at most 1 GiB), and a single read longer
than that fails before anything is fetched: a hostile server cannot
make the page fetch or allocate what a file's lengths claim. `read()`
of a dataset that would take more than 1 GiB while it is decoded
(stored bytes, the values widened to 64 bits, the result) fails,
naming `readHyperslab`; read windows of large datasets. On wasm32 a
buffer past 2 GiB cannot exist and a failed allocation aborts the
whole module (every open file on the page), which these limits keep
from happening; before 2026-09-27 both did abort it.
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
open. The format code turns file offsets into `usize` to use them
(with a clean error past it), which is 32 bits on wasm32, so nothing
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
tested (with a mock server); files above 200 MB have not been served
for real.
- **Cross-origin servers** must allow CORS for the page's origin and
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
answer `HEAD` with `Content-Length`. The file is pinned at open by its
`ETag` or `Last-Modified` (and its length): when the page cannot see
either header, only a change of length is detected.
- **A server without range support** (it answers `200`) costs a whole
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
with `fallback: "error"`. Every body, this one and each `206`, is read
as it arrives and cut off past its limit (the range asked for, or
`maxDownload`): a server cannot make the page buffer more.
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
size is not used. No retries: a failed request fails the call (calling
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included (a page on 127.0.0.1 reading a
file from localhost, with and without exposed headers); not in Firefox
or Safari.
- The native corpus comparison (`tests/lazy.rs` with
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
and which chunk's error is reported depends on the iteration order of
the chunk index (a `HashMap`, seeded per process), so the lazy and
the range-storage reads can name different errors. Both are errors;
not specific to `openUrl` (it predates it).
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`.
+100 -20
View File
@@ -2,10 +2,12 @@
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
file, browse its groups, and look at a dataset's type, shape, attributes
and values (a 50 x 12 window at a time, read as a hyperslab, with the
leading dimensions of a 3-D+ dataset held at chosen indices). The file never
leaves the page.
file or give a URL, browse its groups, and look at a dataset's type, shape,
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
local file never leaves the page. A file given by URL is not downloaded:
only the byte ranges each view needs are fetched (HTTP range requests),
and the header shows the requests and bytes that has cost so far.
## Build and open
@@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
```
Then open <http://localhost:8000/>. `?file=<url>&path=<object>` opens a
file from a URL (same origin, or one serving CORS headers) and selects an
object in it, e.g. `?file=data/run1.h5&path=/results/energy`.
Then open <http://localhost:8000/>. The URL box, or
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
sending CORS headers, see Limits) and selects an object in it, e.g.
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
answer range requests, so a file served by it is downloaded whole (the
header says so); `test/serve.py` does:
```bash
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
```
## JavaScript API
```js
import init, { open } from "./pkg/clawhdf5_wasm.js";
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
await init();
const f = open(new Uint8Array(await blob.arrayBuffer()));
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
@@ -32,8 +42,41 @@ f.attrs("/grid"); // [{ name, value, dtype }]
f.read("/grid"); // { shape, dtype, data }
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
f.free();
// A file on a web server, read by HTTP range requests as needed. The same
// methods, each returning a promise.
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
await r.list("/");
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
r.free();
```
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
between calls, default 64 MiB), `maxFetch` (bytes one call may fetch, and
so the longest single read, up to 1 GiB, default 512 MiB), `fallback`
(`"download"`, the default, reads the whole file when the server ignores
`Range`, up to `maxDownload` bytes, default 512 MiB, at most 1 GiB;
`"error"` refuses such a server), `headers` (a `Headers`, `[name, value]`
pairs or an object) and `credentials` (passed to every request),
`parallel` (requests in flight, 1 to 1024, default 6; when one fails the
others are aborted), `fetch` (a `fetch`-compatible function to use).
How it works: the reader is synchronous and a page cannot block on the
network, so each call runs as a *pass* over the blocks fetched so far. A
pass that needs a block not yet fetched is abandoned, the missing blocks
are fetched (in parallel, adjacent blocks in one request), and the pass is
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
costs one request (the first block, which also gives the file's size);
listing a group whose metadata is in blocks already fetched costs none,
and otherwise a round trip per level of the group's index plus one for
its children's headers, all fetched together;
reading a chunked dataset costs a round trip for its chunk index (a few
for a deep one) and one batch of requests for its chunks. Every answer is
checked — a `206` with exactly the bytes asked for, from the same file
(ETag or Last-Modified, and length) — or the call fails.
`data` is the typed array of the stored width (`Float64Array`,
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
`BigUint64Array`), or an array of strings for fixed- and variable-length
@@ -43,7 +86,17 @@ else throws an `Error` naming the type.
## Limits
- Read-only, and the whole file is held in memory (no range requests).
- Read-only. `open(bytes)` holds the whole file in memory.
- `openUrl`: a cross-origin server must allow CORS and expose
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
`Last-Modified` either, a file replaced on the server is detected only by
a change of length. A call keeps what it reads until it finishes, and
may fetch at most `maxFetch`; `read()` of a dataset that would take more
than 1 GiB to decode is refused (read windows of large datasets with
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in
`docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with
`value: null` and their `dtype`.
@@ -56,28 +109,55 @@ else throws an `Error` naming the type.
## Tests
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
`fixture.nc` (netCDF4) with `test/make_fixture.py`, then:
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
counter, `/norange/...` for a server without range support,
`/noexpose/...` and `/unexposed/...` for one that does not expose its
headers to CORS), then:
- runs `test/test.mjs` under Node: every dataset (whole and a strided
hyperslab), listing and attribute is compared with what libhdf5 reads
back, error paths are checked, and so are the page's DOM-free helpers
(`viewer-lib.js`);
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
and 512 B blocks), calls in flight at once, the request budget of
`big.h5` (listing it and reading small things of it must take at most 8
requests and under 5% of the file), the download fallback and its
limit, and the errors: HTTP status, a file that changes, a server that
sends the wrong bytes or stops honouring `Range`. Also the cross-origin
path with no exposed headers (`/noexpose/`: length from `HEAD`), the
size limits on `make_fixture.py`'s `limits.h5`, `hostile_vl.h5` (a heap
collection claiming 2 GiB) and `far.h5` (data at 3 GiB, served by a
mock), bodies longer than asked for, `headers` forms, `parallel`, and
sibling requests aborted after a failure.
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
to 16 MiB) read by URL with the same file read from bytes;
- runs `test/browser.sh`: loads the page in headless Chromium with
`?file=fixture.h5&path=...` for eight objects and checks the rendered tree,
types, shapes, attribute and value cells, and the error shown for an
unsupported type. Skipped when no Chromium is found (`CHROME` names one;
a Playwright download under `~/.cache/ms-playwright` is picked up).
Drag-and-drop and the file picker are not driven by it; they share
`load()` with the `?file=` path.
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
tree, types, shapes, attribute and value cells, the request counter, and
the error shown for an unsupported type; then the file from another
origin (localhost), with CORS exposing `Content-Range` and exposing
nothing (`/unexposed/`, where the server must see a `HEAD`), a server
without range support, and `big.h5` (a small dataset and a window of the large one,
with a single-digit percentage of the file fetched). Skipped when no
Chromium is found (`CHROME` names one; a Playwright download under
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
and the URL box are not driven by it; they share `setFile()` with the
`?file=` path.
The same expectations are checked natively, without Node, by
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI
container has no Node or browser).
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
corpus too), which is what CI runs (the CI container has no Node or
browser).
## Size
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`:
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
larger now and the table has not been re-measured: the reader has grown
since, and `openUrl` (2026-09-27) made the facade's range-read path
reachable from JavaScript and added promise glue and `remote.js`.
| | raw | gzip -9 |
|---|---:|---:|
+94 -44
View File
@@ -56,27 +56,39 @@
border: 1px solid var(--line); border-radius: 4px; }
.error { color: var(--bad); font-family: var(--mono); white-space: pre-wrap; }
.muted { color: var(--muted); }
form.url { display: flex; gap: 6px; margin-left: auto; flex: 1 1 320px; max-width: 560px; }
form.url input { flex: 1; min-width: 0; font: 13px var(--mono); padding: 4px 8px; background: var(--panel);
color: var(--ink); border: 1px solid var(--line); border-radius: 6px; }
.netstats { flex-basis: 100%; color: var(--muted); font-size: 12.5px; font-family: var(--mono); }
.netstats:empty { display: none; }
</style>
</head>
<body>
<header>
<h1>HDF5 Viewer</h1>
<span class="file" id="filename">no file</span>
<form class="url" id="urlform">
<input type="url" id="url" placeholder="https://…/file.h5 (read by HTTP range requests)" aria-label="File URL">
<button class="button" type="submit">Open URL</button>
</form>
<label class="button">Open file…<input type="file" id="picker" accept=".h5,.hdf5,.he5,.nc,.nc4,.cdf" hidden></label>
<span class="netstats" id="netstats"></span>
</header>
<main>
<nav id="tree"></nav>
<section id="detail">
<div class="drop">
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, or use “Open file…”.</p>
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; it is not uploaded anywhere.</p>
<p><strong>Drop an HDF5 or NetCDF-4 file here</strong>, use “Open file…”, or give a URL.</p>
<p class="muted">The file is read in this page by clawhdf5 compiled to WebAssembly; a local file is not uploaded anywhere.
A file given by URL is not downloaded: only the byte ranges each view needs are fetched (HTTP range requests),
and the requests and bytes it has cost are shown at the top.</p>
<p class="muted" id="version"></p>
</div>
</section>
</main>
<script type="module">
import init, { open, version } from "./pkg/clawhdf5_wasm.js";
import { joinPath, formatValue, viewWindow, toRows, perElement } from "./viewer-lib.js";
import init, { open, openUrl, version } from "./pkg/clawhdf5_wasm.js";
import { joinPath, formatValue, viewWindow, toRows, perElement, formatStats } from "./viewer-lib.js";
const $ = (id) => document.getElementById(id);
const el = (tag, props = {}, ...kids) => {
@@ -85,30 +97,58 @@ const el = (tag, props = {}, ...kids) => {
return e;
};
// The open file: an H5File (a local file, read in memory; synchronous
// methods) or a RemoteFile (openUrl: bytes fetched by range requests as
// needed; methods return promises). Every call below is awaited, so both
// work the same.
let file = null;
let selectedNode = null;
// Bumped by every selection, so a slow read for an object no longer shown
// does not overwrite the current one.
let showSeq = 0;
await init();
$("version").textContent = `clawhdf5 ${version()}`;
async function load(blob) {
const bytes = new Uint8Array(await blob.arrayBuffer());
// Requests and bytes a remote file has cost so far.
function updateStats() {
$("netstats").textContent = file && file.stats ? formatStats(file.stats()) : "";
}
async function setFile(name, opener) {
if (file) file.free();
file = null;
$("filename").textContent = blob.name;
selectedNode = null;
showSeq++;
$("filename").textContent = name;
$("tree").replaceChildren();
$("netstats").textContent = "";
$("detail").replaceChildren(el("p", { className: "muted", textContent: `Opening ${name}…` }));
try {
file = open(bytes);
file = await opener();
} catch (e) {
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${blob.name}: ${e.message}` }));
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot open ${name}: ${e.message}` }));
return;
}
updateStats();
const root = el("ul", { className: "tree" });
root.append(treeNode("/", "/", "group"));
$("tree").append(root);
root.querySelector(".node").click();
await activate(root.querySelector(".node"));
}
async function load(blob) {
const bytes = new Uint8Array(await blob.arrayBuffer());
await setFile(blob.name, async () => open(bytes));
}
async function loadUrl(url) {
$("url").value = url;
await setFile(url, () => openUrl(new URL(url, location.href).href));
}
// A tree row: clicking it selects it (and expands a group); `activate`
// does the same and resolves once the listing and the detail are shown.
function treeNode(path, name, kind) {
const li = el("li");
const icon = el("span", { className: "icon", textContent: kind === "group" ? "▸" : "·" });
@@ -116,34 +156,39 @@ function treeNode(path, name, kind) {
row.dataset.path = path;
li.append(row);
let children = null;
row.addEventListener("click", () => {
row.activate = async () => {
if (selectedNode) selectedNode.classList.remove("selected");
row.classList.add("selected");
selectedNode = row;
const shown = show(path, kind);
if (kind === "group") {
if (children) {
children.hidden = !children.hidden;
} else {
children = el("ul", { className: "tree" });
li.append(children);
try {
for (const c of file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
for (const c of await file.list(path)) children.append(treeNode(joinPath(path, c.name), c.name, c.kind));
} catch (e) {
children.append(el("li", { className: "error", textContent: e.message }));
}
li.append(children);
updateStats();
}
icon.textContent = children.hidden ? "▸" : "▾";
}
show(path, kind);
});
await shown;
};
row.addEventListener("click", () => row.activate());
return li;
}
function attrsTable(path) {
const activate = (row) => row.activate();
async function attrsTable(path) {
let attrs, errors;
try {
attrs = file.attrs(path);
errors = file.attrErrors(path);
attrs = await file.attrs(path);
errors = await file.attrErrors(path);
} catch (e) {
return el("p", { className: "error", textContent: e.message });
}
@@ -157,14 +202,16 @@ function attrsTable(path) {
return el("div", { className: "scroll" }, t);
}
function show(path, kind) {
async function show(path, kind) {
const seq = ++showSeq;
const current = () => seq === showSeq;
const out = [el("h2", { textContent: path })];
if (kind === "dataset") {
let info;
try {
info = file.info(path);
info = await file.info(path);
} catch (e) {
$("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
if (current()) $("detail").replaceChildren(...out, el("p", { className: "error", textContent: e.message }));
return;
}
const max = info.maxshape === null ? "—" : `(${info.maxshape.map((d) => d ?? "∞").join(", ")})`;
@@ -172,15 +219,16 @@ function show(path, kind) {
el("dt", { textContent: "type" }), el("dd", { textContent: info.dtype }),
el("dt", { textContent: "shape" }), el("dd", { textContent: `(${info.shape.join(", ")})` }),
el("dt", { textContent: "max shape" }), el("dd", { textContent: max })));
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
out.push(el("h3", { textContent: "Values" }), valuesView(path, info));
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
out.push(el("h3", { textContent: "Values" }), await valuesView(path, info, current));
} else {
out.push(el("h3", { textContent: "Attributes" }), attrsTable(path));
out.push(el("h3", { textContent: "Attributes" }), await attrsTable(path));
}
$("detail").replaceChildren(...out);
updateStats();
if (current()) $("detail").replaceChildren(...out);
}
function valuesView(path, info) {
async function valuesView(path, info, current) {
const shape = info.shape;
const per = perElement(info.elementShape);
const state = { row: 0, col: 0, rows: 50, cols: 12, fixed: shape.slice(0, Math.max(0, shape.length - 2)).map(() => 0) };
@@ -202,15 +250,19 @@ function valuesView(path, info) {
if (shape.length >= 2) controls.append(num("column", "col"));
state.fixed.forEach((_, i) => controls.append(num(`dim ${i}`, "fixed", i)));
function render() {
let renderSeq = 0;
async function render() {
const seq = ++renderSeq;
const w = viewWindow(shape, state);
let res;
try {
res = w === null ? file.read(path) : file.readHyperslab(path, w.start, w.count);
res = w === null ? await file.read(path) : await file.readHyperslab(path, w.start, w.count);
} catch (e) {
body.replaceChildren(el("p", { className: "error", textContent: e.message }));
if (seq === renderSeq) body.replaceChildren(el("p", { className: "error", textContent: e.message }));
return;
}
updateStats();
if (seq !== renderSeq || !current()) return;
if (w === null) {
body.replaceChildren(el("pre", { textContent: toRows(res.data, 1, 1, per)[0][0] }));
return;
@@ -229,41 +281,39 @@ function valuesView(path, info) {
(shape.length >= 2 ? `, columns ${w.col}–${w.col + w.cols - 1}` : "") + ` of ${total} values` });
body.replaceChildren(t, note);
}
render();
await render();
return box;
}
// Expand the tree down to `path` and select it.
function reveal(path) {
async function reveal(path) {
const rowFor = (p) => [...document.querySelectorAll(".node")].find((n) => n.dataset.path === p);
let cur = "/";
for (const part of path.split("/").filter(Boolean)) {
const row = rowFor(cur);
if (!row) return;
const kids = row.parentElement.querySelector(":scope > ul");
if (!kids || kids.hidden) row.click();
if (!kids || kids.hidden) await activate(row);
cur = joinPath(cur, part);
}
const target = rowFor(cur);
if (target && target !== selectedNode) target.click();
if (target && target !== selectedNode) await activate(target);
}
// ?file=<url>&path=<object> opens a file from a URL (same origin, or one
// that allows CORS) and selects an object in it.
// that allows CORS) by range requests and selects an object in it.
const params = new URLSearchParams(location.search);
if (params.get("file")) {
const url = params.get("file");
try {
const resp = await fetch(url);
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
const blob = await resp.blob();
await load(new File([blob], url.split("/").pop()));
if (file && params.get("path")) reveal(params.get("path"));
} catch (e) {
$("detail").replaceChildren(el("p", { className: "error", textContent: `Cannot fetch ${url}: ${e.message}` }));
}
await loadUrl(params.get("file"));
if (file && params.get("path")) await reveal(params.get("path"));
document.body.dataset.ready = "1";
}
$("urlform").addEventListener("submit", (e) => {
e.preventDefault();
const url = $("url").value.trim();
if (url) loadUrl(url);
});
$("picker").addEventListener("change", (e) => e.target.files[0] && load(e.target.files[0]));
document.addEventListener("dragover", (e) => { e.preventDefault(); document.body.classList.add("dragging"); });
document.addEventListener("dragleave", () => document.body.classList.remove("dragging"));
+60 -30
View File
@@ -3,10 +3,13 @@
#
# browser.sh FIXTURE_DIR
#
# FIXTURE_DIR holds fixture.h5 from make_fixture.py; ../pkg must be built.
# The page is opened with ?file=fixture.h5&path=<object>, which fetches the
# file, builds the tree down to <object> and shows it; the rendered DOM is
# dumped and checked for the values libhdf5 reads.
# FIXTURE_DIR holds fixture.h5 from make_fixture.py (and big.h5 when it was
# run with WASM_BIG_MB); ../pkg must be built. test/serve.py serves the
# page and, under /fix, the fixtures, with HTTP range requests (no copies,
# no links). The page is opened with ?file=fix/<name>&path=<object>, which
# opens the file with openUrl (range requests), builds the tree down to
# <object> and shows it; the rendered DOM is dumped and checked for the
# values libhdf5 reads and for the request counter.
#
# Browser: $CHROME, else chromium/google-chrome on PATH, else a Playwright
# download under ~/.cache/ms-playwright. Exit 3 when none is found.
@@ -31,63 +34,90 @@ if [ -z "$chrome" ] || [ ! -x "$chrome" ]; then
exit 3
fi
root="$(mktemp -d)"
profiles="$(mktemp -d)"
server=""
cleanup() {
[ -n "$server" ] && kill "$server" 2>/dev/null || true
rm -rf "$root"
rm -rf "$profiles"
}
trap cleanup EXIT
ln -s "$HERE/../index.html" "$HERE/../viewer-lib.js" "$HERE/../pkg" "$FIX/fixture.h5" "$root/"
port=$("$PY" -c 'import socket; s = socket.socket(); s.bind(("127.0.0.1", 0)); print(s.getsockname()[1])')
"$PY" -m http.server --bind 127.0.0.1 --directory "$root" "$port" >/dev/null 2>&1 &
"$PY" "$HERE/serve.py" --root "fix=$FIX" --root "=$HERE/.." > "$profiles/port" &
server=$!
for _ in $(seq 50); do
"$PY" -c "import urllib.request; urllib.request.urlopen('http://127.0.0.1:$port/index.html')" 2>/dev/null && break
[ -s "$profiles/port" ] && break
sleep 0.1
done
port="$(head -1 "$profiles/port")"
fails=0
# A fresh profile per page: a second instance on the same profile fails.
render() {
local profile
profile="$(mktemp -d "$root/profile.XXXXXX")"
profile="$(mktemp -d "$profiles/profile.XXXXXX")"
"$chrome" --headless --no-sandbox --disable-gpu --user-data-dir="$profile" \
--virtual-time-budget=20000 \
--dump-dom "http://127.0.0.1:$port/index.html?file=fixture.h5&path=$1" 2>/dev/null
--dump-dom "http://127.0.0.1:$port/index.html?file=$1&path=$2" 2>/dev/null
}
# expect PATH TEXT...: every TEXT appears in the page rendered for PATH.
# expect FILE PATH TEXT...: every TEXT appears in the page rendered for PATH
# of FILE (a URL relative to the page); "re:TEXT" is an extended regex.
expect() {
local path="$1" dom
shift
dom="$(render "$path")"
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$path" | tr / _).html"
local file="$1" path="$2" dom
shift 2
dom="$(render "$file" "$path")"
[ -n "${BROWSER_DEBUG:-}" ] && printf "%s\n" "$dom" > "$BROWSER_DEBUG.$(echo "$file$path" | tr / _).html"
for text in "$@"; do
if ! grep -qF -- "$text" <<<"$dom"; then
echo "FAIL: page for $path lacks: $text" >&2
local flag=-qF
if [ "${text#re:}" != "$text" ]; then flag=-qE; text="${text#re:}"; fi
if ! grep $flag -- "$text" <<<"$dom"; then
echo "FAIL: page for $file $path lacks: $text" >&2
fails=$((fails + 1))
fi
done
echo "rendered $path"
echo "rendered $file $path"
}
H5=fix/fixture.h5
# Tree (root expanded; the group row carries its path) and root attributes.
expect "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
expect $H5 "/" 'data-path="/sensors"' 'data-path="/grid"' '<td>title</td><td>"wasm fixture"</td>' \
'<td>big</td><td>9223372036854775813</td>' '(compound{x: f64, n: i32})'
# A chunked, deflated 2-D dataset: type, shape, the first window of values.
expect "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
'showing rows 0–5, columns 0–9 of 60 values'
# A chunked, deflated 2-D dataset: type, shape, the first window of values;
# the request counter.
expect $H5 "/grid" '<dd>f64</dd>' '<dd>(6, 10)</dd>' '<th>9</th>' '<td>0.25</td>' '<td>14.75</td>' \
'showing rows 0–5, columns 0–9 of 60 values' 'request' 'fetched (100.0%)'
# Nested path revealed through the tree; big-endian float32.
expect "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
expect $H5 "/sensors/temp" 'data-path="/sensors/temp"' '<dd>f32</dd>' '<td>21.5</td>' '<td>22.25</td>'
# 64-bit integers stay exact; strings; array datatype cells.
expect "/u64" '<td>18446744073709551615</td>'
expect "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
expect $H5 "/u64" '<td>18446744073709551615</td>'
expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
# 3-D: leading dimension held at 0, window over the last two.
expect "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
# Unsupported type: an error, not values.
expect "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
# exposes Content-Range, and CORS that exposes nothing (/unexposed/: the
# browser hides Content-Range and ETag, so the page learns the length with
# a HEAD request, which the server's log must show).
expect "http://localhost:$port/$H5" "/grid" '<td>14.75</td>' 'request'
get() { "$PY" -c 'import sys, urllib.request; print(urllib.request.urlopen(sys.argv[1]).read().decode())' "$1"; }
get "http://127.0.0.1:$port/__reset" >/dev/null
expect "http://localhost:$port/unexposed/$H5" "/grid" '<td>0.25</td>' '<td>14.75</td>' 'request'
heads="$(get "http://127.0.0.1:$port/__stats" | "$PY" -c \
'import json, sys; print(sum(1 for l in json.load(sys.stdin)["log"] if l[4] == "HEAD" and l[0].endswith("fixture.h5")))')"
if [ "$heads" -lt 1 ]; then
echo "FAIL: no HEAD request for the length with unexposed headers" >&2
fails=$((fails + 1))
fi
# A server without range support: the file is downloaded whole, and says so.
expect norange/$H5 "/grid" '<td>14.75</td>' 'downloaded whole'
# A large file: a small dataset, and a window of the large one, fetch a few
# blocks (the counter shows a small percentage: "(1.5%)", not "(100.0%)").
if [ -f "$FIX/big.h5" ]; then
expect fix/big.h5 "/small" '<td>1.5</td>' '<td>3.25</td>' 're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
expect fix/big.h5 "/big" '<dd>(25000000)</dd>' '<td>0</td>' '<td>24.5</td>' \
're:MiB of 191 MiB fetched \([0-9]\.[0-9]%\)'
fi
if [ "$fails" -gt 0 ]; then
echo "browser: $fails checks failed" >&2
+106 -1
View File
@@ -3,7 +3,9 @@ reads back from them, for the clawhdf5-wasm tests.
python make_fixture.py OUT_DIR
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json.
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json,
and the limit-test files OUT_DIR/limits.h5 and OUT_DIR/hostile_vl.h5 (see
write_limits).
Both the Rust test (crates/clawhdf5-wasm/tests/h5py_interop.rs, native) and
the Node test (test.mjs, the built wasm package) compare against the same
expected.json, so the two check the same values.
@@ -14,6 +16,7 @@ encoded as strings so JSON.parse keeps 64-bit values exact.
"""
import json
import os
import sys
import warnings
from pathlib import Path
@@ -201,3 +204,105 @@ def slab_for(obj):
json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)},
open(out / "expected.json", "w"), indent=1, ensure_ascii=False)
HUGE_U8 = 2**28 + 1024
# The collection size hostile_vl.h5 claims: past 2 GiB, which a wasm32
# buffer cannot hold.
HOSTILE_GCOL_SIZE = 2**31 + 4096
# The file length a server claims for hostile_vl.h5 (the tests' mock fetch
# answers every range with zeros past the real bytes): 3 GiB, within what
# wasm32 opens, and room for the collection.
HOSTILE_LENGTH = 3 << 30
def write_limits(out):
"""Files for the size limits (the tests must get errors, not aborts):
- limits.h5: /huge_u8, 2^28 + 1024 bytes of u8 in compressed chunks
(a small file): read whole it would take over 2 GiB while decoding;
its last value is 7.
- hostile_vl.h5: a variable-length string dataset /a whose global heap
collection claims HOSTILE_GCOL_SIZE bytes, with the superblock's end
of file set to HOSTILE_LENGTH (libhdf5 cannot read it; it is only
served by a mock that claims that length).
- far.h5 and far.json: /x, 16 float64 values, whose contiguous data
address is moved FAR_SHIFT bytes on (past 2 GiB, the sign bit of a
wasm32 isize) in a file whose end of file is moved as far; the tests'
mock serves the data there, to show offsets up to 4 GiB work on
wasm32.
"""
with h5py.File(out / "limits.h5", "w") as f:
d = f.create_dataset("huge_u8", shape=(HUGE_U8,), dtype="u1",
chunks=(1 << 20,), compression="gzip")
d[-1] = 7
path = out / "hostile_vl.h5"
with h5py.File(path, "w", libver="earliest") as f:
f.create_dataset("a", data=["x", "yy"], dtype=h5py.string_dtype())
b = bytearray(path.read_bytes())
assert b[8] == 0, "a version 0 superblock"
b[40:48] = HOSTILE_LENGTH.to_bytes(8, "little") # end of file address
at = b.index(b"GCOL")
b[at + 8:at + 16] = HOSTILE_GCOL_SIZE.to_bytes(8, "little")
path.write_bytes(bytes(b))
path = out / "far.h5"
values = np.arange(16, dtype="<f8") * 1.5
with h5py.File(path, "w", libver="earliest") as f:
f.create_dataset("x", data=values)
data_at = f["x"].id.get_offset()
b = bytearray(path.read_bytes())
# The layout message: the data's address, then its size.
old = data_at.to_bytes(8, "little") + (values.nbytes).to_bytes(8, "little")
assert b.count(old) == 1
at = b.index(old)
b[at:at + 8] = (data_at + FAR_SHIFT).to_bytes(8, "little")
length = len(b) + FAR_SHIFT
b[40:48] = length.to_bytes(8, "little")
path.write_bytes(bytes(b))
json.dump({"data_at": data_at, "far_at": data_at + FAR_SHIFT,
"nbytes": values.nbytes, "length": length,
"values": [float(x) for x in values]},
open(out / "far.json", "w"))
# far.h5's data moves this far: past 2 GiB, below 4 GiB.
FAR_SHIFT = 3 << 30
write_limits(out)
def write_big(path, megabytes):
"""A large file for the range-request tests (`openUrl`): `/big`, about
`megabytes` MB of float64 in 1 MiB chunks, written after a small
dataset and a group, so listing and reading `/small` touch a few blocks
of the file and a window of `/big` one chunk. Returns what h5py reads
back."""
n = megabytes * 1_000_000 // 8
chunk = 1 << 17
with h5py.File(path, "w") as f:
f.attrs["note"] = "large file for range reads"
f.create_dataset("small", data=np.array([1.5, -2.0, 3.25]))
g = f.create_group("meta")
g.attrs["units"] = "m"
g.create_dataset("ids", data=np.arange(10, dtype="<i4"))
big = f.create_dataset("big", shape=(n,), dtype="<f8", chunks=(chunk,))
for s in range(0, n, 1 << 22):
e = min(n, s + (1 << 22))
big[s:e] = np.arange(s, e, dtype="<f8") * 0.5
start = n // 2 + 12_345
with h5py.File(path, "r") as f:
return {
"size": path.stat().st_size,
"list": {"groups": ["meta"], "datasets": ["big", "small"]},
"small": [float(x) for x in f["small"][()]],
"ids": [str(int(x)) for x in f["meta/ids"][()]],
"big_shape": list(f["big"].shape),
"window": {"start": start, "count": 10,
"values": [float(x) for x in f["big"][start:start + 10]]},
}
big_mb = int(os.environ.get("WASM_BIG_MB", "0"))
if big_mb > 0:
json.dump(write_big(out / "big.h5", big_mb), open(out / "big.json", "w"), indent=1)
+26 -5
View File
@@ -1,11 +1,17 @@
#!/usr/bin/env bash
# Build the wasm package (../build.sh), test it under Node against files
# written by h5py and netCDF4 (make_fixture.py), then load the viewer page
# in headless Chromium if one is found (browser.sh).
# written by h5py and netCDF4 (make_fixture.py) — opened from bytes, and
# opened by URL from a local range-capable HTTP server (serve.py) — then
# load the viewer page in headless Chromium if one is found (browser.sh).
#
# Needs node, the wasm-bindgen CLI (see ../build.sh) and a Python with h5py,
# netCDF4 and numpy: CLAWHDF5_PYTHON names it (default python3). Without that
# Python the test is skipped, unless CLAWHDF5_REQUIRE_INTEROP=1.
#
# WASM_BIG_MB (default 200) sizes the large file of the range-request
# budget test (0 leaves it out); it is written under TMPDIR.
# CLAWHDF5_WASM_CORPUS=DIR also compares every HDF5 file under DIR (up to
# 16 MiB) read over HTTP with the same file read from bytes.
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
@@ -23,9 +29,24 @@ fi
bash "$HERE/../build.sh"
fix="$(mktemp -d)"
trap 'rm -rf "$fix"' EXIT
"$PY" "$HERE/make_fixture.py" "$fix"
node "$HERE/test.mjs" "$HERE/../pkg" "$fix"
server=""
cleanup() {
[ -n "$server" ] && kill "$server" 2>/dev/null || true
rm -rf "$fix"
}
trap cleanup EXIT
WASM_BIG_MB="${WASM_BIG_MB:-200}" "$PY" "$HERE/make_fixture.py" "$fix"
# The fixtures (and the corpus) over HTTP with range support.
roots=(--root "fix=$fix")
[ -n "${CLAWHDF5_WASM_CORPUS:-}" ] && roots+=(--root "corpus=$CLAWHDF5_WASM_CORPUS")
"$PY" "$HERE/serve.py" "${roots[@]}" > "$fix/port" &
server=$!
for _ in $(seq 50); do
[ -s "$fix/port" ] && break
sleep 0.1
done
node "$HERE/test.mjs" "$HERE/../pkg" "$fix" "http://127.0.0.1:$(head -1 "$fix/port")"
# The page itself, in headless Chromium when one is available.
status=0
+223
View File
@@ -0,0 +1,223 @@
"""A static HTTP server for the wasm tests, with HTTP Range support and
request counting.
python serve.py [--root PREFIX=DIR ...]
Serves each DIR under URL PREFIX (the first match wins; PREFIX "" is the
site root), prints the port on its first line of stdout, and runs until
killed. No symlinks or copies are made: files are read where they are.
- `Range: bytes=a-b`, `bytes=a-` and `bytes=-n` get 206 with Content-Range,
an unsatisfiable range 416; every file answer carries an ETag, and CORS
headers exposing Content-Range, so a page on another origin can use it.
- Under `/norange/...` the same files are served but Range is ignored
(200 with the whole file), as by a server without range support.
- Under `/noexpose/...` ranges are served without Content-Range, ETag,
Last-Modified or Accept-Ranges, and CORS exposes none of them: what a
page sees of a cross-origin server that does not list them in
Access-Control-Expose-Headers. The length comes from Content-Length of
a HEAD request (always readable).
- Under `/unexposed/...` ranges are served with all those headers, but
CORS exposes none of them: a browser page on another origin cannot read
them (test/browser.sh loads the page from 127.0.0.1 and the file from
localhost), so it has to take the same path.
- `GET /__stats` returns `{"requests": n, "bytes": n, "log": [...]}` for
file requests since the last `GET /__reset`, which zeroes them.
"""
import argparse
import hashlib
import json
import os
import posixpath
import sys
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import unquote, urlsplit
TYPES = {
".html": "text/html; charset=utf-8",
".js": "text/javascript; charset=utf-8",
".mjs": "text/javascript; charset=utf-8",
".wasm": "application/wasm",
".json": "application/json",
".ts": "text/plain; charset=utf-8",
}
lock = threading.Lock()
stats = {"requests": 0, "bytes": 0, "log": []}
def resolve(roots, path):
"""The file for URL `path`, or None. `..` never leaves a root."""
parts = [p for p in posixpath.normpath(unquote(path)).split("/") if p]
if any(p in (".", "..") for p in parts):
return None
for prefix, root in roots:
pre = [p for p in prefix.split("/") if p]
if parts[: len(pre)] == pre:
rest = parts[len(pre):] or ["index.html"]
f = os.path.join(root, *rest)
if os.path.isfile(f):
return f
return None
def parse_range(header, size):
"""(start, end exclusive) for a single `bytes=` range, "bad" when
unsatisfiable, None when absent or unparsable (served whole)."""
if not header or not header.startswith("bytes=") or "," in header:
return None
a, _, b = header[len("bytes="):].strip().partition("-")
try:
if a == "":
n = int(b)
return (max(0, size - n), size) if n > 0 and size > 0 else "bad"
start = int(a)
end = int(b) + 1 if b else size
except ValueError:
return None
if start >= size or end <= start:
return "bad"
return start, min(end, size)
def make_handler(roots):
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def log_message(self, *args):
pass
def cors(self, expose=True):
self.send_header("Access-Control-Allow-Origin", "*")
if expose:
self.send_header("Access-Control-Expose-Headers",
"Content-Range, Content-Length, ETag, Accept-Ranges")
def do_OPTIONS(self):
self.send_response(204)
self.cors()
self.send_header("Access-Control-Allow-Headers", "Range")
self.send_header("Content-Length", "0")
self.end_headers()
def do_HEAD(self):
self.serve(head=True)
def do_GET(self):
self.serve(head=False)
def json(self, obj):
body = json.dumps(obj).encode()
self.send_response(200)
self.cors()
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.send_header("Cache-Control", "no-store")
self.end_headers()
self.wfile.write(body)
def serve(self, head):
path = urlsplit(self.path).path
if path == "/__stats":
with lock:
return self.json(stats)
if path == "/__reset":
with lock:
stats.update(requests=0, bytes=0, log=[])
return self.json({})
# ranges: honour Range; send_all: send Content-Range, ETag and
# Accept-Ranges; expose: list them for CORS.
ranges = send_all = expose = True
if path.startswith("/norange/"):
ranges = False
path = path[len("/norange"):]
elif path.startswith("/noexpose/"):
expose = send_all = False
path = path[len("/noexpose"):]
elif path.startswith("/unexposed/"):
expose = False
path = path[len("/unexposed"):]
f = resolve(roots, path)
if f is None:
self.send_response(404)
self.cors()
self.send_header("Content-Length", "0")
self.end_headers()
return
size = os.path.getsize(f)
st = os.stat(f)
etag = '"%s"' % hashlib.sha1(
f"{f}:{size}:{st.st_mtime_ns}".encode()).hexdigest()[:16]
r = parse_range(self.headers.get("Range"), size) if ranges else None
if r == "bad":
self.send_response(416)
self.cors()
self.send_header("Content-Range", f"bytes */{size}")
self.send_header("Content-Length", "0")
self.end_headers()
return
start, end = r if r else (0, size)
self.send_response(206 if r else 200)
self.cors(expose)
ext = os.path.splitext(f)[1]
self.send_header("Content-Type", TYPES.get(ext, "application/octet-stream"))
self.send_header("Content-Length", str(end - start))
self.send_header("Cache-Control", "no-store")
if send_all:
self.send_header("ETag", etag)
if ranges:
self.send_header("Accept-Ranges", "bytes")
if r:
self.send_header("Content-Range", f"bytes {start}-{end - 1}/{size}")
self.end_headers()
with lock:
# A HEAD is a request too (openUrl makes one when it cannot
# see Content-Range); it sends no bytes.
stats["requests"] += 1
sent = 0 if head else end - start
stats["bytes"] += sent
stats["log"].append([path, start, end, 206 if r else 200, "HEAD" if head else "GET"])
if not head:
with open(f, "rb") as fh:
fh.seek(start)
left = end - start
try:
while left:
buf = fh.read(min(left, 1 << 20))
if not buf:
break
self.wfile.write(buf)
left -= len(buf)
except (BrokenPipeError, ConnectionResetError):
pass
return Handler
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--root", action="append", default=[],
help="PREFIX=DIR: serve DIR under URL PREFIX")
args = ap.parse_args()
roots = []
for spec in args.root:
prefix, _, d = spec.partition("=")
roots.append((prefix, os.path.abspath(d)))
class Server(ThreadingHTTPServer):
def handle_error(self, request, client_address):
# A client that drops a connection (a cancelled download) is
# not an error of the server.
if not isinstance(sys.exc_info()[1], (ConnectionError, TimeoutError)):
super().handle_error(request, client_address)
httpd = Server(("127.0.0.1", 0), make_handler(roots))
httpd.daemon_threads = True
print(httpd.server_address[1], flush=True)
sys.stdout.close()
httpd.serve_forever()
if __name__ == "__main__":
main()
+531 -32
View File
@@ -1,20 +1,31 @@
// Node test of the built wasm package (the exact pkg/ the viewer page loads)
// and the viewer's DOM-free helpers. Run by test/run.sh:
// node test.mjs PKG_DIR FIXTURE_DIR
// node test.mjs PKG_DIR FIXTURE_DIR [SERVER_URL]
// FIXTURE_DIR holds fixture.h5, fixture.nc and expected.json from
// make_fixture.py (values as libhdf5 reads them back).
// make_fixture.py (values as libhdf5 reads them back), and big.h5/big.json
// when it was run with WASM_BIG_MB. SERVER_URL is test/serve.py serving
// FIXTURE_DIR under /fix (and $CLAWHDF5_WASM_CORPUS under /corpus): with
// it, every check is repeated on files opened with openUrl (HTTP range
// requests), and the request budget, the full-download fallback and the
// error paths of openUrl are tested.
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import { join } from "node:path";
import { existsSync, readFileSync, readdirSync, statSync } from "node:fs";
import { join, relative } from "node:path";
import { pathToFileURL } from "node:url";
const [pkgDir, fixDir] = process.argv.slice(2);
const [pkgDir, fixDir, base] = process.argv.slice(2);
const pkg = await import(pathToFileURL(join(pkgDir, "clawhdf5_wasm.js")));
pkg.initSync({ module: readFileSync(join(pkgDir, "clawhdf5_wasm_bg.wasm")) });
const lib = await import(pathToFileURL(join(import.meta.dirname, "..", "viewer-lib.js")));
let checks = 0;
const eq = (a, b, msg) => { assert.deepEqual(a, b, msg); checks++; };
// A call that must fail: a thrown Error (open) or a rejected promise
// (openUrl), whose message matches `re`.
const fails = async (fn, re, msg) => {
await assert.rejects(async () => fn(), (e) => e instanceof Error && re.test(e.message), msg);
checks++;
};
const ARRAY_TYPES = {
f32: Float32Array, f64: Float64Array, i8: Int8Array, i16: Int16Array, i32: Int32Array,
@@ -54,13 +65,13 @@ function checkAttr(ctx, a, want) {
assert.fail(`${ctx}: unknown expectation ${JSON.stringify(want)}`);
}
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
for (const [name, exp] of Object.entries(expected)) {
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
// Every listing, dataset, error and attribute of `file` against what
// libhdf5 reads (expected.json). `file` is an H5File (synchronous methods)
// or a RemoteFile (promises): every call is awaited.
async function checkFile(name, exp, file) {
for (const [path, want] of Object.entries(exp.lists)) {
eq(file.kind(path), "group", `${name}:${path} kind`);
const list = file.list(path);
eq(await file.kind(path), "group", `${name}:${path} kind`);
const list = await file.list(path);
for (const [kind, key] of [["group", "groups"], ["dataset", "datasets"]]) {
eq(list.filter((c) => c.kind === kind).map((c) => c.name).sort(), want[key], `${name}:${path} ${key}`);
}
@@ -68,58 +79,67 @@ for (const [name, exp] of Object.entries(expected)) {
for (const [path, want] of Object.entries(exp.datasets)) {
const ctx = `${name}:${path}`;
eq(file.kind(path), "dataset", `${ctx} kind`);
eq(await file.kind(path), "dataset", `${ctx} kind`);
if (want.unavailable) {
// The wasm build has no zstd (it links C): a clear error, no data.
assert.throws(() => file.read(path), (e) => e.message.includes(want.unavailable), ctx);
checks++;
await fails(() => file.read(path), new RegExp(want.unavailable), ctx);
continue;
}
const info = file.info(path);
const info = await file.info(path);
eq([...info.shape, ...info.elementShape], want.shape, `${ctx} info shape`);
const r = file.read(path);
const r = await file.read(path);
eq(r.shape, want.shape, `${ctx} shape`);
eq(r.dtype, info.dtype, `${ctx} dtype`);
assert.ok(r.data instanceof ARRAY_TYPES[want.kind], `${ctx}: ${r.data.constructor.name} for ${want.kind}`);
eq(values(want.kind, r.data), want.values, ctx);
if (want.slab) {
const s = want.slab;
const part = file.readHyperslab(path, s.start, s.count, s.stride);
const part = await file.readHyperslab(path, s.start, s.count, s.stride);
eq(part.shape, s.shape, `${ctx} slab shape`);
eq(values(want.kind, part.data), s.values, `${ctx} slab`);
}
}
for (const [path, what] of Object.entries(exp.errors)) {
assert.throws(() => file.read(path), (e) => e instanceof Error && e.message.includes(what), `${name}:${path}`);
checks++;
await fails(() => file.read(path), new RegExp(what), `${name}:${path}`);
}
for (const [path, want] of Object.entries(exp.attrs)) {
const attrs = file.attrs(path);
eq(file.attrErrors(path), [], `${name}:${path} attr errors`);
const attrs = await file.attrs(path);
eq(await file.attrErrors(path), [], `${name}:${path} attr errors`);
const seen = attrs.filter((a) => !a.name.startsWith("_") && !exp.skip_attrs.includes(a.name));
eq(seen.map((a) => a.name).sort(), Object.keys(want).sort(), `${name}:${path} attr names`);
for (const a of seen) checkAttr(`${name}:${path}@${a.name}`, a, want[a.name]);
}
}
// The error paths of the reader, for an H5File or a RemoteFile of
// fixture.h5.
async function checkErrors(h5) {
await fails(() => h5.read("/nope"), /./);
await fails(() => h5.list("/grid"), /not a group/);
await fails(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
await fails(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
await fails(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
await fails(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
// Big integers stay exact.
eq((await h5.read("/u64")).data[0], 18446744073709551615n, "u64 max");
}
const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8"));
for (const [name, exp] of Object.entries(expected)) {
const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name))));
await checkFile(name, exp, file);
file.free();
}
// Errors reach JavaScript as thrown Errors, never as data.
const h5 = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.h5"))));
const throwsMsg = (fn, re) => { assert.throws(fn, (e) => e instanceof Error && re.test(e.message)); checks++; };
throwsMsg(() => pkg.open(new Uint8Array(64)), /./);
throwsMsg(() => h5.read("/nope"), /./);
throwsMsg(() => h5.list("/grid"), /not a group/);
throwsMsg(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/);
throwsMsg(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/);
throwsMsg(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/);
throwsMsg(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/);
await fails(() => pkg.open(new Uint8Array(64)), /./);
await checkErrors(h5);
// Info for a dataset with an unlimited dimension (netCDF "time").
const nc = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.nc"))));
eq(nc.info("/time").maxshape, [null], "unlimited dimension is null");
// Big integers stay exact.
eq(h5.read("/u64").data[0], 18446744073709551615n, "u64 max");
eq(typeof pkg.version(), "string", "version");
// Viewer helpers.
@@ -137,7 +157,486 @@ eq(lib.toRows(pairs.data, 2, 1, lib.perElement([2])), [["[2, 3]"], ["[4, 5]"]],
eq(lib.formatValue(0.1 + 0.2), "0.3", "float formatting");
eq(lib.formatValue(2n ** 64n - 1n), "18446744073709551615", "bigint formatting");
eq(lib.formatValue("x"), '"x"', "string formatting");
eq(lib.formatBytes(0), "0 B", "bytes");
eq(lib.formatBytes(1536), "1.5 KiB", "KiB");
eq(lib.formatBytes(200 * 1024 * 1024), "200 MiB", "MiB");
eq(lib.formatStats({ lazy: true, requests: 3, bytesFetched: 2 << 20, size: 200 << 20 }),
"3 requests, 2 MiB of 200 MiB fetched (1.0%)", "stats line");
eq(lib.formatStats({ lazy: false, requests: 1, bytesFetched: 1024, size: 1024 }),
"downloaded whole (1 KiB): the server does not support range requests", "stats line, no ranges");
h5.free();
nc.free();
console.log(`wasm package: ${checks} checks passed`);
if (base) await remoteTests();
async function serverStats() {
return (await fetch(`${base}/__stats`)).json();
}
async function remoteTests() {
const before = checks;
await fetch(`${base}/__reset`);
// The same checks over HTTP range requests, at the default block size
// and at 512-byte blocks with a 4 KiB cache (almost every structure read
// a miss, evictions between calls).
for (const opts of [undefined, { blockSize: 512, cacheSize: 4096 }]) {
for (const [name, exp] of Object.entries(expected)) {
const f = await pkg.openUrl(`${base}/fix/${name}`, opts);
await checkFile(`${name} (openUrl ${JSON.stringify(opts ?? {})})`, exp, f);
const st = f.stats();
eq(st.lazy, true, "read by ranges");
eq(st.size, statSync(join(fixDir, name)).size, "size");
f.free();
}
}
const remote = await pkg.openUrl(`${base}/fix/fixture.h5`);
await checkErrors(remote);
eq((await (await pkg.openUrl(`${base}/fix/fixture.nc`)).info("/time")).maxshape, [null], "remote unlimited");
// What the page counts is what the server served.
await fetch(`${base}/__reset`);
const counted = await pkg.openUrl(`${base}/fix/fixture.nc`, { blockSize: 1024 });
await counted.read("/temp");
const server = await serverStats();
eq(counted.stats().requests, server.requests, "requests counted");
eq(counted.stats().bytesFetched, server.bytes, "bytes counted");
// A custom fetch is used for every request, with the caller's headers,
// given in any form fetch takes (a caller's Range is not sent).
for (const headers of [{ "X-Test": "1" }, new Headers({ "X-Test": "1", Range: "bytes=0-0" }), [["X-Test", "1"]]]) {
let calls = 0;
const seen = new Set();
const viaCustom = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 4096,
headers,
fetch: (url, init) => {
calls++;
const h = new Headers(init.headers);
seen.add(`${h.get("X-Test")} ${h.get("Range").startsWith("bytes=0-0") ? "caller's range" : "ours"}`);
return fetch(url, init);
},
});
const kind = headers.constructor.name;
eq((await viaCustom.read("/sensors/temp")).data[0], 21.5, `custom fetch values (${kind})`);
eq(calls, viaCustom.stats().requests, `custom fetch calls (${kind})`);
eq([...seen], ["1 ours"], `headers passed (${kind})`);
}
// parallel must be a positive integer.
for (const parallel of [0, -1, 1.5, "4", NaN, 5000]) {
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { parallel }), /parallel/, `parallel ${String(parallel)}`);
}
// When one range request fails, the others in flight are aborted and no
// more start (fetchRanges, called directly: remote.js is in the package).
{
const snippets = join(pkgDir, "snippets");
const dir = readdirSync(snippets).find((d) => existsSync(join(snippets, d, "js", "remote.js")));
const remote = await import(pathToFileURL(join(snippets, dir, "js", "remote.js")));
let started = 0;
const aborted = [];
const slowFetch = (url, init) => {
const k = started++;
if (k === 2) return Promise.resolve(new Response(null, { status: 500 }));
return new Promise((resolve, reject) => {
const t = setTimeout(() => resolve(new Response(new Uint8Array(10), { status: 206 })), 200);
init.signal?.addEventListener("abort", () => {
clearTimeout(t);
aborted.push(k);
reject(new DOMException("aborted", "AbortError"));
});
});
};
const ranges = Array.from({ length: 20 }, (_, i) => [i * 10, i * 10 + 10]).flat();
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: 3 }, null, 200),
/HTTP 500/, "a failed range request is the error");
await new Promise((r) => setTimeout(r, 300));
eq(started, 3, "no request starts after a failure");
eq(aborted.sort(), [0, 1], "requests in flight are aborted");
await fails(() => remote.fetchRanges("http://x.invalid/f.h5", ranges, { fetch: slowFetch, parallel: "x" }, null, 200),
/parallel must be a positive integer/, "fetchRanges checks parallel");
}
// Calls in flight at once share the cache (a 1 KiB budget: nothing is
// evicted while any of them runs) and each gets its own answer.
const both = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, cacheSize: 1024 });
const [g, t, l, a] = await Promise.all([
both.read("/grid"), both.read("/sensors/temp"), both.list("/sensors"), both.attrs("/"),
]);
const exp = expected["fixture.h5"];
eq(Array.from(g.data), exp.datasets["/grid"].values, "concurrent /grid");
eq(Array.from(t.data), exp.datasets["/sensors/temp"].values, "concurrent /sensors/temp");
eq(l.map((c) => c.name).sort(), [...exp.lists["/sensors"].groups, ...exp.lists["/sensors"].datasets].sort(), "concurrent list");
eq(a.length > 0, true, "concurrent attrs");
assert.ok(both.stats().cachedBytes <= 1024, "trimmed to the budget when idle");
// Listing and reading small things of a large file fetches a few blocks,
// not the file.
if (existsSync(join(fixDir, "big.json"))) {
const big = JSON.parse(readFileSync(join(fixDir, "big.json"), "utf8"));
await fetch(`${base}/__reset`);
const f = await pkg.openUrl(`${base}/fix/big.h5`);
const list = await f.list("/");
eq(list.filter((c) => c.kind === "group").map((c) => c.name), big.list.groups, "big: groups");
eq(list.filter((c) => c.kind === "dataset").map((c) => c.name).sort(), big.list.datasets, "big: datasets");
eq(Array.from((await f.read("/small")).data), big.small, "big: /small");
eq(Array.from((await f.read("/meta/ids")).data, String), big.ids, "big: /meta/ids");
eq((await f.info("/big")).shape, big.big_shape, "big: shape");
eq((await f.attrs("/meta"))[0].value, "m", "big: group attribute");
const win = await f.readHyperslab("/big", [big.window.start], [big.window.count]);
eq(Array.from(win.data), big.window.values, "big: window");
const st = f.stats();
const server = await serverStats();
eq(st.requests, server.requests, "big: requests counted");
eq(st.bytesFetched, server.bytes, "big: bytes counted");
console.log(`big.h5 (${big.size} bytes): listed, 2 small reads, attributes, info and a window in ` +
`${server.requests} requests, ${server.bytes} bytes (${(100 * server.bytes / big.size).toFixed(2)}%)`);
assert.ok(server.requests <= 8, `big: ${server.requests} requests`);
assert.ok(server.bytes * 20 < big.size, `big: ${server.bytes} bytes fetched`);
checks += 2;
}
// A cross-origin server that does not expose Content-Range, ETag or
// Last-Modified (serve.py's /noexpose/): the length comes from a HEAD
// request and answers are checked by their length alone. Every fixture
// check, calls in flight at once, and what the page counts.
for (const opts of [undefined, { blockSize: 512, cacheSize: 1024 }]) {
for (const [name, exp] of Object.entries(expected)) {
await fetch(`${base}/__reset`);
const f = await pkg.openUrl(`${base}/noexpose/fix/${name}`, opts);
await checkFile(`${name} (no exposed headers, ${JSON.stringify(opts ?? {})})`, exp, f);
const st = f.stats();
eq(st.lazy, true, "no exposed headers: read by ranges");
eq(st.size, statSync(join(fixDir, name)).size, "no exposed headers: size from HEAD");
const server = await serverStats();
eq(server.log.filter((l) => l[4] === "HEAD").length, 1, "no exposed headers: one HEAD");
eq(st.requests, server.requests, "no exposed headers: requests counted");
eq(st.bytesFetched, server.bytes, "no exposed headers: bytes counted");
f.free();
}
}
{
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, { blockSize: 512, cacheSize: 0, parallel: 3 });
const exp = expected["fixture.h5"];
const paths = ["/grid", "/sensors/temp", "/vlen_str", "/cube"];
const got = await Promise.all([...paths, ...paths].map((p) => f.read(p)));
got.forEach((r, i) => eq(values(exp.datasets[paths[i % 4]].kind, r.data), exp.datasets[paths[i % 4]].values,
`no exposed headers: concurrent ${paths[i % 4]}`));
await checkErrors(f);
}
// Without a validator a changed file cannot be told apart; a short
// answer still can.
await fails(async () => {
const f = await pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
blockSize: 512,
fetch: async (url, init) => {
const r = await fetch(url, init);
if (init.method === "HEAD" || init.headers.Range === "bytes=0-511") return r;
return new Response((await r.arrayBuffer()).slice(1), { status: 206 });
},
});
await f.read("/grid");
}, /got \d+/, "no exposed headers: short answer");
// No HEAD length either: a clear error.
await fails(() => pkg.openUrl(`${base}/noexpose/fix/fixture.h5`, {
fetch: async (url, init) => (init.method === "HEAD" ? new Response(null, { status: 405 }) : fetch(url, init)),
}), /cannot learn the file's size/, "no exposed headers, no HEAD");
// A server without range support: downloaded whole (the default), or
// refused.
const whole = await pkg.openUrl(`${base}/norange/fix/fixture.h5`);
await checkFile("fixture.h5 (no range support)", expected["fixture.h5"], whole);
eq(whole.stats().lazy, false, "downloaded whole");
eq(whole.stats().requests, 1, "one request");
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fallback: "error" }),
/does not support HTTP range requests/, "fallback: error");
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000 }),
/more than maxDownload/, "maxDownload");
// Without a Content-Length the body is streamed, and stopped at the limit.
const undeclared = async (url, init) => {
const r = await fetch(url, init);
return new Response(r.body, { status: r.status });
};
await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000, fetch: undeclared }),
/more than maxDownload/, "maxDownload, streamed");
const streamed = await pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fetch: undeclared });
eq(Array.from((await streamed.read("/sensors/temp")).data), [21.5, 22, 22.25], "streamed download");
// Errors: HTTP status, not HDF5, bad options, a file that changes, a
// server that answers with the wrong bytes.
await fails(() => pkg.openUrl(`${base}/fix/missing.h5`), /HTTP 404/, "404");
await fails(() => pkg.openUrl(`${base}/fix/expected.json`), /./, "not HDF5");
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 100 }), /blockSize/, "blockSize");
const tamper = (edit) => async (url, init) => {
const r = await fetch(url, init);
return init.headers.Range === "bytes=0-511" ? r : edit(r);
};
const withHeaders = async (r, headers) => {
const h = new Headers(r.headers);
for (const [k, v] of Object.entries(headers)) h.set(k, v);
return new Response(await r.arrayBuffer(), { status: r.status, headers: h });
};
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, fetch: tamper((r) => withHeaders(r, { ETag: '"other"' })) });
await f.read("/grid");
}, /changed on the server/, "changed file");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => new Response((await r.arrayBuffer()).slice(1), { status: 206 })),
});
await f.read("/grid");
}, /got \d+/, "short answer");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => new Response(await r.arrayBuffer(), { status: 200 })),
});
await f.read("/grid");
}, /stopped honouring range requests/, "200 mid-file");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
});
await f.read("/grid");
}, /the server sent 0-511/, "wrong range");
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes */25752" })),
});
await f.read("/grid");
}, /unusable Content-Range/, "unusable Content-Range");
await limitTests();
await floodTests();
// Corpus files: what the viewer can show of each is the same read by
// ranges as in memory (an error wherever it gives one).
const corpus = process.env.CLAWHDF5_WASM_CORPUS;
if (corpus) await corpusTests(corpus);
console.log(`openUrl: ${checks - before} checks passed`);
}
// A fetch that serves `buf` as a file of `total` bytes (zeros past the end
// of `buf`, and `extra` = [[offset, bytes], ...] laid over them), counting
// its calls. `hide` leaves out Content-Range and the validators, as a
// cross-origin server that exposes neither does; HEAD then gives the length.
function mockFetch(buf, { total = buf.length, extra = [], hide = false } = {}) {
const f = async (url, init) => {
f.calls++;
if (init.method === "HEAD") {
return new Response(null, { status: 200, headers: { "Content-Length": String(total) } });
}
const m = /^bytes=(\d+)-(\d+)$/.exec(new Headers(init.headers).get("Range"));
const a = Number(m[1]);
const b = Math.min(Number(m[2]) + 1, total);
const out = new Uint8Array(b - a);
for (const [at, bytes] of [[0, buf], ...extra]) {
const from = Math.max(a, at);
const to = Math.min(b, at + bytes.length);
if (from < to) out.set(bytes.subarray(from - at, to - at), from - a);
}
const headers = { "Content-Length": String(b - a) };
if (!hide) headers["Content-Range"] = `bytes ${a}-${b - 1}/${total}`;
return new Response(out, { status: 206, headers });
};
f.calls = 0;
return f;
}
// Sizes a hostile server or a large dataset can name are errors, never an
// allocation that aborts the module (which would take every open file on
// the page with it); see write_limits in make_fixture.py.
async function limitTests() {
// Read whole, /huge_u8 (2^28 + 1024 bytes) would take over 2 GiB while
// decoding: refused before its chunks are fetched; a window reads.
const n = 2 ** 28 + 1024;
const limits = readFileSync(join(fixDir, "limits.h5"));
const local = pkg.open(new Uint8Array(limits));
await fails(() => local.read("/huge_u8"), /readHyperslab/, "huge read, in memory");
const remote = await pkg.openUrl(`${base}/fix/limits.h5`);
const before = remote.stats().requests;
await fails(() => remote.read("/huge_u8"), /readHyperslab/, "huge read, openUrl");
eq(remote.stats().requests, before, "huge read fetched nothing");
eq(Array.from((await remote.readHyperslab("/huge_u8", [n - 4], [4])).data), [0, 0, 0, 7], "huge window");
eq(Array.from(local.readHyperslab("/huge_u8", [n - 4], [4]).data), [0, 0, 0, 7], "huge window, in memory");
local.free();
// A server claiming 3 GiB, a heap collection claiming 2 GiB + 4 KiB: an
// error after a few requests (it used to fetch 2 GiB, then abort).
const hostile = mockFetch(new Uint8Array(readFileSync(join(fixDir, "hostile_vl.h5"))), { total: 3 * 2 ** 30 });
const h = await pkg.openUrl("http://hostile.invalid/h.h5", { fetch: hostile });
await fails(() => h.read("/a"), /maxFetch/, "hostile collection size");
assert.ok(hostile.calls <= 4, `hostile: ${hostile.calls} requests`);
// The module survived: files open and read.
eq((await (await pkg.openUrl(`${base}/fix/fixture.h5`)).read("/sensors/temp")).data[0], 21.5, "alive after hostile");
// A call that would fetch more than maxFetch fails before fetching it.
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, maxFetch: 1024 });
await f.read("/grid");
}, /maxFetch/, "maxFetch");
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxFetch: 2 ** 31 }), /maxFetch/, "maxFetch range");
await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { maxDownload: 2 ** 31 }), /maxDownload/, "maxDownload range");
// Offsets past 2 GiB work on wasm32: far.h5's data sits at 3 GiB.
const far = JSON.parse(readFileSync(join(fixDir, "far.json"), "utf8"));
const farBytes = new Uint8Array(readFileSync(join(fixDir, "far.h5")));
const farData = farBytes.subarray(far.data_at, far.data_at + far.nbytes);
const ff = await pkg.openUrl("http://far.invalid/far.h5", {
fetch: mockFetch(farBytes, { total: far.length, extra: [[far.far_at, farData]] }),
});
eq(Array.from((await ff.read("/x")).data), far.values, "data past 2 GiB");
eq(ff.stats().size, far.length, "size past 2 GiB");
// wasm32's reader holds file offsets in 32 bits: a file of 4 GiB or more
// is refused at open, with the limit in the message (past 2^53 - 1 bytes
// a JavaScript number is not even exact).
for (const total of [2 ** 32, 2 ** 40, 2 ** 53 + 2]) {
await fails(() => pkg.openUrl("http://huge.invalid/x.h5", { fetch: mockFetch(farBytes, { total }) }),
total > 2 ** 53 ? /2\^53 - 1/ : /4 GiB/, `length ${total}`);
}
}
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
// tells how many were.
function flood(total) {
let pulled = 0;
const body = new ReadableStream({
pull(c) {
if (pulled >= total) return c.close();
const n = Math.min(65536, total - pulled);
pulled += n;
c.enqueue(new Uint8Array(n));
},
}, { highWaterMark: 0 });
return { body, pulled: () => pulled };
}
// A 206 whose body runs past the range asked for is cut off as it arrives:
// the page never reads (or holds) more than it asked for, whatever the
// server sends. 64 MiB stands for "gigabytes".
async function floodTests() {
const FLOOD = 64 << 20;
// The probe: asked for the first 1 MiB.
let f = flood(FLOOD);
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
fetch: async () => new Response(f.body, { status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000" } }),
}), /the server sent more/, "flooded probe");
assert.ok(f.pulled() <= (1 << 20) + 65536, `probe: read ${f.pulled()} bytes`);
// A range request after a correct probe.
f = flood(FLOOD);
const file = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: async (url, init) => {
const range = new Headers(init.headers).get("Range");
if (range === "bytes=0-511") return fetch(url, init);
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
},
}).catch((e) => e);
// The open itself may need a second range: flooded either way.
const read = file instanceof Error ? Promise.reject(file) : file.read("/grid");
await fails(() => read, /the server sent more/, "flooded range");
assert.ok(f.pulled() <= 8 * 512 + 65536, `range: read ${f.pulled()} bytes`);
// A declared Content-Length past the range is refused before reading.
f = flood(FLOOD);
await fails(() => pkg.openUrl("http://flood.invalid/x.h5", {
fetch: async () => new Response(f.body, {
status: 206, headers: { "Content-Range": "bytes 0-1048575/2000000", "Content-Length": String(FLOOD) },
}),
}), /the server sent more/, "declared too long");
assert.ok(f.pulled() <= 65536, `declared: read ${f.pulled()} bytes`);
}
function hdf5Files(dir, out) {
for (const e of readdirSync(dir, { withFileTypes: true })) {
const p = join(dir, e.name);
let st;
try {
st = statSync(p);
} catch {
continue;
}
if (st.isDirectory()) hdf5Files(p, out);
else if (st.size <= 16 << 20) {
const b = readFileSync(p);
const sig = [0x89, 0x48, 0x44, 0x46, 0x0d, 0x0a, 0x1a, 0x0a];
for (let at = 0; at + 8 <= b.length; at = at === 0 ? 512 : at * 2) {
if (sig.every((x, i) => b[at + i] === x)) { out.push(p); break; }
}
}
}
return out;
}
function show(v) {
return JSON.stringify(v, (_, x) => {
if (typeof x === "bigint") return `${x}n`;
if (ArrayBuffer.isView(x)) return Array.from(x, (y) => (typeof y === "bigint" ? `${y}n` : Number.isNaN(y) ? "NaN" : y));
return x;
});
}
async function transcript(f) {
const out = [];
const call = async (what, fn) => {
try {
out.push(`${what}: ${show(await fn())}`);
} catch {
out.push(`${what}: Err`);
}
};
const todo = [["/", 0]];
while (todo.length && out.length < 3000) {
const [path, depth] = todo.pop();
let kind = null;
await call(`${path} kind`, async () => (kind = await f.kind(path)));
await call(`${path} attrs`, () => f.attrs(path));
if (kind === "group") {
let list = [];
await call(`${path} list`, async () => (list = await f.list(path)));
if (depth < 12) for (const c of list.reverse()) todo.push([lib.joinPath(path, c.name), depth + 1]);
} else if (kind === "dataset") {
let info = null;
await call(`${path} info`, async () => (info = await f.info(path)));
if (info && [...info.shape, ...info.elementShape].reduce((a, b) => a * b, 1) <= 1 << 20) {
await call(`${path} read`, () => f.read(path));
}
}
}
return out;
}
async function corpusTests(corpus) {
const files = hdf5Files(corpus, []).sort();
let opened = 0;
let bytes = 0;
let size = 0;
for (const p of files) {
const rel = relative(corpus, p).split("/").map(encodeURIComponent).join("/");
let local;
try {
local = pkg.open(new Uint8Array(readFileSync(p)));
} catch {
await fails(() => pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 }), /./, `${rel}: opens in neither`);
continue;
}
const remote = await pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 });
const want = await transcript(local);
const got = await transcript(remote);
// Errors are compared as errors: a malformed file can fail at another
// check, with another message, when read by ranges.
eq(got, want, `${rel}: transcript`);
opened++;
bytes += remote.stats().bytesFetched;
size += remote.stats().size;
local.free();
remote.free();
}
console.log(`corpus: ${opened} of ${files.length} files agree over HTTP (${bytes} of ${size} bytes fetched)`);
}
+21
View File
@@ -75,3 +75,24 @@ export function toRows(data, rows, cols, per = 1) {
export function perElement(elementShape) {
return elementShape.reduce((a, b) => a * b, 1);
}
/** A byte count for people: `1.5 KiB`, `200 MiB`. */
export function formatBytes(n) {
const units = ["B", "KiB", "MiB", "GiB", "TiB"];
let i = 0;
let v = n;
while (v >= 1024 && i < units.length - 1) {
v /= 1024;
i++;
}
const s = i === 0 ? String(v) : v < 10 ? String(Number(v.toFixed(1))) : String(Math.round(v));
return `${s} ${units[i]}`;
}
/** What reading a remote file has cost, from `RemoteFile.stats()`. */
export function formatStats({ lazy, requests, bytesFetched, size }) {
if (!lazy) return `downloaded whole (${formatBytes(size)}): the server does not support range requests`;
const pct = size > 0 ? (100 * bytesFetched) / size : 0;
return `${requests} request${requests === 1 ? "" : "s"}, ${formatBytes(bytesFetched)} of ` +
`${formatBytes(size)} fetched (${pct.toFixed(1)}%)`;
}