wasm: openUrl reads remote files by HTTP range requests (range-read M4)

openUrl(url, opts) returns a RemoteFile with the methods of H5File
(kind, list, info, attrs, attrErrors, read, readHyperslab), each a
promise, and stats(). It runs every call through the restartable
LazyStorage: a pass that misses reports the byte ranges, js/remote.js
fetches them with fetch() and Range headers (six at a time), and the
pass is re-run. This keeps the main thread free without a Worker or
synchronous XHR (h5wasm's lazy files need both), as the design doc
recommends; the cost is re-running a pass per wave of misses.

Every answer is checked: a 206 with exactly the bytes asked for, and
the same ETag/Last-Modified and length as at open, else an error (never
data). A server that ignores Range (200) is downloaded whole, up to
maxDownload (512 MiB), unless fallback: "error". Options: blockSize,
cacheSize, headers, credentials, parallel, fetch.

test/serve.py is a range-capable static server with request counting
(and /norange/ for a server without range support). test.mjs repeats
every fixture check on files opened by URL (1 MiB and 512 B blocks),
checks the request budget on a 200 MB h5py file (list, three small
reads and a window of the big dataset: 5 requests, 6 MiB), the
download fallback, and HTTP errors, changed files and wrong answers;
with CLAWHDF5_WASM_CORPUS every corpus file is compared with open(bytes).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:44:45 -05:00
co-authored by Claude Opus 5.5
parent 076feb089a
commit 5107583b97
8 changed files with 1209 additions and 101 deletions
+37
View File
@@ -14,6 +14,7 @@ encoded as strings so JSON.parse keeps 64-bit values exact.
"""
import json
import os
import sys
import warnings
from pathlib import Path
@@ -201,3 +202,39 @@ def slab_for(obj):
json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)},
open(out / "expected.json", "w"), indent=1, ensure_ascii=False)
def write_big(path, megabytes):
"""A large file for the range-request tests (`openUrl`): `/big`, about
`megabytes` MB of float64 in 1 MiB chunks, written after a small
dataset and a group, so listing and reading `/small` touch a few blocks
of the file and a window of `/big` one chunk. Returns what h5py reads
back."""
n = megabytes * 1_000_000 // 8
chunk = 1 << 17
with h5py.File(path, "w") as f:
f.attrs["note"] = "large file for range reads"
f.create_dataset("small", data=np.array([1.5, -2.0, 3.25]))
g = f.create_group("meta")
g.attrs["units"] = "m"
g.create_dataset("ids", data=np.arange(10, dtype="<i4"))
big = f.create_dataset("big", shape=(n,), dtype="<f8", chunks=(chunk,))
for s in range(0, n, 1 << 22):
e = min(n, s + (1 << 22))
big[s:e] = np.arange(s, e, dtype="<f8") * 0.5
start = n // 2 + 12_345
with h5py.File(path, "r") as f:
return {
"size": path.stat().st_size,
"list": {"groups": ["meta"], "datasets": ["big", "small"]},
"small": [float(x) for x in f["small"][()]],
"ids": [str(int(x)) for x in f["meta/ids"][()]],
"big_shape": list(f["big"].shape),
"window": {"start": start, "count": 10,
"values": [float(x) for x in f["big"][start:start + 10]]},
}
big_mb = int(os.environ.get("WASM_BIG_MB", "0"))
if big_mb > 0:
json.dump(write_big(out / "big.h5", big_mb), open(out / "big.json", "w"), indent=1)