From 13b36e2bd7204ff5c8243095f6faef57bc0b6c08 Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 06:29:48 -0500 Subject: [PATCH 01/30] docs: design for reading files a SWMR writer is appending to (range-read M5) libhdf5's SWMR semantics as they matter to a reader (superblock v3 with the SWMR-write flag and a stale EOF, append-only writer, flush dependencies, refresh, 100 metadata read attempts), what clawhdf5 did with such files, and the plan: data_end bounded by the file length for SWMR-flagged files, a live File::open_swmr over positioned reads without the chunk cache, Dataset::refresh, and bounded whole-operation retries. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/design/swmr.md | 124 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 124 insertions(+) create mode 100644 docs/design/swmr.md diff --git a/docs/design/swmr.md b/docs/design/swmr.md new file mode 100644 index 0000000..339c2d8 --- /dev/null +++ b/docs/design/swmr.md @@ -0,0 +1,124 @@ +# Design: reading files a SWMR writer is still appending to (range-read M5) + +Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader` +(see "Status" at the end). This is milestone M5 of +[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a +refresh". It covers the reader only; clawhdf5 does not write SWMR files. + +## What libhdf5 does + +A SWMR ("single writer, multiple readers") writer is a libhdf5 process that +opened a file with `libver='latest'` and switched to SWMR mode (h5py +`f.swmr_mode = True`, `H5Fstart_swmr_write`). Readers open the same file +with `H5F_ACC_SWMR_READ` (h5py `File(path, 'r', swmr=True)`) while the +writer keeps appending. What the format and the library guarantee: + +- **Superblock v3, flags set.** The writer sets the superblock's + file-consistency flags to write access + SWMR write (`0x05`) and clears + them on close. The superblock's end-of-file address is *not* kept up to + date while writing: a copy of a file taken mid-write records an EOF of a + few hundred bytes while the file is tens of kilobytes (checked on tank, + 2026-09-27, h5py 3.16 / HDF5 2.0: EOF 715 in a 17 857-byte file). A SWMR + reader therefore skips libhdf5's end-of-allocation check for every read + (`H5FD_read`: "allow access to data past the end of the allocated space + … for SWMR read access"), and bounds reads by the file's real length. +- **A non-SWMR open of such a file fails** in libhdf5: "file is already + open for write (may use to clear file consistency flags)". +- **The writer only appends.** New objects and attributes cannot be created + in SWMR mode; datasets grow with `H5Dset_extent` and are written. Chunked + datasets with one unlimited dimension use an Extensible Array index, with + more than one a version-2 B-tree; both are updated in a SWMR-safe way. + Fixed Array and single-chunk indexes are for datasets that cannot grow. +- **Flush ordering.** Every metadata structure the writer uses is + checksummed, and flush dependencies order the writes: a chunk's data is + written before the index entry that points to it, and index blocks before + the object header whose dataspace announces the new extent. A reader + that reads the object header first and the index after sees an index at + least as new as the extent, so every chunk inside the extent it read is + either in the index or never written (then it reads as the fill value, as + it does for libhdf5's reader). +- **Refresh.** A reader sees a dataset's new extent only when it refreshes + it (`H5Drefresh`, h5py `Dataset.refresh()`), which evicts the dataset's + cached metadata and reads the object header again. +- **Retries.** Reads are not atomic against writes on every system, so a + checksum can fail when a structure is read while the writer rewrites it. + A SWMR reader reads checksummed metadata up to 100 times before failing + (`H5Pset_metadata_read_attempts`; default 100 for SWMR access, 1 + otherwise — `H5Ppublic.h` of HDF5 1.14.6). + +## What clawhdf5 did before + +- `File::open` of a SWMR-flagged file bounded every read by the recorded + EOF (`Superblock::data_end` only tolerated an EOF *past* the end of the + file). A file copied or read mid-write therefore listed, but every + chunked read failed ("unexpected EOF: need 787 bytes, have 715"), and + `h5rs check` reported chunk indexes "past the end of the file". +- A `File` is a snapshot: an mmap (or a buffer) of the length at open, and a + per-file chunk cache that keeps each dataset's chunk index and decoded + chunks for the life of the `File`. A reader could not see growth at all, + and a mapping of a file that is being rewritten can change under a read. + +## Design + +1. **Bound SWMR files by their length.** `Superblock::data_end` returns the + file's length for a version-3 superblock with the SWMR-write flag, + whatever EOF it records, as libhdf5's SWMR reader does. Every existing + open path (`File::open`, `open_storage`, `h5rs`) then reads a finished + copy of a live file. We keep opening such files without a SWMR flag + (libhdf5 refuses): it is read-only and the alternative is an error. +2. **A live open: `File::open_swmr(path)` / `File::open_storage_swmr`.** + - The file is read with positioned reads (`FileStorage`, `pread` on + Unix, `seek_read` on Windows), never mapped, and `len()` is the file's + current length, so reads past the length seen at open work. + - The facade's view of the file (`FileData`) is *live*: reads are not + clamped to an end fixed at open, only by the storage's current length. + - The chunk cache is not used: every read reads the chunk index and the + chunks it needs again. A cached index would hide new chunks, and a + cached partial edge chunk would read as fill where the writer has since + written data (unfiltered edge chunks are rewritten in place). + - A storage that caches blocks (`clawhdf5-remote`'s `BlockCache`) would + serve stale bytes; `HttpStorage` also pins a file by ETag and length and + refuses a changed file. Remote SWMR is out of scope. +3. **`Dataset::refresh()`** reads the dataset's object header again (same + address) and replaces the handle's copy, so `shape()` and every later + read use the new extent. Like h5py, a handle that is not refreshed keeps + its extent; its reads still read the index as it is now, and only return + elements inside that extent. +4. **Bounded retries.** In a live file, an operation (open, dataset lookup, + refresh, every read) that fails with an error a concurrent write can + cause — a checksum or Fletcher-32 mismatch, a short read, a bad + signature, an undecodable chunk or chunk index — is run again from the + start, up to `File::swmr_read_attempts()` times (default 100, libhdf5's + default; `set_swmr_read_attempts` changes it), sleeping 1 µs, 2 µs, … + up to 10 ms between attempts (under a second in all). libhdf5 retries + the one structure whose checksum failed; we retry the whole operation, + because the parsers are pure functions of the bytes they read. Results + are only returned from a run where every structure verified, so an + error is returned instead of torn data. Other errors are returned at + once, and non-live files never retry. +5. **Writer state.** `File::swmr_writer_active()` reads the superblock + flags again, so a reader can tell when the writer has closed the file. + +What stays out: SWMR writing, VFD SWMR (HDF5 1.13's page-buffer protocol, +not in 1.14 or 2.0), remote SWMR, refresh of groups/attributes (a SWMR +writer cannot add them), and `MmapFile`/`LazyFile`. + +## Tests + +- `crates/clawhdf5-format`: `data_end` of a SWMR-flagged v3 superblock whose + EOF is below the file length. +- `crates/clawhdf5/tests/swmr_interop.rs`: + - a copy of a file taken mid-write (fixture) reads like h5py's SWMR reader; + - a live test: an h5py writer (`swmr_mode = True`) appends to a 1-D and a + 2-D dataset with one unlimited dimension (Extensible Array, one of them + gzip) and a 2-D dataset with two (v2 B-tree), flushing after every step, + while a Rust reader refreshes and reads them in a loop, and an h5py SWMR + reader does the same as the reference. Every value read must be the value + the writer wrote (a deterministic function of its position), extents + never shrink, and after the writer closes both readers must read the + same data as h5py. + +## Status + +See the end of this file's history in `CHANGELOG.md` ("Range reads, +milestone M5"). -- 2.54.0 From e39737682031c49fbbb90719dc17b70531f8e5da Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 06:31:15 -0500 Subject: [PATCH 02/30] format: a SWMR-flagged superblock's data ends at the end of the file A SWMR writer does not keep the superblock's end-of-file address up to date: a copy h5py made of its own file mid-write records 715 in a 6 030-byte file. Superblock::data_end only ignored the recorded end when it lay past the end of the file, so every open path bounded reads at 715: the file listed, but every chunked read failed ("unexpected EOF: need 787 bytes, have 715") and h5rs check reported the chunk indexes past the end of the file. libhdf5's SWMR reader skips the end-of-allocation check for every read (H5FD_read); for a v3 superblock with the SWMR-write flag the data now ends at the end of the file. Test: tests/swmr_interop.rs over the mid-write copy (fixture), through File::open, open_buffered and from_bytes, and against h5py's SWMR reader. Co-Authored-By: Claude Opus 5.5 (1M context) --- crates/clawhdf5-format/src/superblock.rs | 26 +++-- .../clawhdf5/tests/fixtures/swmr_mid_write.h5 | Bin 0 -> 6030 bytes crates/clawhdf5/tests/swmr_interop.rs | 107 ++++++++++++++++++ 3 files changed, 126 insertions(+), 7 deletions(-) create mode 100644 crates/clawhdf5/tests/fixtures/swmr_mid_write.h5 create mode 100644 crates/clawhdf5/tests/swmr_interop.rs diff --git a/crates/clawhdf5-format/src/superblock.rs b/crates/clawhdf5-format/src/superblock.rs index dc21f7c..5ae4daa 100644 --- a/crates/clawhdf5-format/src/superblock.rs +++ b/crates/clawhdf5-format/src/superblock.rs @@ -116,22 +116,27 @@ impl Superblock { /// [`FormatError::TruncatedFile`]. Bytes past that address are not part /// of the file: libhdf5 fails any read of them ("addr overflow" / /// "address plus size exceeds file eoa"), so a reader should parse only - /// the data up to the returned end. As libhdf5 does for a SWMR reader, - /// the check is skipped for a version-3 superblock whose writer is still - /// writing it in SWMR mode (it extends the file as it goes); the data - /// then ends at the end of the file. + /// the data up to the returned end. + /// + /// A version-3 superblock with the SWMR-write flag set belongs to a file + /// a SWMR writer has open (or had, and did not close). That writer does + /// not keep the recorded end of file up to date — a copy taken mid-write + /// can record an end of a few hundred bytes in a file of tens of + /// kilobytes — and libhdf5's SWMR reader skips its end-of-allocation + /// check for every read (`H5FD_read`). For such a superblock the data + /// ends at the end of the file, whatever end it records. /// /// When the superblock's recorded base address differs from where the /// superblock actually is (a user block added or removed after the file /// was written), libhdf5 moves the recorded end of file by the same /// amount, and so does this. pub fn data_end(&self, user_block: u64, file_len: u64) -> Result { + if self.version >= 3 && self.is_swmr_write() { + return Ok(file_len.saturating_sub(user_block)); + } let eof = i128::from(self.eof_address) - i128::from(self.base_address) + i128::from(user_block); if eof < 0 || eof > i128::from(file_len) { - if self.version >= 3 && self.is_swmr_write() { - return Ok(file_len.saturating_sub(user_block)); - } return Err(FormatError::TruncatedFile { stored_eof: u64::try_from(eof).unwrap_or(self.eof_address), actual_len: file_len, @@ -617,6 +622,13 @@ mod tests { let mut swmr = Superblock::parse(&build_v2_bytes(8, 3), 0).unwrap(); swmr.consistency_flags = swmr_flags::WRITE_ACCESS | swmr_flags::SWMR_WRITE; assert_eq!(swmr.data_end(0, 1000), Ok(1000)); + // ... nor bounded by its recorded end, which the writer does not + // keep up to date (2048 here). + assert_eq!(swmr.data_end(0, 17_857), Ok(17_857)); + assert_eq!(swmr.data_end(512, 17_857), Ok(17_345)); + // Without the SWMR-write flag the recorded end bounds the data. + swmr.consistency_flags = swmr_flags::WRITE_ACCESS; + assert_eq!(swmr.data_end(0, 17_857), Ok(2048)); } #[test] diff --git a/crates/clawhdf5/tests/fixtures/swmr_mid_write.h5 b/crates/clawhdf5/tests/fixtures/swmr_mid_write.h5 new file mode 100644 index 0000000000000000000000000000000000000000..28cbb9487345a2fa7d1d9653c61c48cab8d5d79b GIT binary patch literal 6030 zcmeI#c~nzZ8UXMY2qq8+VNoe-aJR^&D3rwziv&ekWE49{6mb-5WnTg@pg^EV1rZP( z2%sH{)2X7h6+{6cAQd;LC|FsH?6L|}7D0Mn@;wiK%sF$;{AEwQN8j(=`|f-1%a@yb zFG)@gjw(vpO7c{y0tRIk$~`$*gBi3hlE2Gzb#mAyhwJOhL@vniVTha%b$LyUq{d{u&y58Q zvzV2~AV)NmIrI?x+r#iXs0*Z5>!h#|ZKPKXzwp(rgvNOZtKjpU8&O&c(b7-sjNYXw8r zllW%->RKFi;z=__BF0vqiY!BEGNi{NbXg^7PMeFWtdcONtpVB+=d`uK#P(OqbeOMs z>PEK12JC%i@uWyJWT}~(AwL+tU(d-}G@38*Ld+=mxjLE;@RBb*$__t5Tkf1Tw}rCG z`Z;Yx_UAOcUqAj#ZNGS!p!Mxz`Iv!^8Tha=K&;Kg+D>d$h`kT7znuqRQ$X215yKRr zPORxPs1rYTSe7tYiZEDeFj&wrHRvzh8l!WVXfa}M023+(6D9_eB&G%TeFAEOI-oA7 z2kL_cAQLnM7lDhxCE!wU8Mqu=0j>l;1&zQ}pfP9ynu5fi2j);;4Xyz#z_s8y&=Ms6 z9QX|C)}Re&3)+Fis=|W0J;(;prb}8r2hpZbS~h}?AO~~;H-Vc$XK)L+6?6ezK{wDH z+y;7pe*?FJJ3voxC-?>UchC#m1$u*Aa5uOI^a1yRzTiHP$6Z^qiB59Ybd<#S`6)Us z+WY<4Cz(w4`Q&BC-5MMRx2@_lu2x^gF76iH&00J&r5Uq?rrlJvF8oTeU4J{4XjLg^ z9NKUuZ`ebBZF~20EXR0qetDOBw%OrHZ@-?&aoc4L$5a&z-sD6^6n2F_xM0tG24$t5E4BYSUY+tUg;x&D&CMCGT-Nk>y3D)SDt9N$uEWUc} zsc%zz|J&x~cObKes&!+uSHtNic8&bQ1cxG9b|(k{Cy84jM$xj#{nqEq{F!4Jjf zji(fzE#^pgT;DyI+%a_uqdFhEppyTSd`2sFLDjyO<~Bvjpsx_tDY}(ln=S2oRo(u1 zPI&{({-DANJJM#;qrd9pkC9f`Df!LA8hHyHcwFOZXEk|4hJ8iaQSuU|&k`m{QTJ5K z^mQdgCwn>dlX#=yrU!k_$Q(koG;gFoRlKq^XzZ!bE23g66ft}Q(7lWH~XI5fZ28HfX$%qz{55r$98B+?Ue zJe2()P^&a!Ocq zCkI;N%DK0qMh)JQa>G}z{0EHg{on!64?GBd2_6E!0uO`!U;r2h27$p~2p9^6f#F~T z$O9w6D3A{x0R><*7y}*!kAbmZ92gG@!2~c7{2ELGkAunJH{b~{1xy7`f@$C>FdfVQ zKk{K`0G|x>v?}c}tWJ#o?zHq1kcmEZk0gOs%3}PGC^q|{jPJ^&vXY&r7oO%u4wR~{ zFBuWLMhwmb6?ZM!-O6v^9x93Lw* zJ>hLWT06R1j4b2x`ZZ^T!mNj-K?9c}1_k3U5?S=_&@S}zZ!LcFK*VpedpnRSURoL+ z+FQ#hD!Vb-k#(Tld)H&5LcP-Q3QkdU)L2JW)>i8XNoG-NoR|k|01a%_p{$CoeOWh zsSu``2*NtiEIQGf)#g8xy*NUb-;-A3GF?Nd_u5>KXJt+_WHw_AW22zhX83<Wmim^=H9S6p1RU`#hlV*j#9FDL95i#Kofdqz}D)w9nFY)we+!{-e#8ec# z)yLo2I?@apYF&EYlPu6Co%v?0N8{c_e!XTYqLTB)mb z=T9{HjO6^G%Qq$8p(pewHMa8lMCb{F9L;wm14U0|JU+wIY!u&Z18H!Kc$HWtwNe#7 z%##JkZn{Fbt$|;q$O0xwd){8-v5eZfNIdMp2t(0`LRu9Xr~68D(09g~nD~qciP%rta@+JOsqp&R4iMx z6HV{>y8kHMO2nQ%NM{~DC%<{zzE{yLN(J9d6xsLcxlJi$l#wnB+4oZ1rc^WbpP`@o SBHx0Np|2o659{JAsQLp{_WNA` literal 0 HcmV?d00001 diff --git a/crates/clawhdf5/tests/swmr_interop.rs b/crates/clawhdf5/tests/swmr_interop.rs new file mode 100644 index 0000000..f22e7cd --- /dev/null +++ b/crates/clawhdf5/tests/swmr_interop.rs @@ -0,0 +1,107 @@ +//! Files a libhdf5 SWMR writer (h5py `f.swmr_mode = True`) has open: a copy +//! taken mid-write (fixture), and a live file appended to by an h5py writer +//! process while clawhdf5 and h5py's own SWMR reader read it (see +//! `docs/design/swmr.md`). +//! +//! The live tests need python3 with h5py; they are skipped without it, +//! unless `CLAWHDF5_REQUIRE_INTEROP=1`. + +use std::path::{Path, PathBuf}; +use std::process::Command; + +use clawhdf5::File; + +fn python() -> String { + std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string()) +} + +fn interop_required() -> bool { + std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1") +} + +fn python_available() -> bool { + Command::new(python()) + .args(["-c", "import h5py, numpy"]) + .output() + .map(|o| o.status.success()) + .unwrap_or(false) +} + +macro_rules! skip_if_no_python { + () => { + if !python_available() { + assert!( + !interop_required(), + "CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available" + ); + eprintln!("SKIP: python3 with h5py not available"); + return; + } + }; +} + +fn fixture(name: &str) -> PathBuf { + Path::new(env!("CARGO_MANIFEST_DIR")) + .join("tests/fixtures") + .join(name) +} + +/// `swmr_mid_write.h5`: a copy h5py 3.16 (HDF5 2.0) made of its own file +/// while writing it in SWMR mode, after 4 appends of 37 rows to +/// `/a` (int64, chunks of 100, no filter: `a[i] = i + 1`) and `/b` (float64 +/// `(n, 4)`, chunks of 16 x 4, gzip: `b[i, j] = 10 i + j + 1`), each +/// followed by a flush. Its superblock (v3) still has the SWMR-write flag +/// set and records an end of file of 715 in a 6 030-byte file. +fn check_mid_write_copy(f: &File) { + let sb = f.superblock(); + assert_eq!(sb.version, 3); + assert!(sb.is_swmr_write()); + let a = f.dataset("a").unwrap(); + assert_eq!(a.shape().unwrap(), vec![148]); + let want_a: Vec = (1..=148).collect(); + assert_eq!(a.read_i64().unwrap(), want_a); + let b = f.dataset("b").unwrap(); + assert_eq!(b.shape().unwrap(), vec![148, 4]); + let want_b: Vec = (0..148) + .flat_map(|i| (0..4).map(move |j| (10 * i + j + 1) as f64)) + .collect(); + assert_eq!(b.read_f64().unwrap(), want_b); +} + +#[test] +fn a_copy_taken_mid_write_reads_past_its_recorded_end_of_file() { + let path = fixture("swmr_mid_write.h5"); + // The recorded end of file (715) is far below the file's length; the + // chunk indexes and chunks lie past it. libhdf5's SWMR reader does not + // bound reads by it, and neither does any open path here. + check_mid_write_copy(&File::open(&path).unwrap()); + check_mid_write_copy(&File::open_buffered(&path).unwrap()); + check_mid_write_copy(&File::from_bytes(std::fs::read(&path).unwrap()).unwrap()); +} + +#[test] +fn a_copy_taken_mid_write_reads_as_h5py_swmr_reader_reads_it() { + skip_if_no_python!(); + let path = fixture("swmr_mid_write.h5"); + // libhdf5 refuses a non-SWMR open of this file ("file is already open + // for write"); its SWMR reader reads the values checked above. + let script = format!( + r#" +import h5py, numpy as np +with h5py.File("{p}", "r", swmr=True, locking=False) as f: + a = f["a"][()] + b = f["b"][()] +assert np.array_equal(a, np.arange(148) + 1), a +assert np.array_equal(b, (np.arange(148)[:, None] * 10 + np.arange(4) + 1).astype("f8")), b +print("ok") +"#, + p = path.display() + ); + let out = Command::new(python()).args(["-c", &script]).output().unwrap(); + assert!( + out.status.success(), + "h5py failed:\n{}", + String::from_utf8_lossy(&out.stderr) + ); + check_mid_write_copy(&File::open(&path).unwrap()); +} -- 2.54.0 From 076feb089a79e0141bcaf27cd041fa55ed6b4596 Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 06:35:21 -0500 Subject: [PATCH 03/30] wasm: restartable NeedBytes storage for lazy reads (range-read M4, core) LazyStorage holds the blocks of a remote file fetched so far. An operation runs as passes over it: a read that misses records the missing blocks and fails, the pass's result is dropped whatever it is (a parser may have caught the error and carried on), and the caller fetches the reported ranges and re-runs the pass. No block is evicted while an operation is in flight, so every pass that does not finish asks for at least one new block and the operation ends. Blocks sit in an LRU with a byte budget, trimmed between operations, bulk (raw data) blocks first. Reader::open_storage opens a file through any Storage, and variable- length strings resolve through the file's storage instead of File::as_bytes, which panics for a file not in memory. tests/lazy.rs compares, file by file, what the viewer can show (kinds, listings, attributes, info, whole reads and hyperslabs) read lazily with the facade's range-storage path and the in-memory reader: files built here at 512 B to 1 MiB blocks, the h5py/netCDF4 fixture, and with CLAWHDF5_WASM_CORPUS the conformance corpus (656 files agree). A listing plus a small read of a 48 MB file fetches 3 ranges. Co-Authored-By: Claude Opus 5.5 (1M context) --- crates/clawhdf5-wasm/src/core.rs | 28 +- crates/clawhdf5-wasm/src/lazy.rs | 672 +++++++++++++++++++++++++++++ crates/clawhdf5-wasm/src/lib.rs | 1 + crates/clawhdf5-wasm/tests/lazy.rs | 391 +++++++++++++++++ 4 files changed, 1086 insertions(+), 6 deletions(-) create mode 100644 crates/clawhdf5-wasm/src/lazy.rs create mode 100644 crates/clawhdf5-wasm/tests/lazy.rs diff --git a/crates/clawhdf5-wasm/src/core.rs b/crates/clawhdf5-wasm/src/core.rs index 4c6beab..23b4b1d 100644 --- a/crates/clawhdf5-wasm/src/core.rs +++ b/crates/clawhdf5-wasm/src/core.rs @@ -6,9 +6,12 @@ //! with no typed-array mapping (compound, reference, opaque, ...) is refused //! with a message naming it, never returned as reinterpreted bytes. +use std::sync::Arc; + use clawhdf5::{AttrValue, File, Selection}; use clawhdf5_format::data_read; use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder}; +use clawhdf5_format::storage::Storage; use clawhdf5_format::vl_data::{VlResolver, check_element_size}; /// Errors are reported to JavaScript as messages. @@ -119,7 +122,9 @@ pub struct Attr { pub value: AttrValue, } -/// An open file, held in memory. +/// An open file: held in memory ([`Reader::open`]) or read through a +/// [`Storage`] ([`Reader::open_storage`], such as a +/// [`LazyStorage`](crate::lazy::LazyStorage)). pub struct Reader { file: File, } @@ -132,6 +137,14 @@ impl Reader { }) } + /// Open a file read through `storage` (the file's bytes from offset 0, + /// user block included, as [`File::open_storage`] takes them). + pub fn open_storage(storage: Arc) -> Result { + Ok(Self { + file: File::open_storage(storage).map_err(err)?, + }) + } + /// Whether `path` names a group or a dataset. pub fn kind(&self, path: &str) -> Result { match self.file.dataset(path) { @@ -281,11 +294,14 @@ impl Reader { // string ends at its first NUL and a heap object of the // wrong size is an error, as in libhdf5 and h5py. let sb = self.file.superblock(); - Data::Strings( - VlResolver::new(self.file.as_bytes(), sb.offset_size, sb.length_size) - .strings(raw) - .map_err(err)?, - ) + let strings = match self.file.contiguous_bytes() { + Some(bytes) => { + VlResolver::new(bytes, sb.offset_size, sb.length_size).strings(raw) + } + None => VlResolver::new_in(self.file.storage(), sb.offset_size, sb.length_size) + .strings(raw), + }; + Data::Strings(strings.map_err(err)?) } Datatype::Enumeration { .. } if !is_array => { Data::Strings(data_read::read_enum_names(raw, dt).map_err(err)?) diff --git a/crates/clawhdf5-wasm/src/lazy.rs b/crates/clawhdf5-wasm/src/lazy.rs new file mode 100644 index 0000000..193f61f --- /dev/null +++ b/crates/clawhdf5-wasm/src/lazy.rs @@ -0,0 +1,672 @@ +//! Reading a file that is not all here, when no read may wait for the +//! network: the restartable "NeedBytes" mode of `docs/design/range-reads.md` +//! (milestone M4). +//! +//! A browser's main thread cannot block on `fetch`, and the parsers are +//! synchronous. So an operation (open, list a group, read a dataset) runs +//! as a *pass* over a [`LazyStorage`] that holds the blocks fetched so far: +//! +//! 1. [`LazyStorage::attempt`] runs the operation. A read whose blocks are +//! all present is served; a read that misses records the missing blocks +//! and fails with a storage error. +//! 2. If the pass missed anything, its result is thrown away — whatever it +//! is, since a parser may have caught the error and carried on (a +//! listing skips a link it cannot resolve) — and the caller gets the +//! byte ranges to fetch ([`Step::Need`]). +//! 3. The caller fetches them (asynchronously, with HTTP `Range` requests), +//! hands them over with [`LazyStorage::supply`] and runs the operation +//! again. +//! +//! A pass is pure over the storage: the facade only caches what completed +//! reads decoded (its chunk cache), so re-running it is safe. Every pass +//! that does not finish asks for at least one block not yet present, and no +//! block is evicted while an operation is in flight +//! ([`LazyStorage::operation`]), so an operation finishes after at most one +//! pass per block it needs. In practice it is one pass per *wave* of +//! misses: a chunked read asks for all the chunks of a batch at once. +//! +//! Blocks are kept in an LRU cache with a byte budget, trimmed only when no +//! operation is in flight. Blocks fetched for bulk reads (raw data: a +//! `read_ranges` call, or a read longer than a block) go first, so reading +//! a large dataset does not evict the metadata. + +use std::borrow::Cow; +use std::collections::{BTreeSet, HashMap}; +use std::ops::Range; +use std::sync::{Arc, Mutex, MutexGuard}; + +use clawhdf5_format::error::FormatError; +use clawhdf5_format::storage::Storage; + +/// Default block size: 1 MiB, as `clawhdf5-remote`'s block cache (the size +/// `docs/design/range-reads.md` §2 measured). +pub const DEFAULT_BLOCK_SIZE: u64 = 1 << 20; + +/// The message of the error a read that misses returns. It never reaches +/// the caller of [`LazyStorage::attempt`]: a pass that missed is re-run. +pub const NEED_BYTES: &str = "bytes not fetched yet (restartable read)"; + +/// Settings of a [`LazyStorage`]. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct LazyConfig { + /// Size of a block in bytes (at least 512); fetches are whole, aligned + /// blocks (the file's last block is shorter). + pub block_size: u64, + /// Byte budget of cached blocks between operations. An operation keeps + /// every block it needs until it finishes, whatever the budget. + pub capacity: u64, + /// Largest single range asked for, in bytes (whole blocks, at least + /// one); longer runs are split so they can be fetched in parallel. + pub max_request: u64, +} + +impl Default for LazyConfig { + fn default() -> Self { + LazyConfig { + block_size: DEFAULT_BLOCK_SIZE, + capacity: 64 << 20, + max_request: 8 << 20, + } + } +} + +/// What a [`LazyStorage`] has done so far. +#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)] +pub struct LazyStats { + /// Passes run by [`LazyStorage::attempt`]. + pub passes: u64, + /// Ranges handed to [`LazyStorage::supply`]: one HTTP request each. + pub requests: u64, + /// Bytes handed to [`LazyStorage::supply`]. + pub bytes_fetched: u64, + /// Blocks evicted to stay within the budget. + pub evictions: u64, + /// Bytes cached now. + pub cached_bytes: u64, +} + +/// The outcome of one pass. +#[derive(Debug)] +pub enum Step { + /// The pass read only bytes that were present: its result stands. + Done(T), + /// The pass missed: fetch these byte ranges (sorted, disjoint, block + /// aligned), [`supply`](LazyStorage::supply) them and run it again. + Need(Vec>), +} + +struct Block { + data: Arc<[u8]>, + /// Eviction order: bulk blocks (`false`) before metadata (`true`), + /// then least recently used first. + key: (bool, u64), +} + +#[derive(Default)] +struct State { + blocks: HashMap, + /// `(metadata?, tick, block index)`, in eviction order. + order: BTreeSet<(bool, u64, u64)>, + tick: u64, + bytes: u64, + /// Blocks the current pass missed, and whether a small read wanted + /// them (metadata). + missing: HashMap, + /// Blocks a bulk read missed that have not been supplied yet: kept + /// as bulk when they arrive. + bulk_pending: BTreeSet, + /// Operations in flight: no eviction while any is. + active: u32, + stats: LazyStats, +} + +/// A [`Storage`] over the blocks of a file fetched so far; a read of +/// anything else fails and is recorded, so the pass can be re-run once the +/// bytes arrive. See the [module documentation](self). +pub struct LazyStorage { + len: u64, + config: LazyConfig, + state: Mutex, +} + +fn lock(m: &Mutex) -> MutexGuard<'_, State> { + m.lock().unwrap_or_else(std::sync::PoisonError::into_inner) +} + +/// Keeps an operation's blocks cached until it is dropped; see +/// [`LazyStorage::operation`]. +pub struct Operation<'a> { + storage: &'a LazyStorage, +} + +impl Drop for Operation<'_> { + fn drop(&mut self) { + let mut st = lock(&self.storage.state); + st.active = st.active.saturating_sub(1); + if st.active == 0 { + self.storage.evict(&mut st); + } + } +} + +impl LazyStorage { + /// An empty cache for a file of `len` bytes. + pub fn new(len: u64, mut config: LazyConfig) -> Self { + config.block_size = config.block_size.max(512); + config.max_request = (config.max_request / config.block_size).max(1) * config.block_size; + LazyStorage { + len, + config, + state: Mutex::new(State::default()), + } + } + + /// The settings in use (after rounding). + pub fn config(&self) -> &LazyConfig { + &self.config + } + + /// Counters since the storage was made. + pub fn stats(&self) -> LazyStats { + let st = lock(&self.state); + LazyStats { + cached_bytes: st.bytes, + ..st.stats + } + } + + /// Mark an operation in flight until the guard is dropped: no block is + /// evicted meanwhile, so re-running its passes always makes progress. + /// Hold it across every pass of one operation. + pub fn operation(&self) -> Operation<'_> { + lock(&self.state).active += 1; + Operation { storage: self } + } + + /// Run one pass of `f` over this storage. `Done` when `f` read nothing + /// that is missing; otherwise `Need` with the ranges to fetch, and `f`'s + /// result is dropped (it may be an error caused by the miss, or a + /// result built around one). + pub fn attempt(&self, f: impl FnOnce() -> T) -> Step { + { + let mut st = lock(&self.state); + st.missing.clear(); + st.stats.passes += 1; + } + let out = f(); + let missing = std::mem::take(&mut lock(&self.state).missing); + if missing.is_empty() { + return Step::Done(out); + } + drop(out); + Step::Need(self.runs(missing)) + } + + /// The bytes of the file at `offset`, fetched for a range a pass asked + /// for. `offset` must be block aligned and the bytes whole blocks (the + /// last block of the file may be short) inside the file, or this is an + /// error and nothing is kept. Blocks already present are left alone. + pub fn supply(&self, offset: u64, bytes: &[u8]) -> Result<(), String> { + let bs = self.config.block_size; + let end = offset + .checked_add(bytes.len() as u64) + .filter(|&e| e <= self.len) + .ok_or_else(|| { + format!( + "{} bytes at offset {offset} run past the end of the {}-byte file", + bytes.len(), + self.len + ) + })?; + if !offset.is_multiple_of(bs) || (!end.is_multiple_of(bs) && end != self.len) { + return Err(format!( + "{} bytes at offset {offset} are not whole {bs}-byte blocks", + bytes.len() + )); + } + let mut st = lock(&self.state); + st.stats.requests += 1; + st.stats.bytes_fetched += bytes.len() as u64; + let mut start = offset; + while start < end { + let i = start / bs; + let stop = (start + bs).min(end); + if !st.blocks.contains_key(&i) { + let rel = (start - offset) as usize..(stop - offset) as usize; + let metadata = !st.bulk_pending.remove(&i); + self.keep(&mut st, i, Arc::from(&bytes[rel]), metadata); + } + start = stop; + } + if st.active == 0 { + self.evict(&mut st); + } + Ok(()) + } + + /// [`supply`](Self::supply) the bytes fetched for `range`, one of the + /// ranges a [`Step::Need`] asked for: anything but exactly its length + /// (a server that answered with more or less) is an error. + pub fn supply_range(&self, range: &Range, bytes: &[u8]) -> Result<(), String> { + let want = range.end.saturating_sub(range.start); + if bytes.len() as u64 != want { + return Err(format!( + "asked for {want} bytes at offset {}, got {}", + range.start, + bytes.len() + )); + } + self.supply(range.start, bytes) + } + + /// Run `f` to completion, fetching what its passes miss with `fetch` + /// (a byte range to its bytes). The blocking driver, for native code + /// and tests; the browser's is the same loop with an `await` between + /// passes. + pub fn run_blocking( + &self, + mut f: impl FnMut() -> T, + mut fetch: impl FnMut(Range) -> Result, String>, + ) -> Result { + let _op = self.operation(); + loop { + match self.attempt(&mut f) { + Step::Done(v) => return Ok(v), + Step::Need(ranges) => { + for r in ranges { + let bytes = fetch(r.clone())?; + self.supply_range(&r, &bytes)?; + } + } + } + } + } + + /// Cache block `i`. + fn keep(&self, st: &mut State, i: u64, data: Arc<[u8]>, metadata: bool) { + st.tick += 1; + let key = (metadata, st.tick); + st.bytes += <[u8]>::len(&data) as u64; + st.order.insert((key.0, key.1, i)); + if let Some(old) = st.blocks.insert(i, Block { data, key }) { + st.order.remove(&(old.key.0, old.key.1, i)); + st.bytes -= <[u8]>::len(&old.data) as u64; + } + } + + fn evict(&self, st: &mut State) { + while st.bytes > self.config.capacity { + let Some((_, _, i)) = st.order.pop_first() else { + break; + }; + if let Some(b) = st.blocks.remove(&i) { + st.bytes -= <[u8]>::len(&b.data) as u64; + st.stats.evictions += 1; + } + } + } + + /// Byte ranges covering the missing blocks: runs of consecutive + /// blocks, a one-block hole between two runs filled so they merge, + /// each at most `max_request` long. + fn runs(&self, missing: HashMap) -> Vec> { + let bs = self.config.block_size; + let mut wanted: Vec = missing.keys().copied().collect(); + wanted.sort_unstable(); + { + // Remember which blocks only bulk reads asked for: they are + // kept as bulk once supplied. + let mut st = lock(&self.state); + for (&i, &metadata) in &missing { + if metadata { + st.bulk_pending.remove(&i); + } else { + st.bulk_pending.insert(i); + } + } + } + let per_request = self.config.max_request / bs; + let mut runs: Vec<(u64, u64)> = Vec::new(); + for i in wanted { + match runs.last_mut() { + Some((first, last)) if i <= *last + 2 && i - *first < per_request => *last = i, + _ => runs.push((i, i)), + } + } + runs.into_iter() + .map(|(a, b)| a * bs..((b + 1) * bs).min(self.len)) + .collect() + } + + /// Block indices covering `[offset, offset + len)`, clamped to the file. + fn span(&self, offset: u64, len: u64) -> Option> { + let end = offset.saturating_add(len).min(self.len); + if offset >= end { + return None; + } + let bs = self.config.block_size; + Some(offset / bs..(end - 1) / bs + 1) + } + + /// The blocks of `spans` if all are present (touching them), else + /// record the missing ones and fail. + fn blocks( + &self, + spans: &[Range], + metadata: bool, + ) -> Result>, FormatError> { + let mut st = lock(&self.state); + let mut have = HashMap::new(); + let mut missed = false; + for span in spans { + for i in span.clone() { + if have.contains_key(&i) { + continue; + } + match st.blocks.get(&i) { + Some(b) => { + have.insert(i, b.data.clone()); + } + None => { + missed = true; + let m = st.missing.entry(i).or_insert(metadata); + *m |= metadata; + } + } + } + } + if missed { + return Err(FormatError::Storage(NEED_BYTES.into())); + } + // Touch: most recently used last; a small read promotes a bulk + // block to metadata. + for &i in have.keys() { + st.tick += 1; + let tick = st.tick; + let Some(b) = st.blocks.get_mut(&i) else { + continue; + }; + let old = b.key; + b.key = (old.0 || metadata, tick); + let new = b.key; + st.order.remove(&(old.0, old.1, i)); + st.order.insert((new.0, new.1, i)); + } + Ok(have) + } + + fn assemble(&self, offset: u64, end: u64, blocks: &HashMap>) -> Vec { + let bs = self.config.block_size; + let mut out = Vec::with_capacity(usize::try_from(end - offset).unwrap_or(0)); + let mut pos = offset; + while pos < end { + let i = pos / bs; + let block = &blocks[&i]; + let from = (pos - i * bs) as usize; + let to = ((end - i * bs) as usize).min(<[u8]>::len(block)); + out.extend_from_slice(&block[from..to]); + pos = i * bs + to as u64; + } + out + } +} + +impl Storage for LazyStorage { + fn read_at(&self, offset: u64, len: usize) -> Result, FormatError> { + let Some(span) = self.span(offset, len as u64) else { + return Ok(Cow::Owned(Vec::new())); + }; + let metadata = len as u64 <= self.config.block_size; + let blocks = self.blocks(std::slice::from_ref(&span), metadata)?; + let end = offset.saturating_add(len as u64).min(self.len); + Ok(Cow::Owned(self.assemble(offset, end, &blocks))) + } + + fn len(&self) -> u64 { + self.len + } + + fn read_ranges(&self, ranges: &[Range]) -> Result>, FormatError> { + let mut spans = Vec::with_capacity(ranges.len()); + for r in ranges { + if r.end < r.start { + return Err(FormatError::Storage( + "read range ends before it starts".into(), + )); + } + spans.extend(self.span(r.start, r.end - r.start)); + } + let blocks = self.blocks(&spans, false)?; + Ok(ranges + .iter() + .map(|r| { + let end = r.end.min(self.len); + if r.start >= end { + Cow::Owned(Vec::new()) + } else { + Cow::Owned(self.assemble(r.start, end, &blocks)) + } + }) + .collect()) + } +} + +#[cfg(test)] +#[allow(clippy::single_range_in_vec_init)] +mod tests { + use super::*; + + fn file(n: usize) -> Vec { + (0..n).map(|i| (i * 7 + i / 251) as u8).collect() + } + + fn config(block: u64, capacity: u64) -> LazyConfig { + LazyConfig { + block_size: block, + capacity, + max_request: 4 * block, + } + } + + /// Supply every range of `need` from `data`. + fn serve(s: &LazyStorage, data: &[u8], need: &[Range]) { + for r in need { + s.supply_range(r, &data[r.start as usize..r.end as usize]) + .unwrap(); + } + } + + fn owned(r: Result, FormatError>) -> Result, FormatError> { + r.map(Cow::into_owned) + } + + #[test] + fn a_miss_asks_for_whole_blocks_then_the_rerun_reads_them() { + let data = file(10_000); + let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20)); + let Step::Need(need) = s.attempt(|| owned(s.read_at(1500, 1000))) else { + panic!("nothing is cached yet"); + }; + assert_eq!(need, vec![1024..3072]); + serve(&s, &data, &need); + let Step::Done(got) = s.attempt(|| owned(s.read_at(1500, 1000))) else { + panic!("the blocks were supplied"); + }; + assert_eq!(got.unwrap(), &data[1500..2500]); + // Past the end: short, then empty, as for a slice. + serve(&s, &data, &[9216..10_000]); + let Step::Done(tail) = s.attempt(|| owned(s.read_at(9_990, 100))) else { + panic!("the last block was supplied"); + }; + assert_eq!(tail.unwrap(), &data[9_990..]); + assert!(matches!( + s.attempt(|| s.read_at(20_000, 10).map(|c| c.len())), + Step::Done(Ok(0)) + )); + let st = s.stats(); + assert_eq!( + (st.passes, st.requests, st.bytes_fetched), + (4, 2, 2048 + 784) + ); + } + + #[test] + fn a_pass_that_swallowed_the_miss_is_still_rerun() { + // A parser that catches the error and returns something anyway + // (a listing skipping a link it cannot resolve) must not have its + // result used. + let data = file(4096); + let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20)); + let step = s.attempt(|| s.read_at(0, 4).map(|b| b.len()).unwrap_or(0)); + assert!( + matches!(step, Step::Need(ref n) if n == &vec![0..1024]), + "{step:?}" + ); + } + + #[test] + fn read_ranges_asks_for_every_missing_block_in_one_pass() { + let data = file(64 * 1024); + let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20)); + let ranges = [ + 100..200, + 5000..5100, + 5200..5300, + 30_000..33_000, + 60_000..60_010, + ]; + let read = || { + s.read_ranges(&ranges) + .map(|v| v.into_iter().map(Cow::into_owned).collect::>()) + }; + let Step::Need(need) = s.attempt(read) else { + panic!("nothing is cached yet"); + }; + // 5000..5300 is blocks 4 and 5; 30_000..33_000 is blocks 29..=32. + assert_eq!( + need, + vec![ + 0..1024, + 4096..6144, + 29 * 1024..33 * 1024, + 58 * 1024..59 * 1024 + ] + ); + serve(&s, &data, &need); + let Step::Done(got) = s.attempt(read) else { + panic!("one pass fetched everything"); + }; + for (r, g) in ranges.iter().zip(&got.unwrap()) { + assert_eq!(g, &data[r.start as usize..r.end as usize]); + } + } + + #[test] + fn runs_merge_one_block_holes_and_split_long_runs() { + let data = file(32 * 1024); + let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20)); + // Blocks 0 and 2 (a hole of one: merged), 5..=14 (split in fours). + let Step::Need(need) = s.attempt(|| { + let _ = s.read_at(0, 10); + let _ = s.read_at(2048, 10); + s.read_ranges(&[5120..15 * 1024]).map(|_| ()) + }) else { + panic!("nothing is cached yet"); + }; + assert_eq!( + need, + vec![0..3072, 5120..9216, 9216..13_312, 13_312..15_360] + ); + } + + #[test] + fn supply_refuses_what_was_not_asked_for() { + let data = file(10_000); + let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20)); + assert!(s.supply(1, &data[1..1025]).unwrap_err().contains("whole")); + assert!(s.supply(0, &data[..1000]).unwrap_err().contains("whole")); + assert!( + s.supply(9216, &[0u8; 1024]) + .unwrap_err() + .contains("past the end") + ); + for wrong in [&data[..1023], &data[..2048]] { + assert!( + s.supply_range(&(0..1024), wrong) + .unwrap_err() + .contains("asked for 1024 bytes") + ); + } + assert_eq!(s.stats().cached_bytes, 0); + // The file's last, short block is whole. + s.supply(9216, &data[9216..]).unwrap(); + assert_eq!(s.stats().cached_bytes, 784); + } + + #[test] + fn an_operation_keeps_its_blocks_whatever_the_budget() { + // A budget of one block, an operation that needs eight: without + // the operation guard every pass would evict what the last one + // fetched and never finish. + let data = file(8 * 1024); + let s = LazyStorage::new(data.len() as u64, config(1024, 1024)); + let got = s + .run_blocking( + || { + (0..8) + .map(|i| owned(s.read_at(i * 1024 + 3, 10))) + .collect::, _>>() + }, + |r| Ok(data[r.start as usize..r.end as usize].to_vec()), + ) + .unwrap() + .unwrap(); + for (i, g) in got.iter().enumerate() { + assert_eq!(g, &data[i * 1024 + 3..i * 1024 + 13]); + } + // Trimmed to the budget once the operation is over. + let st = s.stats(); + assert_eq!(st.cached_bytes, 1024); + assert_eq!(st.evictions, 7); + assert_eq!(st.passes, 9, "one pass per block, then the one that ends"); + } + + #[test] + fn bulk_blocks_are_evicted_before_metadata() { + let data = file(16 * 1024); + let s = LazyStorage::new(data.len() as u64, config(1024, 4 * 1024)); + let fetch = |r: Range| Ok(data[r.start as usize..r.end as usize].to_vec()); + // Metadata: a small read of block 0. + s.run_blocking(|| s.read_at(0, 16).map(|_| ()), fetch) + .unwrap() + .unwrap(); + // Bulk: raw data over blocks 4..12, more than the budget. + s.run_blocking(|| s.read_ranges(&[4096..12 * 1024]).map(|_| ()), fetch) + .unwrap() + .unwrap(); + // The metadata block survived: reading it again is a hit. + let before = s.stats(); + assert!(before.cached_bytes <= 4 * 1024); + assert!(matches!( + s.attempt(|| s.read_at(0, 16).map(|_| ())), + Step::Done(Ok(())) + )); + assert_eq!(s.stats().requests, before.requests); + } + + #[test] + fn a_failed_fetch_is_an_error_not_data() { + let data = file(4096); + let s = LazyStorage::new(data.len() as u64, config(1024, 1 << 20)); + let e = s + .run_blocking(|| owned(s.read_at(0, 8)), |_| Err("HTTP 500".into())) + .unwrap_err(); + assert_eq!(e, "HTTP 500"); + let e = s + .run_blocking(|| owned(s.read_at(0, 8)), |_| Ok(vec![0; 10])) + .unwrap_err(); + assert!(e.contains("asked for 1024 bytes"), "{e}"); + assert_eq!(s.stats().cached_bytes, 0); + drop(data); + } +} diff --git a/crates/clawhdf5-wasm/src/lib.rs b/crates/clawhdf5-wasm/src/lib.rs index 7c6b4a9..2fed048 100644 --- a/crates/clawhdf5-wasm/src/lib.rs +++ b/crates/clawhdf5-wasm/src/lib.rs @@ -21,6 +21,7 @@ //! The logic lives in [`core`], which is plain Rust and tested natively. pub mod core; +pub mod lazy; use clawhdf5::AttrValue; use js_sys::{Array, Object, Reflect}; diff --git a/crates/clawhdf5-wasm/tests/lazy.rs b/crates/clawhdf5-wasm/tests/lazy.rs new file mode 100644 index 0000000..eed0ed5 --- /dev/null +++ b/crates/clawhdf5-wasm/tests/lazy.rs @@ -0,0 +1,391 @@ +//! The restartable ("NeedBytes") reader against the in-memory one: every +//! file must list, describe and read the same through a [`LazyStorage`] +//! that starts empty and is fed only the ranges its passes ask for, as the +//! browser's `openUrl` feeds it from HTTP range requests. +//! +//! - Files written here with `FileBuilder`, at several block sizes (512 B +//! blocks make almost every structure read a miss). +//! - The h5py/netCDF4 fixture of `examples/wasm-viewer/test/make_fixture.py` +//! (skipped without h5py, unless `CLAWHDF5_REQUIRE_INTEROP=1`; +//! `CLAWHDF5_PYTHON` names the interpreter). +//! - `CLAWHDF5_WASM_CORPUS=dir[:dir...]`: every HDF5 file under those +//! directories up to 64 MiB (e.g. `conformance/.cache/corpus`). +//! +//! Also the request budget: listing and reading one small dataset of a large +//! file fetches a few blocks, not the file. + +use std::ops::Range; +use std::path::{Path, PathBuf}; +use std::process::Command; +use std::sync::Arc; + +use clawhdf5::{AttrValue, FileBuilder}; +use clawhdf5_format::storage::CountingStorage; +use clawhdf5_wasm::core::{Hyperslab, Kind, Reader}; +use clawhdf5_wasm::lazy::{LazyConfig, LazyStorage}; + +/// The API the JavaScript side calls, one operation at a time. +trait Api { + fn call(&self, op: impl Fn(&Reader) -> T) -> T; +} + +struct Local(Reader); + +impl Api for Local { + fn call(&self, op: impl Fn(&Reader) -> T) -> T { + op(&self.0) + } +} + +/// A lazily read file and the "server" it fetches from. +struct Lazy { + data: Arc>, + storage: Arc, + reader: Reader, +} + +fn fetch(data: &[u8], r: Range) -> Result, String> { + Ok(data[r.start as usize..r.end as usize].to_vec()) +} + +impl Lazy { + /// Open as `openUrl` does: the first block comes with the probe that + /// learns the length, then the open is run until it has its bytes. + fn open(data: Vec, config: LazyConfig) -> Result { + let data = Arc::new(data); + let storage = Arc::new(LazyStorage::new(data.len() as u64, config)); + let first = (storage.config().block_size as usize).min(data.len()); + storage.supply(0, &data[..first])?; + let s = storage.clone(); + let reader = + storage.run_blocking(|| Reader::open_storage(s.clone()), |r| fetch(&data, r))??; + Ok(Lazy { + data, + storage, + reader, + }) + } +} + +impl Api for Lazy { + fn call(&self, op: impl Fn(&Reader) -> T) -> T { + self.storage + .run_blocking(|| op(&self.reader), |r| fetch(&self.data, r)) + .expect("serving from memory cannot fail") + } +} + +/// Everything the viewer can show of a file, as text: each object's kind, +/// listing, attributes (and attribute errors), dataset info, whole value +/// and a hyperslab — or the error each gives. +fn transcript(api: &impl Api) -> Vec { + let mut out = Vec::new(); + let mut todo = vec![("/".to_string(), 0usize)]; + while let Some((path, depth)) = todo.pop() { + if out.len() > 4000 { + out.push("... (truncated)".into()); + break; + } + let kind = api.call(|r| r.kind(&path)); + out.push(format!("{path}: {kind:?}")); + out.push(format!("{path} attrs: {:?}", api.call(|r| r.attrs(&path)))); + match kind { + Ok(Kind::Group) => { + let list = api.call(|r| r.list(&path)); + out.push(format!("{path} list: {list:?}")); + if let Ok(children) = list + && depth < 12 + { + for c in children.into_iter().rev() { + let child = if path == "/" { + format!("/{}", c.name) + } else { + format!("{path}/{}", c.name) + }; + todo.push((child, depth + 1)); + } + } + } + Ok(Kind::Dataset) => { + let info = api.call(|r| r.info(&path)); + out.push(format!("{path} info: {info:?}")); + let Ok(info) = info else { continue }; + let n = info + .shape + .iter() + .chain(&info.element_shape) + .try_fold(1u64, |a, &d| a.checked_mul(d)); + if n.is_none_or(|n| n > 4_000_000) { + out.push(format!("{path}: not read ({n:?} values)")); + continue; + } + out.push(format!( + "{path} read: {:?}", + api.call(|r| r.read(&path, None)) + )); + if !info.shape.is_empty() && info.shape.iter().all(|&d| d > 1) { + let slab = Hyperslab { + start: info.shape.iter().map(|_| 1).collect(), + count: info.shape.iter().map(|&d| d / 2).collect(), + stride: None, + block: None, + }; + let part = api.call(|r| r.read(&path, Some(&slab))); + out.push(format!("{path} slab: {part:?}")); + } + } + Err(_) => {} + } + } + out +} + +/// The lazy transcript of `data` at `block` bytes per block equals the +/// transcript of the same file through a range storage that has every byte +/// (`CountingStorage`: the facade's `Storage` path, the one the lazy reader +/// takes), and agrees with the in-memory one: the same values, and an error +/// wherever it has one (a malformed file can fail at a different check, +/// with a different message, when read by ranges). Returns what the lazy +/// reader fetched and its transcript. +fn check_equal(name: &str, data: &[u8], block: u64) -> (u64, u64, Vec) { + let ctx = format!("{name} (blocks of {block} B)"); + let ranged = Reader::open_storage(Arc::new(CountingStorage::new(data.to_vec()))); + let local = Reader::open(data.to_vec()); + let lazy = Lazy::open(data.to_vec(), config(block)); + let (ranged, local, lazy) = match (ranged, local, lazy) { + (Ok(r), Ok(l), Ok(z)) => (r, l, z), + (Err(r), Err(_), Err(z)) => { + assert_eq!(z, r, "{ctx}: open error"); + return (0, 0, Vec::new()); + } + (r, l, z) => panic!( + "{ctx}: opens differently: ranged {:?}, in memory {:?}, lazily {:?}", + r.err(), + l.err(), + z.err() + ), + }; + let got = transcript(&lazy); + let want = transcript(&Local(ranged)); + for (i, (w, g)) in want.iter().zip(&got).enumerate() { + assert_eq!(g, w, "{ctx}, line {i}"); + } + assert_eq!(got.len(), want.len(), "{ctx}: transcript length"); + let local = transcript(&Local(local)); + for (i, (l, g)) in local.iter().zip(&got).enumerate() { + let both_errors = match (l.split_once("Err("), g.split_once("Err(")) { + (Some((a, _)), Some((b, _))) => a == b, + _ => false, + }; + assert!( + l == g || both_errors, + "{ctx}, line {i}: in memory\n {l}\nlazily\n {g}" + ); + } + assert_eq!(got.len(), local.len(), "{ctx}: transcript length"); + let st = lazy.storage.stats(); + (st.requests, st.bytes_fetched, got) +} + +fn config(block: u64) -> LazyConfig { + LazyConfig { + block_size: block, + // A small budget, so eviction between operations is exercised. + capacity: 16 * block, + max_request: 8 * block, + } +} + +fn builder_file() -> Vec { + let mut b = FileBuilder::new(); + b.create_dataset("grid") + .with_f64_data(&(0..20_000).map(f64::from).collect::>()) + .with_shape(&[100, 200]) + .with_chunks(&[10, 25]) + .with_deflate(4); + b.create_dataset("contiguous") + .with_i32_data(&(0..50_000).collect::>()); + b.create_dataset("bytes").with_u8_data(&[1, 2, 250]); + let mut g = b.create_group("sensors"); + for i in 0..40 { + g.create_dataset(&format!("t{i}")) + .with_f32_data(&[i as f32, 1.5, -2.25]); + } + g.set_attr("location", AttrValue::String("lab".into())); + b.add_group(g.finish()); + b.set_attr("version", AttrValue::I64(3)); + b.set_attr("scale", AttrValue::F64Array(vec![0.5, 2.0])); + b.finish().unwrap() +} + +#[test] +fn builder_files_read_the_same_at_every_block_size() { + let data = builder_file(); + for block in [512, 4096, 1 << 20] { + let (requests, _, lines) = check_equal("builder", &data, block); + assert!(requests > 0); + // The transcript covers every object, values included. + assert!(lines.iter().any(|l| l.starts_with("/grid read: Ok"))); + assert!(lines.iter().any(|l| l.starts_with("/grid slab: Ok"))); + assert!(lines.iter().any(|l| l.starts_with("/sensors/t39 read: Ok"))); + } +} + +#[test] +fn garbage_fails_to_open_as_in_memory() { + check_equal("zeros", &[0u8; 5000], 512); + check_equal("empty", &[], 512); + let mut cut = builder_file(); + cut.truncate(cut.len() / 3); + check_equal("truncated", &cut, 512); +} + +/// Listing a large file and reading one small dataset fetches a few blocks, +/// not the file. +#[test] +fn a_small_read_of_a_large_file_fetches_a_few_blocks() { + let mut b = FileBuilder::new(); + b.create_dataset("small").with_f64_data(&[1.0, 2.0, 3.0]); + // 48 MB of raw data, written after the small dataset's metadata. + b.create_dataset("big") + .with_f64_data(&(0..6_000_000).map(f64::from).collect::>()); + let mut g = b.create_group("group"); + g.create_dataset("inner").with_i32_data(&[7, 8]); + b.add_group(g.finish()); + let data = b.finish().unwrap(); + let lazy = Lazy::open(data.clone(), LazyConfig::default()).unwrap(); + let list = lazy.call(|r| r.list("/")).unwrap(); + assert_eq!(list.len(), 3); + assert_eq!( + format!("{:?}", lazy.call(|r| r.read("/small", None)).unwrap().data), + "F64([1.0, 2.0, 3.0])" + ); + assert_eq!( + format!( + "{:?}", + lazy.call(|r| r.read("/group/inner", None)).unwrap().data + ), + "I32([7, 8])" + ); + // A window of the big dataset reads only its block(s). + let slab = Hyperslab { + start: vec![3_000_000], + count: vec![4], + stride: None, + block: None, + }; + assert_eq!( + format!( + "{:?}", + lazy.call(|r| r.read("/big", Some(&slab))).unwrap().data + ), + "F64([3000000.0, 3000001.0, 3000002.0, 3000003.0])" + ); + let st = lazy.storage.stats(); + eprintln!("{} bytes: {st:?}", data.len()); + assert!(st.requests <= 6, "{st:?}"); + assert!(st.bytes_fetched <= 6 << 20, "{st:?}"); + assert!(st.bytes_fetched * 8 < data.len() as u64, "{st:?}"); +} + +fn python() -> String { + std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string()) +} + +fn python_available() -> bool { + Command::new(python()) + .args(["-c", "import h5py, netCDF4, numpy"]) + .output() + .is_ok_and(|o| o.status.success()) +} + +#[test] +fn h5py_and_netcdf4_files_read_the_same_lazily() { + if !python_available() { + assert!( + !std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"), + "CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy", + python() + ); + eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python()); + return; + } + let dir = tempfile::tempdir().unwrap(); + let generator = Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../../examples/wasm-viewer/test/make_fixture.py"); + let out = Command::new(python()) + .arg(&generator) + .arg(dir.path()) + .output() + .unwrap(); + assert!( + out.status.success(), + "{}", + String::from_utf8_lossy(&out.stderr) + ); + for name in ["fixture.h5", "fixture.nc"] { + let data = std::fs::read(dir.path().join(name)).unwrap(); + for block in [512, 64 * 1024] { + let (_, _, lines) = check_equal(name, &data, block); + assert!(lines.iter().filter(|l| l.contains(" read: Ok")).count() >= 2); + } + } +} + +fn hdf5_files(dir: &Path, out: &mut Vec) { + let Ok(entries) = std::fs::read_dir(dir) else { + return; + }; + for e in entries.flatten() { + let p = e.path(); + if p.is_dir() { + hdf5_files(&p, out); + } else if std::fs::read(&p) + .ok() + .is_some_and(|b| b.len() <= 64 << 20 && is_hdf5(&b)) + { + out.push(p); + } + } +} + +/// The HDF5 signature at 0 or a power-of-two user-block offset. +fn is_hdf5(b: &[u8]) -> bool { + const SIG: &[u8] = b"\x89HDF\r\n\x1a\n"; + let mut at = 0usize; + loop { + if b.get(at..at + 8) == Some(SIG) { + return true; + } + at = if at == 0 { 512 } else { at * 2 }; + if at >= b.len() { + return false; + } + } +} + +#[test] +fn corpus_files_read_the_same_lazily() { + let Ok(dirs) = std::env::var("CLAWHDF5_WASM_CORPUS") else { + eprintln!("CLAWHDF5_WASM_CORPUS not set; skipping the corpus"); + return; + }; + let mut files = Vec::new(); + for d in std::env::split_paths(&dirs) { + hdf5_files(&d, &mut files); + } + files.sort(); + assert!(!files.is_empty(), "no HDF5 files under {dirs}"); + let (mut requests, mut bytes, mut total) = (0u64, 0u64, 0u64); + for f in &files { + let data = std::fs::read(f).unwrap(); + total += data.len() as u64; + let (r, b, _) = check_equal(&f.display().to_string(), &data, 64 * 1024); + requests += r; + bytes += b; + } + eprintln!( + "{} files ({total} bytes): {requests} requests, {bytes} bytes fetched", + files.len() + ); +} -- 2.54.0 From 910d81904c1ef5e48977a8095d1b6049967c9032 Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 06:40:44 -0500 Subject: [PATCH 04/30] py: remote files (clawhdf5.File(url), File.open_url) through File::storage() The Python bindings could not open a remote file: they parsed through File::as_bytes() in eight places (path lookups, object headers, dataspaces, attributes, group listings, the global heap of variable-length data), which a storage-backed file does not have. - Every object of a File now shares one handle (src/handle.rs) that runs all file access, metadata included, with the GIL released and parses through File::storage() and the clawhdf5_format *_in functions. Local files take the same path (their storage is the mmap). - clawhdf5.File(url) opens any scheme://... through clawhdf5_remote::storage_for_url (read-only; another mode is a ValueError). File.open_url(url, **options) takes the cache and HTTP options (block_size, cache_size, headers, retries, timeout, allow_full_download, max_full_download, require_validator, max_redirects, max_parallel); File.remote_stats gives the block cache's counters. - Default build: plain HTTP only, no C. https (rustls/ring) and s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and ci-test.sh's no-C check now covers the crate. - A failed storage read (network error, file changed on the server) is an OSError, never KeyError/ValueError and never data; `key in group` raises it instead of answering False. Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and 1 KiB blocks) against a range-capable http.server in the test process (conftest.RangeServer); test_remote.py covers request counts, cache hits, a server without Range support, a changed file, a server that hangs up, 16 threads, and a spinning thread that keeps running while a read waits on 0.2 s requests. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 30 +++ README.md | 12 + crates/clawhdf5-py/Cargo.toml | 9 + crates/clawhdf5-py/README.md | 35 ++- crates/clawhdf5-py/src/attrs.rs | 42 +-- crates/clawhdf5-py/src/convert.rs | 70 +++-- crates/clawhdf5-py/src/dataset.rs | 144 +++++++---- crates/clawhdf5-py/src/file.rs | 206 ++++++++++++--- crates/clawhdf5-py/src/group.rs | 151 +++++------ crates/clawhdf5-py/src/handle.rs | 130 ++++++++++ crates/clawhdf5-py/src/lib.rs | 7 +- crates/clawhdf5-py/src/node.rs | 165 ++++++------ crates/clawhdf5-py/tests/conftest.py | 135 ++++++++++ crates/clawhdf5-py/tests/test_read_vs_h5py.py | 25 +- crates/clawhdf5-py/tests/test_remote.py | 244 ++++++++++++++++++ docs/design/range-reads.md | 18 +- docs/known-issues.md | 15 +- scripts/ci-test.sh | 7 +- 18 files changed, 1146 insertions(+), 299 deletions(-) create mode 100644 crates/clawhdf5-py/src/handle.rs create mode 100644 crates/clawhdf5-py/tests/test_remote.py diff --git a/CHANGELOG.md b/CHANGELOG.md index 664a831..6b153b2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,36 @@ ## Unreleased +### Python bindings: remote files (2026-09-27) +- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`, + `gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`, + `azure` features) through `clawhdf5-remote`'s `open_url`: range requests + through the block cache, the whole read API (groups, attributes, every + dataset type and index the local reader handles). A URL is any + `scheme://…`; a remote file is read-only (another mode is a + `ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes + `block_size`, `cache_size`, `headers`, `retries`, `timeout`, + `allow_full_download`, `max_full_download`, `require_validator`, + `max_redirects` and `max_parallel`; `File.remote_stats` gives the block + cache's counters. The default wheel builds plain HTTP only (no C: rustls + needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C + check now covers `clawhdf5-py`. +- **Every read parses through `File::storage()`** instead of + `File::as_bytes()` (path lookups, object headers, dataspaces, attributes, + group listings, variable-length data through the global heap), inside one + shared file handle that releases the GIL for all file access, not only + dataset reads: a read waiting on the network lets other Python threads + run. A failed read of the storage (a network error, a file changed on + the server) is an `OSError`, never a `KeyError`/`ValueError` and never + data; `key in group` raises it instead of answering `False`. +- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB + blocks and 1 KiB blocks) against a range-capable `http.server` in the + test process, plus `tests/test_remote.py`: request counts of a small + read, cache hits, a server without `Range` support (refused, or a + whole download when allowed), a file changed on the server, a server + that hangs up, 16 threads on one remote file, and a thread that keeps + running while a read waits on 0.2 s requests. + ### Range reads, milestone M3: remote files (2026-09-26) - **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives a `clawhdf5::File` (through `File::open_storage`) that reads the file by diff --git a/README.md b/README.md index 3301b09..dccee48 100644 --- a/README.md +++ b/README.md @@ -533,8 +533,20 @@ with clawhdf5.File("data.h5", "r") as f: records = f["table"] # compound -> numpy structured array ids = records["id"] # one field + +# A file on a web server: range requests through a block cache, nothing +# downloaded up front; the same read API. The GIL is released while waiting. +with clawhdf5.File("http://data.example.org/run42.h5") as f: + first = f["group/temperatures"][0] +f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024, + headers={"Authorization": "Bearer ..."}) ``` +The default build reads `http://` URLs only; build with +`maturin develop --release --features https` (rustls with ring, which +compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for +object-store URLs. + Reads cover integers and IEEE floats of every width in either byte order, `bool`, enums, complex, fixed and variable-length strings, variable-length sequences, opaque, HDF5 array types and compounds; other types (references, diff --git a/crates/clawhdf5-py/Cargo.toml b/crates/clawhdf5-py/Cargo.toml index 3513a62..8246440 100644 --- a/crates/clawhdf5-py/Cargo.toml +++ b/crates/clawhdf5-py/Cargo.toml @@ -17,11 +17,20 @@ crate-type = ["cdylib", "rlib"] [dependencies] clawhdf5_rs = { path = "../clawhdf5", version = "2.7.0", package = "clawhdf5" } clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" } +# Remote files (`clawhdf5.File(url)`): plain HTTP by default, which builds +# no C. HTTPS and the object stores are opt-in features below. +clawhdf5-remote = { path = "../clawhdf5-remote", version = "2.7.0" } pyo3 = "0.29" numpy = "0.29" [features] extension-module = ["pyo3/extension-module"] +# https:// URLs (rustls with ring, which compiles C and assembly). +https = ["clawhdf5-remote/https"] +# s3://, gs://, az:// URLs (object_store; its cloud clients build aws-lc-rs, C). +s3 = ["clawhdf5-remote/s3"] +gcs = ["clawhdf5-remote/gcs"] +azure = ["clawhdf5-remote/azure"] [package.metadata.docs.rs] features = [] diff --git a/crates/clawhdf5-py/README.md b/crates/clawhdf5-py/README.md index 95cae12..16e4169 100644 --- a/crates/clawhdf5-py/README.md +++ b/crates/clawhdf5-py/README.md @@ -61,6 +61,36 @@ with clawhdf5.File("data.h5", "r") as f: - Attributes return what h5py returns; `clawhdf5.Empty` stands for a null dataspace (h5py's `Empty`). +## Remote files + +A URL instead of a path reads the file where it is, through +`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks, +64 MiB budget by default), fetching only the blocks a read needs. The whole +read API works the same, and the GIL is released while waiting on the +network. + +```python +f = clawhdf5.File("http://host/data.h5") # default options +f = clawhdf5.File.open_url( + "http://host/data.h5", + block_size=256 * 1024, cache_size=128 << 20, # the block cache + headers={"Authorization": "Bearer ..."}, # sent to this origin only + retries=3, timeout=30.0, max_redirects=5, max_parallel=8, + allow_full_download=False, # a server without Range support: refuse + require_validator=False, # refuse servers without ETag/Last-Modified +) +f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...} +``` + +- The file is pinned when opened (ETag or Last-Modified, and length): if it + changes on the server, reads raise `OSError` instead of mixing versions. + Network failures are `OSError` too. +- Remote files are read-only. +- Schemes: the default build (no C) reads `http://`. `https://` needs + `maturin develop --release --features https` (rustls with ring, which + compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and + `azure` features (credentials from the environment; aws-lc-rs, C). + ## Writing `clawhdf5.File(path, "w")` with `create_dataset(name, data=array, @@ -76,7 +106,10 @@ pytest crates/clawhdf5-py/tests ``` `tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py -writes. `scripts/ci-test.sh` builds the wheel and runs these in CI. +writes, opened locally and over HTTP (an in-process range server, +`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests, +failures, the GIL). `scripts/ci-test.sh` builds the wheel and runs these +in CI. ## License diff --git a/crates/clawhdf5-py/src/attrs.rs b/crates/clawhdf5-py/src/attrs.rs index fd3f0b3..615214b 100644 --- a/crates/clawhdf5-py/src/attrs.rs +++ b/crates/clawhdf5-py/src/attrs.rs @@ -8,13 +8,14 @@ use pyo3::prelude::*; use pyo3::types::{PyList, PyTuple}; use crate::convert::{Converter, Elements, resolve_vl}; +use crate::handle::Handle; use crate::{OwnedAttrValue, PyEmpty, attr_value_to_py, node, py_to_attr_value}; /// Backing storage for attributes. enum AttrsInner { /// Attributes of an object in a file opened for reading, sorted by name. Read { - file: Arc, + handle: Arc, attrs: Vec, }, /// Writable attribute list shared with a parent (PyFile or PyGroup). @@ -36,10 +37,15 @@ pub struct PyAttrs { impl PyAttrs { /// The attributes of the object at `addr` (whose path is `path`) in a /// file opened for reading. - pub(crate) fn read(file: Arc, addr: u64, path: &str) -> PyResult { - let attrs = node::attributes(&file, addr, path)?; + pub(crate) fn read( + py: Python<'_>, + handle: Arc, + addr: u64, + path: &str, + ) -> PyResult { + let attrs = handle.with(py, |f| node::attributes(f, addr, path))?; Ok(Self { - inner: AttrsInner::Read { file, attrs }, + inner: AttrsInner::Read { handle, attrs }, }) } @@ -55,8 +61,8 @@ impl PyAttrs { impl PyAttrs { fn __getitem__(&self, py: Python<'_>, key: &str) -> PyResult> { match &self.inner { - AttrsInner::Read { file, attrs } => match attrs.iter().find(|a| a.name == key) { - Some(attr) => Ok(attr_to_py(py, file, attr)?.unbind()), + AttrsInner::Read { handle, attrs } => match attrs.iter().find(|a| a.name == key) { + Some(attr) => Ok(attr_to_py(py, handle, attr)?.unbind()), None => Err(PyKeyError::new_err(format!( "Can't open attribute (can't locate attribute: '{key}')" ))), @@ -146,9 +152,9 @@ impl PyAttrs { /// Return attribute values as a list. fn values(&self, py: Python<'_>) -> PyResult> { let vals: Vec> = match &self.inner { - AttrsInner::Read { file, attrs } => attrs + AttrsInner::Read { handle, attrs } => attrs .iter() - .map(|a| attr_to_py(py, file, a).map(Bound::unbind)) + .map(|a| attr_to_py(py, handle, a).map(Bound::unbind)) .collect::>()?, AttrsInner::Write(store) => store .lock() @@ -167,9 +173,9 @@ impl PyAttrs { /// Return attribute (key, value) pairs as a list of tuples. fn items(&self, py: Python<'_>) -> PyResult> { let pairs: Vec<(String, Py)> = match &self.inner { - AttrsInner::Read { file, attrs } => attrs + AttrsInner::Read { handle, attrs } => attrs .iter() - .map(|a| Ok((a.name.clone(), attr_to_py(py, file, a)?.unbind()))) + .map(|a| Ok((a.name.clone(), attr_to_py(py, handle, a)?.unbind()))) .collect::>()?, AttrsInner::Write(store) => store .lock() @@ -189,12 +195,11 @@ impl PyAttrs { /// An attribute's value as h5py returns it. fn attr_to_py<'py>( py: Python<'py>, - file: &clawhdf5_rs::File, + handle: &Handle, attr: &AttributeMessage, ) -> PyResult> { crate::no_panic(|| { - let sb = file.superblock(); - let conv = Converter::new(py, &attr.datatype, sb.offset_size) + let conv = Converter::new(py, &attr.datatype, handle.offset_size) .map_err(|e| prefix_err(py, &attr.name, e))?; if node::is_null(&attr.dataspace) { return Ok(PyEmpty::new(conv.dtype).into_pyobject(py)?.into_any()); @@ -216,12 +221,11 @@ fn attr_to_py<'py>( ))); } let raw = &attr.raw_data[..want]; - let file_data = file.as_bytes(); - let (osz, lsz, unit) = (sb.offset_size, sb.length_size, conv.vl_unit); - Elements::Vl( - py.detach(|| resolve_vl(file_data, raw, n, osz, lsz, unit)) - .map_err(|e| PyValueError::new_err(format!("attribute {}: {e}", attr.name)))?, - ) + let (osz, lsz, unit) = (handle.offset_size, handle.length_size, conv.vl_unit); + let what = format!("attribute {}", attr.name); + Elements::Vl(handle.with(py, |f| { + resolve_vl(f.storage(), raw, n, osz, lsz, unit).map_err(|e| e.into_py(&what)) + })?) } else { Elements::Bytes(attr.raw_data.clone()) }; diff --git a/crates/clawhdf5-py/src/convert.rs b/crates/clawhdf5-py/src/convert.rs index 4ca7546..c8c3d68 100644 --- a/crates/clawhdf5-py/src/convert.rs +++ b/crates/clawhdf5-py/src/convert.rs @@ -17,6 +17,7 @@ use std::collections::HashMap; use clawhdf5_format::datatype::{CharacterSet, Datatype, DatatypeByteOrder}; use clawhdf5_format::global_heap::GlobalHeapCollection; +use clawhdf5_format::storage::Storage; use numpy::PyArray1; use pyo3::exceptions::{PyTypeError, PyValueError}; use pyo3::prelude::*; @@ -538,16 +539,16 @@ fn object_array<'py>( /// their bytes: each element's stored length times `unit` (1 for strings, /// the base type's size for sequences). Pure Rust, so it runs without the /// GIL. -pub(crate) fn resolve_vl( - file_data: &[u8], +pub(crate) fn resolve_vl( + file: &S, raw: &[u8], count: usize, offset_size: u8, length_size: u8, unit: usize, -) -> Result>, String> { +) -> Result>, VlError> { let refs = clawhdf5_format::vl_data::parse_vl_references(raw, count as u64, offset_size) - .map_err(|e| e.to_string())?; + .map_err(|e| VlError::Invalid(e.to_string()))?; let undefined = match offset_size { 2 => 0xFFFF, 4 => 0xFFFF_FFFF, @@ -558,43 +559,70 @@ pub(crate) fn resolve_vl( for vl in &refs { if vl.collection_address == 0 || vl.collection_address == undefined { if vl.length != 0 { - return Err(format!( + return Err(VlError::Invalid(format!( "variable-length element of length {} has no heap address", vl.length - )); + ))); } out.push(Vec::new()); continue; } let coll = match collections.entry(vl.collection_address) { std::collections::hash_map::Entry::Occupied(e) => e.into_mut(), - std::collections::hash_map::Entry::Vacant(e) => { - let addr = usize::try_from(vl.collection_address) - .map_err(|_| "global heap address out of range".to_string())?; - e.insert( - GlobalHeapCollection::parse(file_data, addr, length_size) - .map_err(|e| e.to_string())?, - ) - } + std::collections::hash_map::Entry::Vacant(e) => e.insert( + GlobalHeapCollection::parse_in(file, vl.collection_address, length_size) + .map_err(VlError::from_format)?, + ), }; - let index = u16::try_from(vl.object_index) - .map_err(|_| format!("global heap object index {} out of range", vl.object_index))?; + let index = u16::try_from(vl.object_index).map_err(|_| { + VlError::Invalid(format!( + "global heap object index {} out of range", + vl.object_index + )) + })?; let obj = coll.get_object(index).ok_or_else(|| { - format!( + VlError::Invalid(format!( "global heap object {index} not found in the collection at {}", vl.collection_address - ) + )) })?; let need = (vl.length as usize) .checked_mul(unit) - .ok_or("variable-length element too long")?; + .ok_or_else(|| VlError::Invalid("variable-length element too long".into()))?; if need > obj.data.len() { - return Err(format!( + return Err(VlError::Invalid(format!( "variable-length element of {need} bytes in a {}-byte heap object", obj.data.len() - )); + ))); } out.push(obj.data[..need].to_vec()); } Ok(out) } + +/// Why variable-length elements could not be resolved. +#[derive(Debug)] +pub(crate) enum VlError { + /// Reading the file failed (a network error on a remote file). + Storage(String), + /// The references or the heap are not valid. + Invalid(String), +} + +impl VlError { + fn from_format(e: clawhdf5_format::error::FormatError) -> Self { + match e { + clawhdf5_format::error::FormatError::Storage(_) => VlError::Storage(e.to_string()), + e => VlError::Invalid(e.to_string()), + } + } + + /// As a Python exception, the message prefixed with `what`: a storage + /// failure is an `OSError`, anything else a `ValueError`. + pub(crate) fn into_py(self, what: &str) -> PyErr { + match self { + VlError::Storage(m) => pyo3::exceptions::PyOSError::new_err(format!("{what}: {m}")), + VlError::Invalid(m) => PyValueError::new_err(format!("{what}: {m}")), + } + } +} diff --git a/crates/clawhdf5-py/src/dataset.rs b/crates/clawhdf5-py/src/dataset.rs index bcfd852..24c9b51 100644 --- a/crates/clawhdf5-py/src/dataset.rs +++ b/crates/clawhdf5-py/src/dataset.rs @@ -6,21 +6,53 @@ //! whole dataset instead); the //! bytes it returns become the numpy array's buffer without a copy (see //! `convert`). All file access and decoding runs with the GIL released, so -//! Python threads reading the same or different datasets run in parallel. +//! Python threads reading the same or different datasets run in parallel, +//! and a remote file's network reads never hold the GIL. use std::sync::Arc; use clawhdf5_format::datatype::Datatype; use clawhdf5_format::object_header::ObjectHeader; +use clawhdf5_rs::File; use pyo3::exceptions::{PyTypeError, PyValueError}; use pyo3::prelude::*; use pyo3::types::{PyList, PyTuple}; use crate::attrs::PyAttrs; -use crate::convert::{Converter, Elements, resolve_vl}; +use crate::convert::{Converter, Elements, VlError, resolve_vl}; +use crate::handle::Handle; use crate::select::{self, Plan}; use crate::{PyEmpty, node, to_py_err}; +/// What opening a dataset reads from the file (without the GIL). +pub(crate) struct DatasetMeta { + /// `None` for a dataset with a null dataspace (h5py's `Empty`). + shape: Option>, + chunks: Option>, + datatype: Datatype, +} + +impl DatasetMeta { + pub(crate) fn load(f: &File, addr: u64, hdr: &ObjectHeader, path: &str) -> PyResult { + let null = node::is_null(&node::dataspace(f, hdr, path)?); + let ds = f.dataset_at(addr).map_err(to_py_err)?; + let shape = if null { + None + } else { + Some(ds.shape().map_err(to_py_err)?) + }; + let datatype = ds.raw_datatype().map_err(to_py_err)?; + let chunks = shape + .as_ref() + .and_then(|s| node::chunk_shape(f, hdr, s.len())); + Ok(Self { + shape, + chunks, + datatype, + }) + } +} + /// A dataset in a file opened for reading. /// /// ```python @@ -30,7 +62,7 @@ use crate::{PyEmpty, node, to_py_err}; /// ``` #[pyclass(name = "Dataset")] pub struct PyDataset { - file: Arc, + handle: Arc, path: String, /// Where the dataset's object header is: reads open it from here rather /// than resolve `path` again. @@ -45,39 +77,24 @@ pub struct PyDataset { } impl PyDataset { - pub(crate) fn open( + pub(crate) fn new( py: Python<'_>, - file: Arc, + handle: Arc, path: String, addr: u64, - hdr: &ObjectHeader, - ) -> PyResult { - crate::no_panic(|| { - let null = node::is_null(&node::dataspace(&file, hdr)?); - let (shape, datatype) = { - let ds = file.dataset_at(addr).map_err(to_py_err)?; - let shape = if null { - None - } else { - Some(ds.shape().map_err(to_py_err)?) - }; - (shape, ds.raw_datatype().map_err(to_py_err)?) - }; - let conv = Converter::new(py, &datatype, file.superblock().offset_size) - .map_err(|e| e.value(py).to_string()); - let chunks = shape - .as_ref() - .and_then(|s| node::chunk_shape(&file, hdr, s.len())); - Ok(Self { - file, - path, - addr, - shape, - chunks, - datatype, - conv, - }) - }) + meta: DatasetMeta, + ) -> Self { + let conv = crate::no_panic(|| Converter::new(py, &meta.datatype, handle.offset_size)) + .map_err(|e| e.value(py).to_string()); + Self { + handle, + path, + addr, + shape: meta.shape, + chunks: meta.chunks, + datatype: meta.datatype, + conv, + } } fn converter(&self) -> PyResult<&Converter> { @@ -103,10 +120,10 @@ impl PyDataset { }; let (reads, list_axis) = plan.reads(dims, chunk_len, elem_size); let read_shape = plan.read_shape(); - let file = &*self.file; + let handle = &*self.handle; let addr = self.addr; // Everything below touches only Rust data: release the GIL. - let read = || -> Result { + let read = |file: &File| -> Result { let ds = file.dataset_at(addr)?; let mut blocks = Vec::with_capacity(reads.len()); for read in reads { @@ -147,7 +164,7 @@ impl PyDataset { let sb = file.superblock(); let n = read_shape.iter().product(); resolve_vl( - file.as_bytes(), + file.storage(), &raw, n, sb.offset_size, @@ -155,12 +172,20 @@ impl PyDataset { unit, ) .map(Elements::Vl) - .map_err(ReadError::Other) + .map_err(ReadError::Vl) }; let data = py .detach(|| { - std::panic::catch_unwind(std::panic::AssertUnwindSafe(read)) - .unwrap_or_else(|p| Err(ReadError::Panic(crate::panic_text(&*p)))) + handle + .with_detached(|f| { + Ok( + std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| read(f))) + .unwrap_or_else(|p| { + Err(ReadError::Panic(crate::panic_text(&*p))) + }), + ) + }) + .unwrap_or_else(|e| Err(ReadError::Py(e))) }) .map_err(|e| e.into_py(&self.path))?; let joined = conv.to_array(py, data, &read_shape, false)?; @@ -183,6 +208,8 @@ impl PyDataset { /// An error from the read closure, turned into a Python error with the GIL. enum ReadError { Lib(clawhdf5_rs::Error), + Vl(VlError), + Py(PyErr), Other(String), Panic(String), } @@ -197,6 +224,8 @@ impl ReadError { fn into_py(self, path: &str) -> PyErr { match self { ReadError::Lib(e) => to_py_err(e), + ReadError::Vl(e) => e.into_py(&node::name(path)), + ReadError::Py(e) => e, ReadError::Other(msg) => PyValueError::new_err(format!("{}: {msg}", node::name(path))), ReadError::Panic(msg) => crate::InternalError::new_err(format!( "{}: clawhdf5 internal error (please report it): {msg}", @@ -252,22 +281,23 @@ impl PyDataset { /// The maximum shape (`None` per unlimited dimension), like h5py. #[getter] fn maxshape<'py>(&self, py: Python<'py>) -> PyResult> { - crate::no_panic(|| { - let Some(shape) = &self.shape else { - return Ok(py.None().into_bound(py)); - }; - let max = self - .file - .dataset_at(self.addr) - .and_then(|ds| ds.max_dimensions()) - .map_err(to_py_err)? - .unwrap_or_else(|| shape.clone()); - let items: Vec> = max - .into_iter() - .map(|d| (d != u64::MAX).then_some(d)) - .collect(); - Ok(PyTuple::new(py, items)?.into_any()) - }) + let Some(shape) = &self.shape else { + return Ok(py.None().into_bound(py)); + }; + let addr = self.addr; + let max = self + .handle + .with(py, |f| { + f.dataset_at(addr) + .and_then(|ds| ds.max_dimensions()) + .map_err(to_py_err) + })? + .unwrap_or_else(|| shape.clone()); + let items: Vec> = max + .into_iter() + .map(|d| (d != u64::MAX).then_some(d)) + .collect(); + Ok(PyTuple::new(py, items)?.into_any()) } /// The dataset's numpy dtype, as h5py reports it. @@ -295,8 +325,8 @@ impl PyDataset { /// The dataset's attributes (read-only, dict-like). #[getter] - fn attrs(&self) -> PyResult { - PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path) + fn attrs(&self, py: Python<'_>) -> PyResult { + PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path) } /// Read with h5py indexing: integers, slices with positive steps, diff --git a/crates/clawhdf5-py/src/file.rs b/crates/clawhdf5-py/src/file.rs index 5f7f51b..48d94c0 100644 --- a/crates/clawhdf5-py/src/file.rs +++ b/crates/clawhdf5-py/src/file.rs @@ -1,13 +1,17 @@ //! PyFile — the main entry point for opening and creating HDF5 files. +use std::collections::HashMap; use std::path::PathBuf; use std::sync::{Arc, Mutex}; +use std::time::Duration; +use pyo3::exceptions::PyValueError; use pyo3::prelude::*; -use pyo3::types::PyList; +use pyo3::types::{PyDict, PyList}; use crate::attrs::PyAttrs; use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group}; +use crate::handle::Handle; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err}; /// Internal state for write mode. @@ -23,8 +27,10 @@ struct WriteState { /// Mirrors the h5py.File interface: /// /// ```python -/// # Reading +/// # Reading, a local file or a URL (range requests, nothing downloaded +/// # up front) /// f = clawhdf5.File('data.h5', 'r') +/// f = clawhdf5.File('https://example.org/data.h5') /// ds = f['dataset'] /// f.close() /// @@ -44,27 +50,51 @@ enum FileInner { Write(WriteState), } +/// Whether `s` is a URL (`scheme://…`) rather than a path: the scheme is a +/// letter followed by letters, digits, `+`, `-` or `.` (RFC 3986). +fn is_url(s: &str) -> bool { + let Some((scheme, _)) = s.split_once("://") else { + return false; + }; + let mut chars = scheme.chars(); + chars.next().is_some_and(|c| c.is_ascii_alphabetic()) + && chars.all(|c| c.is_ascii_alphanumeric() || matches!(c, '+' | '-' | '.')) +} + +impl PyFile { + fn from_handle(handle: Arc, filename: String) -> Self { + let root = handle.root; + Self { + inner: Some(FileInner::Read(ReadGroup::new(handle, String::new(), root))), + filename, + } + } +} + #[pymethods] impl PyFile { /// Open or create an HDF5 file. /// /// Parameters: - /// path: file path + /// path: file path, or a URL (`http://`, `https://`, `s3://`, `gs://`, + /// `az://`; which schemes work depends on how the wheel was built) + /// to read the file remotely with default options (see `open_url`) /// mode: 'r' for read (default), 'w' for write #[new] #[pyo3(signature = (path, mode="r"))] fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult { let filename = path.to_string(); - match mode { - "r" => { - let file = py.detach(|| { - crate::no_panic(|| clawhdf5_rs::File::open(path).map_err(to_py_err)) - })?; - Ok(Self { - inner: Some(FileInner::Read(root_group(Arc::new(file)))), - filename, - }) + if is_url(path) { + if mode != "r" { + return Err(PyValueError::new_err(format!( + "remote files are read-only: mode '{mode}' is not supported for a URL" + ))); } + let handle = Handle::open_url(py, path, &clawhdf5_remote::Options::default())?; + return Ok(Self::from_handle(handle, filename)); + } + match mode { + "r" => Ok(Self::from_handle(Handle::open_local(py, path)?, filename)), "w" => Ok(Self { filename, inner: Some(FileInner::Write(WriteState { @@ -74,12 +104,122 @@ impl PyFile { groups: Vec::new(), })), }), - other => Err(PyErr::new::(format!( + other => Err(PyValueError::new_err(format!( "unsupported mode '{other}'; expected 'r' or 'w'" ))), } } + /// Open a remote file for reading, with options. + /// + /// The file is read through a block cache with range requests: opening + /// costs one request (it also fetches the first block), and a read + /// fetches only the blocks it needs. The GIL is released while waiting + /// on the network. + /// + /// Parameters (all optional): + /// block_size: bytes per cached block (default 1 MiB) + /// cache_size: byte budget of the block cache (default 64 MiB) + /// headers: dict of extra HTTP headers (e.g. Authorization), sent only + /// to the URL's own origin + /// retries: retries of a request that failed transiently (default 3) + /// timeout: seconds to connect and receive response headers (default 30) + /// allow_full_download: when the server ignores Range requests, + /// download the whole file once instead of failing (default False) + /// max_full_download: largest file such a download may fetch + /// (default 1 GiB) + /// require_validator: refuse a server that sends neither ETag nor + /// Last-Modified (default False) + /// max_redirects: redirects followed per request (default 5) + /// max_parallel: requests of one read in flight at once (default 8) + #[staticmethod] + #[allow(clippy::too_many_arguments)] + #[pyo3(signature = (url, *, block_size=None, cache_size=None, headers=None, retries=None, + timeout=None, allow_full_download=None, max_full_download=None, + require_validator=None, max_redirects=None, max_parallel=None))] + fn open_url( + py: Python<'_>, + url: &str, + block_size: Option, + cache_size: Option, + headers: Option>, + retries: Option, + timeout: Option, + allow_full_download: Option, + max_full_download: Option, + require_validator: Option, + max_redirects: Option, + max_parallel: Option, + ) -> PyResult { + let mut options = clawhdf5_remote::Options::default(); + if let Some(b) = block_size { + if b == 0 { + return Err(PyValueError::new_err("block_size must be positive")); + } + options.cache.block_size = b; + options.cache.coalesce_gap = b; + // The opening request fetches the first block, not 1 MiB. + options.http.first_request = b; + } + if let Some(c) = cache_size { + options.cache.capacity = c; + } + let http = &mut options.http; + if let Some(h) = headers { + http.headers = h.into_iter().collect(); + } + if let Some(r) = retries { + http.retries = r; + } + if let Some(t) = timeout { + if !(t.is_finite() && t > 0.0) { + return Err(PyValueError::new_err("timeout must be a positive number")); + } + http.timeout = Duration::from_secs_f64(t); + } + if let Some(a) = allow_full_download { + http.allow_full_download = a; + } + if let Some(m) = max_full_download { + http.max_full_download = m; + } + if let Some(v) = require_validator { + http.require_validator = v; + } + if let Some(r) = max_redirects { + http.max_redirects = r; + } + if let Some(p) = max_parallel { + if p == 0 { + return Err(PyValueError::new_err("max_parallel must be positive")); + } + http.max_parallel = p; + } + let handle = Handle::open_url(py, url, &options)?; + Ok(Self::from_handle(handle, url.to_string())) + } + + /// For a remote file, what its block cache has done so far (reads, + /// hits, misses, requests, bytes fetched, ...); `None` for a local file. + #[getter] + fn remote_stats<'py>(&self, py: Python<'py>) -> PyResult>> { + let Some(storage) = self.read_file()?.handle.remote_storage() else { + return Ok(None); + }; + let s = storage.stats(); + let d = PyDict::new(py); + d.set_item("reads", s.reads)?; + d.set_item("hits", s.hits)?; + d.set_item("misses", s.misses)?; + d.set_item("waits", s.waits)?; + d.set_item("requests", s.requests)?; + d.set_item("fetch_calls", s.fetch_calls)?; + d.set_item("bytes_fetched", s.bytes_fetched)?; + d.set_item("evictions", s.evictions)?; + d.set_item("cached_bytes", s.cached_bytes)?; + Ok(Some(d)) + } + /// Close the file. In write mode, this finalizes and writes the file. fn close(&mut self) -> PyResult<()> { let inner = self.inner.take().ok_or_else(|| { @@ -121,7 +261,7 @@ impl PyFile { /// List the names of all children in the root group. fn keys(&self, py: Python<'_>) -> PyResult> { - let names = self.read_file()?.member_names()?; + let names = self.read_file()?.member_names(py)?; Ok(PyList::new(py, names)?.into_any().unbind()) } @@ -139,8 +279,8 @@ impl PyFile { self.keys(py)?.call_method0(py, "__iter__") } - fn __len__(&self) -> PyResult { - Ok(self.read_file()?.member_names()?.len()) + fn __len__(&self, py: Python<'_>) -> PyResult { + Ok(self.read_file()?.member_names(py)?.len()) } /// The root group's name, `/`. @@ -149,7 +289,7 @@ impl PyFile { "/" } - /// The path the file was opened with. + /// The path (or URL) the file was opened with. #[getter] fn filename(&self) -> &str { &self.filename @@ -204,9 +344,9 @@ impl PyFile { /// Attribute access. In read mode, returns attributes of the root group. /// In write mode, returns a writable attrs handle. #[getter] - fn attrs(&self) -> PyResult { + fn attrs(&self, py: Python<'_>) -> PyResult { match self.inner.as_ref() { - Some(FileInner::Read(root)) => root.attrs(), + Some(FileInner::Read(root)) => root.attrs(py), Some(FileInner::Write(state)) => Ok(PyAttrs::from_write(Arc::clone(&state.root_attrs))), None => Err(PyErr::new::( "file is closed", @@ -216,9 +356,10 @@ impl PyFile { fn __repr__(&self) -> String { match &self.inner { - Some(FileInner::Read(root)) => { - format!("", root.file.as_bytes().len()) - } + Some(FileInner::Read(root)) => match root.handle.redacted_url() { + Some(url) => format!(""), + None => format!("", self.filename), + }, Some(FileInner::Write(s)) => { format!("", s.path.display()) } @@ -226,8 +367,8 @@ impl PyFile { } } - fn __contains__(&self, key: &str) -> PyResult { - Ok(self.read_file()?.contains(key)) + fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult { + self.read_file()?.contains(py, key) } } @@ -271,11 +412,6 @@ fn parse_compression( } } -fn root_group(file: Arc) -> ReadGroup { - let root = file.superblock().root_group_address; - ReadGroup::new(file, String::new(), root) -} - /// Build and write the HDF5 file from accumulated write state. fn finalize_write(state: WriteState) -> PyResult<()> { crate::no_panic(|| { @@ -309,6 +445,18 @@ fn finalize_write(state: WriteState) -> PyResult<()> { mod tests { use super::*; + #[test] + fn urls_and_paths() { + assert!(is_url("http://h/f.h5")); + assert!(is_url("s3://bucket/key.h5")); + assert!(is_url("git+https://x")); + assert!(!is_url("data.h5")); + assert!(!is_url("/tmp/a://b.h5")); + assert!(!is_url("dir/x://y")); + assert!(!is_url("1http://x")); + assert!(!is_url("://x")); + } + #[test] fn parse_gzip_compression() { assert_eq!(parse_compression(Some("gzip"), Some(6)).unwrap(), Some(6)); diff --git a/crates/clawhdf5-py/src/group.rs b/crates/clawhdf5-py/src/group.rs index 7585e57..36506b9 100644 --- a/crates/clawhdf5-py/src/group.rs +++ b/crates/clawhdf5-py/src/group.rs @@ -3,11 +3,12 @@ use std::collections::HashMap; use std::sync::{Arc, Mutex, OnceLock}; -use pyo3::exceptions::{PyIOError, PyKeyError, PyValueError}; +use pyo3::exceptions::{PyIOError, PyKeyError, PyOSError, PyValueError}; use pyo3::prelude::*; use pyo3::types::PyList; use crate::attrs::PyAttrs; +use crate::handle::Handle; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, node}; /// Shared state for a group being written. @@ -34,9 +35,9 @@ enum GroupInner { } impl PyGroup { - pub(crate) fn from_read(file: Arc, path: String, addr: u64) -> Self { + pub(crate) fn from_read(handle: Arc, path: String, addr: u64) -> Self { Self { - inner: GroupInner::Read(ReadGroup::new(file, path, addr)), + inner: GroupInner::Read(ReadGroup::new(handle, path, addr)), } } @@ -60,9 +61,10 @@ impl PyGroup { /// h5py). It keeps its own address and, once listed, its links, so looking /// up a child neither resolves the path from the root nor scans the group's /// links again: visiting every member of a large group is linear, not -/// quadratic. +/// quadratic. (Edits never add or remove links, so these stay valid in a +/// file open for editing.) pub(crate) struct ReadGroup { - pub file: Arc, + pub handle: Arc, pub path: String, pub addr: u64, /// Link name -> object address (soft links resolved), filled on first use. @@ -72,9 +74,9 @@ pub(crate) struct ReadGroup { } impl ReadGroup { - pub(crate) fn new(file: Arc, path: String, addr: u64) -> Self { + pub(crate) fn new(handle: Arc, path: String, addr: u64) -> Self { Self { - file, + handle, path, addr, links: OnceLock::new(), @@ -82,17 +84,14 @@ impl ReadGroup { } } - fn links(&self) -> PyResult<&HashMap> { + fn links(&self, py: Python<'_>) -> PyResult<&HashMap> { if let Some(links) = self.links.get() { return Ok(links); } - let entries = crate::no_panic(|| { - clawhdf5_format::group_v2::resolve_group_children( - self.file.as_bytes(), - self.file.superblock(), - self.addr, - ) - .map_err(|e| PyValueError::new_err(format!("{}: {e}", node::name(&self.path)))) + let (addr, path) = (self.addr, &self.path); + let entries = self.handle.with(py, |f| { + clawhdf5_format::group_v2::resolve_group_children_in(f.storage(), f.superblock(), addr) + .map_err(|e| node::format_err(path, e, PyValueError::new_err)) })?; let map = entries .into_iter() @@ -102,7 +101,7 @@ impl ReadGroup { } /// The path and address of `key` (a name, a relative or an absolute path). - fn locate(&self, key: &str) -> PyResult<(String, u64)> { + fn locate(&self, py: Python<'_>, key: &str) -> PyResult<(String, u64)> { let path = node::join(&self.path, key); let rel = if self.path.is_empty() { Some(path.as_str()) @@ -112,24 +111,24 @@ impl ReadGroup { path.strip_prefix(self.path.as_str()) .and_then(|r| r.strip_prefix('/')) }; - let addr = match rel { - // A direct child: the link table, when it has the name. - Some(name) if !name.is_empty() && !name.contains('/') => { - match self.links()?.get(name) { - Some(&a) => a, - None => node::resolve_from(&self.file, self.addr, name, &path)?, - } - } - Some(rel) => node::resolve_from(&self.file, self.addr, rel, &path)?, - None => node::address(&self.file, &path)?, - }; - Ok((path, addr)) + // A direct child: the link table, when it has the name. + if let Some(name) = rel.filter(|n| !n.is_empty() && !n.contains('/')) + && let Some(&a) = self.links(py)?.get(name) + { + return Ok((path, a)); + } + let addr = self.addr; + let found = self.handle.with(py, |f| match rel { + Some(rel) => node::resolve_from(f, addr, rel, &path), + None => node::address(f, &path), + })?; + Ok((path, found)) } /// `group[key]`. pub(crate) fn get_item(&self, py: Python<'_>, key: &str) -> PyResult> { - let (path, addr) = self.locate(key)?; - node::open(py, &self.file, path, addr) + let (path, addr) = self.locate(py, key)?; + node::open(py, &self.handle, path, addr) } /// `group.get(key, default)`. @@ -148,48 +147,57 @@ impl ReadGroup { } /// Names of the group's datasets and subgroups, sorted (h5py's order). - pub(crate) fn member_names(&self) -> PyResult<&[String]> { + pub(crate) fn member_names(&self, py: Python<'_>) -> PyResult<&[String]> { if let Some(m) = self.members.get() { return Ok(m); } - let mut names = Vec::new(); - for (name, &addr) in self.links()? { - let hdr = node::header_at(&self.file, addr, &node::join(&self.path, name))?; - if matches!( - node::kind(&hdr), - Some(node::Kind::Dataset | node::Kind::Group) - ) { - names.push(name.clone()); + let links = self.links(py)?; + let path = &self.path; + let mut names = self.handle.with(py, |f| { + let mut names = Vec::new(); + for (name, &addr) in links { + if matches!( + node::kind_at(f, addr, &node::join(path, name))?, + Some(node::Kind::Dataset | node::Kind::Group) + ) { + names.push(name.clone()); + } } - } + Ok(names) + })?; names.sort_by(|a, b| a.as_bytes().cmp(b.as_bytes())); Ok(self.members.get_or_init(|| names)) } - pub(crate) fn contains(&self, key: &str) -> bool { - self.locate(key) - .and_then(|(path, addr)| node::header_at(&self.file, addr, &path)) - .ok() - .and_then(|h| node::kind(&h)) - .is_some_and(|k| k != node::Kind::Datatype) + /// `key in group`: whether `key` names a dataset or group. A failed + /// read of the file (a network error) is raised, not `False`. + pub(crate) fn contains(&self, py: Python<'_>, key: &str) -> PyResult { + let found = self + .locate(py, key) + .and_then(|(path, addr)| self.handle.with(py, |f| node::kind_at(f, addr, &path))); + match found { + Ok(kind) => Ok(kind.is_some_and(|k| k != node::Kind::Datatype)), + Err(e) if e.is_instance_of::(py) => Err(e), + Err(_) => Ok(false), + } } pub(crate) fn values(&self, py: Python<'_>) -> PyResult>> { - self.member_names()? + self.member_names(py)? .iter() .map(|n| self.get_item(py, n)) .collect() } pub(crate) fn items(&self, py: Python<'_>) -> PyResult)>> { - self.member_names()? + self.member_names(py)? .iter() .map(|n| Ok((n.clone(), self.get_item(py, n)?))) .collect() } - pub(crate) fn attrs(&self) -> PyResult { - PyAttrs::read(Arc::clone(&self.file), self.addr, &self.path) + pub(crate) fn attrs(&self, py: Python<'_>) -> PyResult { + PyAttrs::read(py, Arc::clone(&self.handle), self.addr, &self.path) } } @@ -210,7 +218,7 @@ impl PyGroup { fn keys(&self, py: Python<'_>) -> PyResult> { match &self.inner { GroupInner::Read(g) => { - let list = PyList::new(py, g.member_names()?)?; + let list = PyList::new(py, g.member_names(py)?)?; Ok(list.into_any().unbind()) } GroupInner::Write(state) => { @@ -236,9 +244,9 @@ impl PyGroup { self.keys(py)?.call_method0(py, "__iter__") } - fn __len__(&self) -> PyResult { + fn __len__(&self, py: Python<'_>) -> PyResult { match &self.inner { - GroupInner::Read(g) => Ok(g.member_names()?.len()), + GroupInner::Read(g) => Ok(g.member_names(py)?.len()), GroupInner::Write(state) => Ok(state.lock().unwrap().datasets.len()), } } @@ -301,9 +309,9 @@ impl PyGroup { /// Attribute access. #[getter] - fn attrs(&self) -> PyResult { + fn attrs(&self, py: Python<'_>) -> PyResult { match &self.inner { - GroupInner::Read(g) => g.attrs(), + GroupInner::Read(g) => g.attrs(py), GroupInner::Write(state) => { let store = Arc::clone(&state.lock().unwrap().attrs); Ok(PyAttrs::from_write(store)) @@ -311,10 +319,10 @@ impl PyGroup { } } - fn __repr__(&self) -> String { + fn __repr__(&self, py: Python<'_>) -> String { match &self.inner { GroupInner::Read(g) => { - let n = g.member_names().map_or(0, |m| m.len()); + let n = g.member_names(py).map_or(0, |m| m.len()); format!("", node::name(&g.path)) } GroupInner::Write(state) => { @@ -324,9 +332,9 @@ impl PyGroup { } } - fn __contains__(&self, key: &str) -> PyResult { + fn __contains__(&self, py: Python<'_>, key: &str) -> PyResult { match &self.inner { - GroupInner::Read(g) => Ok(g.contains(key)), + GroupInner::Read(g) => g.contains(py, key), GroupInner::Write(state) => { let guard = state.lock().unwrap(); Ok(guard.datasets.iter().any(|d| d.name == key)) @@ -357,31 +365,6 @@ pub(crate) fn finalize_write_group( mod tests { use super::*; - #[test] - fn member_names_are_sorted() { - let mut b = clawhdf5_rs::FileBuilder::new(); - b.create_dataset("zeta").with_f64_data(&[1.0]); - b.create_dataset("alpha").with_f64_data(&[1.0]); - let mut g = b.create_group("mid"); - g.create_dataset("x").with_f64_data(&[1.0]); - let finished = g.finish(); - b.add_group(finished); - let bytes = b.finish().unwrap(); - let file = Arc::new(clawhdf5_rs::File::from_bytes(bytes).unwrap()); - let root = file.superblock().root_group_address; - let top = ReadGroup::new(Arc::clone(&file), String::new(), root); - assert_eq!(top.member_names().unwrap(), ["alpha", "mid", "zeta"]); - let (path, addr) = top.locate("mid").unwrap(); - assert_eq!(path, "mid"); - let mid = ReadGroup::new(Arc::clone(&file), path, addr); - assert_eq!(mid.member_names().unwrap(), ["x"]); - assert!(top.contains("mid/x")); - assert!(mid.contains("/alpha")); - assert!(mid.contains("x") && mid.contains("./x")); - assert!(!top.contains("nope")); - assert!(!mid.contains("alpha")); - } - #[test] fn finalize_group() { let state = WriteGroupState { diff --git a/crates/clawhdf5-py/src/handle.rs b/crates/clawhdf5-py/src/handle.rs new file mode 100644 index 0000000..8649f74 --- /dev/null +++ b/crates/clawhdf5-py/src/handle.rs @@ -0,0 +1,130 @@ +//! The open file every object of a `File` shares. +//! +//! Every read goes through [`Handle::with`], which releases the GIL and +//! parses through `File::storage()`, so the same code serves a local file +//! (memory-mapped) and a remote one (`clawhdf5-remote`: range requests +//! through a block cache, so a network read never holds the GIL). +//! +//! Lock discipline (no deadlock with the GIL): the file lock is only taken +//! with the GIL released, and code that holds it never touches Python. + +use std::path::PathBuf; +use std::sync::{Arc, PoisonError, RwLock}; + +use clawhdf5_rs::File; +use pyo3::exceptions::PyOSError; +use pyo3::prelude::*; + +use crate::to_py_err; + +/// Where the file's bytes come from. +pub(crate) enum Source { + /// A local path (memory-mapped). + Local(#[allow(dead_code)] PathBuf), + /// A URL, read through `clawhdf5-remote`'s block cache. + Remote { + url: String, + storage: Arc, + }, +} + +pub(crate) struct Handle { + file: RwLock>, + source: Source, + pub offset_size: u8, + pub length_size: u8, + pub root: u64, +} + +fn closed_after_failed_reopen() -> PyErr { + PyOSError::new_err("the file could not be reopened after an edit; open it again") +} + +impl Handle { + fn new(file: File, source: Source) -> Arc { + let sb = file.superblock(); + let (offset_size, length_size, root) = + (sb.offset_size, sb.length_size, sb.root_group_address); + Arc::new(Self { + file: RwLock::new(Some(file)), + source, + offset_size, + length_size, + root, + }) + } + + /// A local file, read-only. + pub(crate) fn open_local(py: Python<'_>, path: &str) -> PyResult> { + let file = py.detach(|| crate::no_panic(|| File::open(path).map_err(to_py_err)))?; + Ok(Self::new(file, Source::Local(PathBuf::from(path)))) + } + + /// A remote file (`http(s)://`, `s3://`, ...). + pub(crate) fn open_url( + py: Python<'_>, + url: &str, + options: &clawhdf5_remote::Options, + ) -> PyResult> { + let (file, storage) = py.detach(|| { + crate::no_panic(|| { + let storage = clawhdf5_remote::storage_for_url(url, options).map_err(remote_err)?; + let file = File::open_storage(storage.clone()).map_err(to_py_err)?; + Ok((file, storage)) + }) + })?; + Ok(Self::new( + file, + Source::Remote { + url: url.to_string(), + storage, + }, + )) + } + + /// Run `f` on the file with the GIL released (a remote read may wait + /// on the network; other Python threads run meanwhile). `f` must not + /// touch Python. + pub(crate) fn with( + &self, + py: Python<'_>, + f: impl FnOnce(&File) -> PyResult + Send, + ) -> PyResult { + py.detach(|| self.with_detached(f)) + } + + /// [`with`](Self::with) for code that already runs without the GIL. + pub(crate) fn with_detached(&self, f: impl FnOnce(&File) -> PyResult) -> PyResult { + crate::no_panic(|| { + let guard = self.file.read().unwrap_or_else(PoisonError::into_inner); + let file = guard.as_ref().ok_or_else(closed_after_failed_reopen)?; + f(file) + }) + } + + /// The remote file's block cache. + pub(crate) fn remote_storage(&self) -> Option<&clawhdf5_remote::RemoteStorage> { + match &self.source { + Source::Remote { storage, .. } => Some(storage), + Source::Local(_) => None, + } + } + + /// The URL of a remote file, credentials and query values redacted. + pub(crate) fn redacted_url(&self) -> Option { + match &self.source { + Source::Remote { url, .. } => Some(clawhdf5_remote::redact_url(url)), + Source::Local(_) => None, + } + } +} + +/// A `clawhdf5_remote::Error` as a Python exception: the network side +/// (unreachable, a status, no range support, a changed file) is `OSError`, +/// a file that is not HDF5 is what `to_py_err` makes of it. +pub(crate) fn remote_err(e: clawhdf5_remote::Error) -> PyErr { + match e { + clawhdf5_remote::Error::Hdf5(e) => to_py_err(e), + other => PyOSError::new_err(other.to_string()), + } +} diff --git a/crates/clawhdf5-py/src/lib.rs b/crates/clawhdf5-py/src/lib.rs index 4055b82..db89d8d 100644 --- a/crates/clawhdf5-py/src/lib.rs +++ b/crates/clawhdf5-py/src/lib.rs @@ -14,6 +14,7 @@ mod convert; mod dataset; mod file; mod group; +mod handle; mod node; mod select; @@ -63,7 +64,7 @@ fn _panic_for_test() -> PyResult<()> { /// Convert a `clawhdf5_rs::Error` into a `PyErr`. /// /// Maps different error variants to more specific Python exception types: -/// - I/O errors -> `PyIOError` +/// - I/O errors, and failed reads of a remote file -> `PyIOError`/`PyOSError` /// - Format/parsing errors -> `PyValueError` /// - Missing dataset/path errors -> `PyKeyError` /// - Invalid arguments -> `PyValueError` @@ -73,6 +74,10 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr { use clawhdf5_rs::Error; match &e { Error::Io(_) => PyErr::new::(e.to_string()), + // A failed read of the storage: a network error on a remote file. + Error::Format(clawhdf5_format::error::FormatError::Storage(_)) => { + PyErr::new::(e.to_string()) + } Error::Format(_) => PyErr::new::(e.to_string()), Error::NotADataset(_) | Error::MissingMessage(_) => { PyErr::new::(e.to_string()) diff --git a/crates/clawhdf5-py/src/node.rs b/crates/clawhdf5-py/src/node.rs index fa9ede7..35e5dec 100644 --- a/crates/clawhdf5-py/src/node.rs +++ b/crates/clawhdf5-py/src/node.rs @@ -1,16 +1,24 @@ //! Resolving paths to objects in a file opened for reading. +//! +//! Everything here parses through `File::storage()` (the `clawhdf5_format` +//! `*_in` functions), never `File::as_bytes()`, so it works the same on a +//! memory-mapped local file and on a remote one; and it runs inside +//! `Handle::with`, without the GIL. use std::sync::Arc; use clawhdf5_format::attribute::AttributeMessage; use clawhdf5_format::dataspace::{Dataspace, DataspaceType}; +use clawhdf5_format::error::FormatError; use clawhdf5_format::message_type::MessageType; use clawhdf5_format::object_header::ObjectHeader; -use pyo3::exceptions::{PyKeyError, PyTypeError, PyValueError}; +use clawhdf5_rs::File; +use pyo3::exceptions::{PyKeyError, PyOSError, PyTypeError, PyValueError}; use pyo3::prelude::*; -use crate::dataset::PyDataset; +use crate::dataset::{DatasetMeta, PyDataset}; use crate::group::PyGroup; +use crate::handle::Handle; /// Join `key` onto the group path `base` the way h5py does: an absolute key /// starts from the root, a relative one from `base`. Paths are kept without @@ -33,42 +41,46 @@ pub(crate) fn name(path: &str) -> String { format!("/{path}") } +/// A format error met at `path`: a failed read of the storage (a network +/// error on a remote file) is an `OSError`, anything else `other(message)`. +pub(crate) fn format_err(path: &str, e: FormatError, other: fn(String) -> PyErr) -> PyErr { + let msg = format!("{}: {e}", name(path)); + match e { + FormatError::Storage(_) => PyOSError::new_err(msg), + _ => other(msg), + } +} + +fn value_err(msg: String) -> PyErr { + PyValueError::new_err(msg) +} + /// The address of the object at `path`, resolved from the root group. -pub(crate) fn address(file: &clawhdf5_rs::File, path: &str) -> PyResult { +pub(crate) fn address(file: &File, path: &str) -> PyResult { resolve_from(file, file.superblock().root_group_address, path, path) } /// The address of `rel` resolved from the group at `group` (`full` is the /// resulting path, for the error message). -pub(crate) fn resolve_from( - file: &clawhdf5_rs::File, - group: u64, - rel: &str, - full: &str, -) -> PyResult { +pub(crate) fn resolve_from(file: &File, group: u64, rel: &str, full: &str) -> PyResult { if rel.is_empty() { return Ok(group); } - crate::no_panic(|| { - clawhdf5_format::group_v2::resolve_path_from(file.as_bytes(), file.superblock(), group, rel) - .map_err(|e| { - PyKeyError::new_err(format!( - "Unable to open object (object '{}' doesn't exist): {e}", - name(full) - )) - }) - }) + clawhdf5_format::group_v2::resolve_path_from_in(file.storage(), file.superblock(), group, rel) + .map_err(|e| match e { + FormatError::Storage(_) => format_err(full, e, value_err), + e => PyKeyError::new_err(format!( + "Unable to open object (object '{}' doesn't exist): {e}", + name(full) + )), + }) } /// The object header at `addr` (the object at `path`). -pub(crate) fn header_at(file: &clawhdf5_rs::File, addr: u64, path: &str) -> PyResult { - crate::no_panic(|| { - let sb = file.superblock(); - let at = usize::try_from(addr) - .map_err(|_| PyValueError::new_err(format!("{}: address out of range", name(path))))?; - ObjectHeader::parse(file.as_bytes(), at, sb.offset_size, sb.length_size) - .map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path)))) - }) +pub(crate) fn header_at(file: &File, addr: u64, path: &str) -> PyResult { + let sb = file.superblock(); + ObjectHeader::parse_in(file.storage(), addr, sb.offset_size, sb.length_size) + .map_err(|e| format_err(path, e, value_err)) } /// What an object header describes. @@ -96,29 +108,50 @@ pub(crate) fn kind(hdr: &ObjectHeader) -> Option { } } +/// The kind of the object at `addr`, from its header. +pub(crate) fn kind_at(file: &File, addr: u64, path: &str) -> PyResult> { + Ok(kind(&header_at(file, addr, path)?)) +} + +/// What opening an object found, read without the GIL. +enum Found { + Dataset(DatasetMeta), + Group, + Datatype, + Other, +} + /// Open the object at `addr` (whose path is `path`) as a `Dataset` or /// `Group`. Both keep the address, so later reads resolve nothing. pub(crate) fn open( py: Python<'_>, - file: &Arc, + handle: &Arc, path: String, addr: u64, ) -> PyResult> { - let hdr = header_at(file, addr, &path)?; - match kind(&hdr) { - Some(Kind::Dataset) => Ok(PyDataset::open(py, Arc::clone(file), path, addr, &hdr)? + let found = handle.with(py, |f| { + let hdr = header_at(f, addr, &path)?; + Ok(match kind(&hdr) { + Some(Kind::Dataset) => Found::Dataset(DatasetMeta::load(f, addr, &hdr, &path)?), + Some(Kind::Group) => Found::Group, + Some(Kind::Datatype) => Found::Datatype, + None => Found::Other, + }) + })?; + match found { + Found::Dataset(meta) => Ok(PyDataset::new(py, Arc::clone(handle), path, addr, meta) .into_pyobject(py)? .into_any() .unbind()), - Some(Kind::Group) => Ok(PyGroup::from_read(Arc::clone(file), path, addr) + Found::Group => Ok(PyGroup::from_read(Arc::clone(handle), path, addr) .into_pyobject(py)? .into_any() .unbind()), - Some(Kind::Datatype) => Err(PyTypeError::new_err(format!( + Found::Datatype => Err(PyTypeError::new_err(format!( "{}: committed (named) datatypes are not supported by clawhdf5", name(&path) ))), - None => Err(PyValueError::new_err(format!( + Found::Other => Err(PyValueError::new_err(format!( "{}: not a dataset, group or datatype", name(&path) ))), @@ -126,32 +159,26 @@ pub(crate) fn open( } /// The dataspace message of an object header. -pub(crate) fn dataspace(file: &clawhdf5_rs::File, hdr: &ObjectHeader) -> PyResult { - crate::no_panic(|| { - let sb = file.superblock(); - let msg = hdr - .messages - .iter() - .find(|m| m.msg_type == MessageType::Dataspace) - .ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?; - let data = clawhdf5_format::shared_message::message_data( - file.as_bytes(), - msg, - sb.offset_size, - sb.length_size, - ) - .map_err(|e| PyValueError::new_err(e.to_string()))?; - Dataspace::parse(&data, sb.length_size).map_err(|e| PyValueError::new_err(e.to_string())) - }) +pub(crate) fn dataspace(file: &File, hdr: &ObjectHeader, path: &str) -> PyResult { + let sb = file.superblock(); + let msg = hdr + .messages + .iter() + .find(|m| m.msg_type == MessageType::Dataspace) + .ok_or_else(|| PyValueError::new_err("object has no dataspace message"))?; + let data = clawhdf5_format::shared_message::message_data_in( + file.storage(), + msg, + sb.offset_size, + sb.length_size, + ) + .map_err(|e| format_err(path, e, value_err))?; + Dataspace::parse(&data, sb.length_size).map_err(|e| format_err(path, e, value_err)) } /// The chunk shape of a chunked dataset (one entry per dataset dimension), /// or `None` for other layouts or a layout message that does not parse. -pub(crate) fn chunk_shape( - file: &clawhdf5_rs::File, - hdr: &ObjectHeader, - rank: usize, -) -> Option> { +pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Option> { let sb = file.superblock(); let msg = hdr .messages @@ -179,24 +206,18 @@ pub(crate) fn is_null(space: &Dataspace) -> bool { /// The attributes of the object at `addr` (whose path is `path`), sorted by /// name (h5py's order). Attributes whose messages cannot be parsed are left /// out, as the facade's `attrs()` does. -pub(crate) fn attributes( - file: &clawhdf5_rs::File, - addr: u64, - path: &str, -) -> PyResult> { +pub(crate) fn attributes(file: &File, addr: u64, path: &str) -> PyResult> { let hdr = header_at(file, addr, path)?; - crate::no_panic(|| { - let sb = file.superblock(); - let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant( - file.as_bytes(), - &hdr, - sb.offset_size, - sb.length_size, - ) - .map_err(|e| PyValueError::new_err(format!("{}: {e}", name(path))))?; - attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes())); - Ok(attrs) - }) + let sb = file.superblock(); + let (mut attrs, _errors) = clawhdf5_format::attribute::extract_attributes_tolerant_in( + file.storage(), + &hdr, + sb.offset_size, + sb.length_size, + ) + .map_err(|e| format_err(path, e, value_err))?; + attrs.sort_by(|a, b| a.name.as_bytes().cmp(b.name.as_bytes())); + Ok(attrs) } #[cfg(test)] diff --git a/crates/clawhdf5-py/tests/conftest.py b/crates/clawhdf5-py/tests/conftest.py index 967f1e0..88113e8 100644 --- a/crates/clawhdf5-py/tests/conftest.py +++ b/crates/clawhdf5-py/tests/conftest.py @@ -1,6 +1,10 @@ """Shared fixtures for the clawhdf5 Python binding tests.""" import os +import re +import threading +import time +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer import pytest @@ -16,3 +20,134 @@ def h5py(): pytest.fail("h5py is required (CLAWHDF5_REQUIRE_INTEROP=1) but not importable") pytest.skip("h5py not installed") return mod + + +# --------------------------------------------------------------------------- +# An HTTP server for remote reads +# --------------------------------------------------------------------------- + +_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$") + + +class RangeServer: + """A static file server on 127.0.0.1, in a thread of this process, that + answers `Range: bytes=a-b` with 206 and `Content-Range` (the way S3 and + common web servers do), sends an ETag and honours `If-Match`. + + - `ranges=False`: ignores `Range` and answers 200 with the whole file, + like a server without range support. + - `down` (set by `close()`): hang up on every request. + - `delay`: seconds to wait before answering each request after the + first `delay_after` ones (a slow network). + - `log`: every request as `(method, path, range header)`. + """ + + def __init__(self, root, ranges=True): + self.root = str(root) + self.ranges = ranges + self.delay = 0.0 + self.delay_after = 0 + self.down = False + self.log = [] + self._lock = threading.Lock() + server = self + + class Handler(BaseHTTPRequestHandler): + protocol_version = "HTTP/1.1" + + def log_message(self, *args): # quiet + pass + + def do_HEAD(self): + self._serve(body=False) + + def do_GET(self): + self._serve(body=True) + + def _serve(self, body): + with server._lock: + server.log.append((self.command, self.path, self.headers.get("Range"))) + n = len(server.log) + if server.delay and n > server.delay_after: + time.sleep(server.delay) + if server.down: + # Hang up without an answer (open keep-alive + # connections outlive shutdown(), so close() sets this). + self.close_connection = True + return + path = os.path.join(server.root, self.path.lstrip("/").split("?")[0]) + if not os.path.isfile(path): + self.send_response(404) + self.send_header("Content-Length", "0") + self.end_headers() + return + with open(path, "rb") as fh: + data = fh.read() + st = os.stat(path) + etag = f'"{st.st_mtime_ns:x}-{st.st_size:x}"' + want = self.headers.get("If-Match") + if want is not None and want != etag and want != "*": + self.send_response(412) + self.send_header("Content-Length", "0") + self.end_headers() + return + rng = self.headers.get("Range") if server.ranges else None + m = _RANGE.match(rng.strip()) if rng else None + if m and (m.group(1) or m.group(2)): + size = len(data) + if m.group(1): + start = int(m.group(1)) + end = int(m.group(2)) if m.group(2) else size - 1 + else: + start = max(0, size - int(m.group(2))) + end = size - 1 + if start >= size: + self.send_response(416) + self.send_header("Content-Range", f"bytes */{size}") + self.send_header("Content-Length", "0") + self.end_headers() + return + end = min(end, size - 1) + part = data[start : end + 1] + self.send_response(206) + self.send_header("Content-Range", f"bytes {start}-{end}/{size}") + else: + part = data + self.send_response(200) + if server.ranges: + self.send_header("Accept-Ranges", "bytes") + self.send_header("ETag", etag) + self.send_header("Content-Length", str(len(part))) + self.send_header("Content-Type", "application/x-hdf5") + self.end_headers() + if body: + try: + self.wfile.write(part) + except (BrokenPipeError, ConnectionResetError): + pass + + self.httpd = ThreadingHTTPServer(("127.0.0.1", 0), Handler) + self.httpd.daemon_threads = True + self.port = self.httpd.server_address[1] + self.thread = threading.Thread(target=self.httpd.serve_forever, daemon=True) + self.thread.start() + + def url(self, name): + return f"http://127.0.0.1:{self.port}/{name}" + + def requests(self): + with self._lock: + return len(self.log) + + def close(self): + self.down = True + self.httpd.shutdown() + self.httpd.server_close() + + +@pytest.fixture +def range_server(tmp_path): + """A range-capable server over `tmp_path`.""" + server = RangeServer(tmp_path) + yield server + server.close() diff --git a/crates/clawhdf5-py/tests/test_read_vs_h5py.py b/crates/clawhdf5-py/tests/test_read_vs_h5py.py index 6a3b6cb..f3132a7 100644 --- a/crates/clawhdf5-py/tests/test_read_vs_h5py.py +++ b/crates/clawhdf5-py/tests/test_read_vs_h5py.py @@ -178,15 +178,32 @@ def _write_fixture(h5py, path): g.attrs["depth"] = np.int8(3) -@pytest.fixture(scope="module") -def pair(h5py, tmp_path_factory): - path = str(tmp_path_factory.mktemp("h5") / "fixture.h5") +@pytest.fixture(scope="module", params=["local", "http", "http-1k-blocks"]) +def pair(request, h5py, tmp_path_factory): + """The fixture file through h5py and through clawhdf5: opened locally, + and over HTTP range requests (a local server in this process) with the + default 1 MiB blocks and with 1 KiB blocks, so every structure is read + through many small ranges.""" + from conftest import RangeServer + + root = tmp_path_factory.mktemp("h5") + path = str(root / "fixture.h5") _write_fixture(h5py, path) theirs = h5py.File(path, "r") - ours = clawhdf5.File(path, "r") + server = None + if request.param == "local": + ours = clawhdf5.File(path, "r") + else: + server = RangeServer(root) + if request.param == "http": + ours = clawhdf5.File(server.url("fixture.h5")) + else: + ours = clawhdf5.File.open_url(server.url("fixture.h5"), block_size=1024) yield ours, theirs, path theirs.close() ours.close() + if server is not None: + server.close() def _all_datasets(h5py, f): diff --git a/crates/clawhdf5-py/tests/test_remote.py b/crates/clawhdf5-py/tests/test_remote.py new file mode 100644 index 0000000..dbb24b8 --- /dev/null +++ b/crates/clawhdf5-py/tests/test_remote.py @@ -0,0 +1,244 @@ +"""Remote files: `clawhdf5.File(url)` / `File.open_url(url, ...)` read over +HTTP range requests (clawhdf5-remote's block cache), against a server in this +process (conftest.RangeServer). Values are compared with h5py reading the +same file locally; the rest checks what the server saw (only the blocks a +read needs are fetched), the failure modes (no range support, a missing +file, a file that changes, a server that goes away: errors, never wrong +data), and that the GIL is released while a read waits on the network.""" + +import os +import sys +import threading +import time + +import numpy as np +import pytest + +import clawhdf5 +from conftest import RangeServer + + +def _write(h5py, path): + rng = np.random.default_rng(7) + with h5py.File(path, "w") as f: + f.create_dataset("contig", data=rng.standard_normal((400, 300))) + f.create_dataset( + "chunked", + data=rng.integers(0, 1000, size=(512, 512), dtype="= 2 + + +def test_a_small_read_fetches_only_its_blocks(h5py, remote_file): + """With 4 KiB blocks, opening and reading one chunk of a 1 MB chunked + dataset costs a handful of requests and a few blocks, not the file.""" + path, url, server = remote_file + size = os.path.getsize(path) + f = clawhdf5.File.open_url(url, block_size=4096) + opened = server.requests() + assert opened == 1, server.log + ds = f["chunked"] + got = ds[0:10, 0:10] + with h5py.File(path, "r") as theirs: + np.testing.assert_array_equal(got, theirs["chunked"][0:10, 0:10]) + stats = f.remote_stats + assert stats["bytes_fetched"] < size / 4, (stats, size) + assert server.requests() - opened <= 12, server.log + # A second read of the same region is served by the cache. + before = server.requests() + ds[0:10, 0:10] + assert server.requests() == before + assert f.remote_stats["hits"] > stats["hits"] + assert clawhdf5.File(str(path), "r").remote_stats is None + + +def test_server_without_range_support(h5py, tmp_path): + """A server that ignores Range answers 200 with the whole file: that is + an OSError by default, and a whole download when allowed.""" + path = tmp_path / "remote.h5" + _write(h5py, str(path)) + server = RangeServer(tmp_path, ranges=False) + try: + url = server.url("remote.h5") + with pytest.raises(OSError, match="range"): + clawhdf5.File(url) + with clawhdf5.File.open_url(url, allow_full_download=True) as ours, h5py.File(path, "r") as theirs: + np.testing.assert_array_equal(ours["chunked"][...], theirs["chunked"][...]) + np.testing.assert_array_equal(ours["contig"][5], theirs["contig"][5]) + with pytest.raises(OSError): + clawhdf5.File.open_url(url, allow_full_download=True, max_full_download=1000) + finally: + server.close() + + +def test_errors_are_oserrors(remote_file): + _, url, server = remote_file + with pytest.raises(OSError, match="404"): + clawhdf5.File(server.url("missing.h5")) + with pytest.raises(ValueError, match="read-only"): + clawhdf5.File(url, "r+") + with pytest.raises(ValueError, match="read-only"): + clawhdf5.File(url, "w") + with pytest.raises(OSError, match="unsupported URL"): + clawhdf5.File("nosuchscheme://x/y.h5") + with pytest.raises(ValueError): + clawhdf5.File.open_url(url, block_size=0) + with pytest.raises(TypeError): + clawhdf5.File.open_url(url, no_such_option=1) + + +def test_object_store_urls_need_their_features(): + """The default wheel has no S3/GCS/Azure clients (aws-lc-rs builds C): + such a URL is an OSError naming the build feature.""" + for url, feature in [("s3://bucket/k.h5", "s3"), ("gs://b/k.h5", "gcs"), ("az://c/k.h5", "azure")]: + try: + clawhdf5.File(url) + except OSError as e: + if "feature" in str(e): + assert f"`{feature}`" in str(e), str(e) + else: + pytest.fail(f"{url} opened") + + +def test_https_needs_the_https_feature(): + """The default wheel has no TLS stack (rustls needs ring, which builds C): + an https URL is an OSError that names the build feature.""" + with pytest.raises(OSError) as e: + clawhdf5.File("https://127.0.0.1:1/x.h5") + msg = str(e.value) + # Built with `--features https` the error is the refused connection. + assert "https" in msg or "connect" in msg.lower() or "refused" in msg.lower(), msg + + +def test_a_changed_file_is_an_error_not_mixed_data(h5py, remote_file): + path, url, _ = remote_file + f = clawhdf5.File.open_url(url, block_size=1024) + first = f["grp/small"][...] + # Rewrite the file with other values: new ETag, same name. + time.sleep(0.01) + with h5py.File(path, "w") as g: + g.create_dataset("contig", data=np.zeros((400, 300))) + with pytest.raises(OSError, match="changed"): + f["contig"][...] + np.testing.assert_array_equal(first, np.arange(10, dtype=" before + assert took >= 0.2, took + assert progress["n"] > 1000, progress + # Held across a 0.2 s request, the spinner would stall that long. + assert progress["worst"] < 0.1, (progress, took) + + +def test_a_clawhdf5_written_file_reads_the_same_remotely(tmp_path, range_server): + path = tmp_path / "ours.h5" + data = np.arange(3000, dtype=" Promise` backed by `fetch` with a diff --git a/docs/known-issues.md b/docs/known-issues.md index e97f602..bf4ae74 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -844,9 +844,9 @@ cache, but: `read_*_zerocopy`) need the file in memory and answer `FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()` panics for such a file (`File::contiguous_bytes()` is the fallible form). - `LazyFile`, `MmapFile` and the Python and wasm bindings still read a - whole file (`h5rs` reads through `File::storage`, and takes URLs with its - `remote` feature). + `LazyFile`, `MmapFile` and the wasm bindings still read a whole file + (`h5rs` and the Python bindings read through `File::storage`, and take + URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`). - The file's length is read once, at open: a growing file (SWMR) is not followed (milestone M5). A remote file is pinned at open, so one that grows is `RemoteError::FileChanged`. @@ -862,9 +862,12 @@ cache, but: **Status:** open (added 2026-09-26, milestone M3 of `docs/design/range-reads.md`). -- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3) - parses through `File::as_bytes`, which a remote file does not have; the - wasm reader's `openUrl` is milestone M4. +- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is + milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but + the default wheel reads plain `http://` only: `https://` needs a wheel + built with `--features https` (rustls with ring, which compiles C), and + `s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs). + The Python tests run against an in-process `http.server` only. - **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise). The design's policy of using a paged file's page size as the block size is not implemented, and only the first block is read ahead. diff --git a/scripts/ci-test.sh b/scripts/ci-test.sh index 4f678ad..5731358 100755 --- a/scripts/ci-test.sh +++ b/scripts/ci-test.sh @@ -131,14 +131,17 @@ run_step "cargo clippy (h5rs remote)" cargo clippy \ # js-sys (clawhdf5-wasm's bindings to JavaScript) builds no C. # clawhdf5-remote is checked by default (plain HTTP) and with its # object-store feature, and h5rs with URL support (remote); the https -# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. +# (ring) and s3/gcs/azure (aws-lc-rs) features build C and are opt-in. The +# Python bindings (clawhdf5-py, remote reads over plain HTTP) are checked too: +# their https/s3/gcs/azure features are opt-in for the same reason. no_c_in_default_build() { local entry crate features found=0 for entry in clawhdf5-format clawhdf5-io clawhdf5-filters clawhdf5 \ clawhdf5-agent clawhdf5-ann clawhdf5-accel clawhdf5-netcdf4 clawhdf5-cli \ clawhdf5-tools \ clawhdf5-wasm \ - clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote; do + clawhdf5-remote clawhdf5-remote:object-store clawhdf5-tools:remote \ + clawhdf5-py; do crate=${entry%%:*} features=() [ "$entry" != "$crate" ] && features=(--features "${entry#*:}") -- 2.54.0 From 5107583b97c8e722b9e6c51e422091f3f4a2018c Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 06:44:45 -0500 Subject: [PATCH 05/30] wasm: openUrl reads remote files by HTTP range requests (range-read M4) openUrl(url, opts) returns a RemoteFile with the methods of H5File (kind, list, info, attrs, attrErrors, read, readHyperslab), each a promise, and stats(). It runs every call through the restartable LazyStorage: a pass that misses reports the byte ranges, js/remote.js fetches them with fetch() and Range headers (six at a time), and the pass is re-run. This keeps the main thread free without a Worker or synchronous XHR (h5wasm's lazy files need both), as the design doc recommends; the cost is re-running a pass per wave of misses. Every answer is checked: a 206 with exactly the bytes asked for, and the same ETag/Last-Modified and length as at open, else an error (never data). A server that ignores Range (200) is downloaded whole, up to maxDownload (512 MiB), unless fallback: "error". Options: blockSize, cacheSize, headers, credentials, parallel, fetch. test/serve.py is a range-capable static server with request counting (and /norange/ for a server without range support). test.mjs repeats every fixture check on files opened by URL (1 MiB and 512 B blocks), checks the request budget on a 200 MB h5py file (list, three small reads and a window of the big dataset: 5 requests, 6 MiB), the download fallback, and HTTP errors, changed files and wrong answers; with CLAWHDF5_WASM_CORPUS every corpus file is compared with open(bytes). Co-Authored-By: Claude Opus 5.5 (1M context) --- crates/clawhdf5-wasm/Cargo.toml | 2 + crates/clawhdf5-wasm/js/remote.js | 182 ++++++++ crates/clawhdf5-wasm/src/lib.rs | 498 +++++++++++++++++++--- examples/wasm-viewer/test/make_fixture.py | 37 ++ examples/wasm-viewer/test/run.sh | 31 +- examples/wasm-viewer/test/serve.py | 201 +++++++++ examples/wasm-viewer/test/test.mjs | 338 +++++++++++++-- examples/wasm-viewer/viewer-lib.js | 21 + 8 files changed, 1209 insertions(+), 101 deletions(-) create mode 100644 crates/clawhdf5-wasm/js/remote.js create mode 100644 examples/wasm-viewer/test/serve.py diff --git a/crates/clawhdf5-wasm/Cargo.toml b/crates/clawhdf5-wasm/Cargo.toml index 303837a..6d222bc 100644 --- a/crates/clawhdf5-wasm/Cargo.toml +++ b/crates/clawhdf5-wasm/Cargo.toml @@ -24,6 +24,8 @@ clawhdf5-format = { path = "../clawhdf5-format", version = "2.7.0" } # Must match the wasm-bindgen CLI exactly; build.sh checks. wasm-bindgen = "0.2.129" js-sys = "0.3.106" +# Promises for openUrl and RemoteFile (pure Rust over js-sys). +wasm-bindgen-futures = "0.4.79" [dev-dependencies] serde_json = "1" diff --git a/crates/clawhdf5-wasm/js/remote.js b/crates/clawhdf5-wasm/js/remote.js new file mode 100644 index 0000000..fdbe619 --- /dev/null +++ b/crates/clawhdf5-wasm/js/remote.js @@ -0,0 +1,182 @@ +// HTTP for clawhdf5-wasm's openUrl (see src/lib.rs and src/lazy.rs). +// +// The Rust side decides which byte ranges a read needs; this file fetches +// them with `fetch` and `Range` headers and checks every answer, so a server +// that ignores the range, answers with other bytes, or serves a file that +// changed since it was opened is an error, never data. wasm-bindgen copies +// it into the package (pkg/snippets/...). + +const DEFAULT_MAX_DOWNLOAD = 512 * 1024 * 1024; +const DEFAULT_PARALLEL = 6; + +function fetcher(opts) { + const f = opts?.fetch ?? globalThis.fetch; + if (typeof f !== "function") { + throw new Error("openUrl: no fetch() in this environment (pass opts.fetch)"); + } + return f; +} + +function init(opts, extra, method = "GET") { + return { method, headers: { ...(opts?.headers ?? {}), ...extra }, credentials: opts?.credentials }; +} + +// "bytes a-b/total" -> { start, end (exclusive), total | null }; null when +// the page cannot see the header (cross-origin, not exposed). +function contentRange(resp, url) { + const v = resp.headers.get("Content-Range"); + if (v === null) return null; + const m = /^bytes (\d+)-(\d+)\/(\d+|\*)$/.exec(v.trim()); + if (!m) throw new Error(`${url}: the server sent an unusable Content-Range: ${v}`); + return { start: Number(m[1]), end: Number(m[2]) + 1, total: m[3] === "*" ? null : Number(m[3]) }; +} + +// What pins the file: its ETag, else its Last-Modified (null if neither is +// visible to this page). +function validatorOf(resp) { + return resp.headers.get("ETag") ?? resp.headers.get("Last-Modified"); +} + +async function discard(resp) { + try { + await resp.body?.cancel(); + } catch { + // Nothing to release. + } +} + +// The whole body, refusing more than `limit` bytes as they arrive. +async function readAll(resp, limit, url) { + const tooBig = (n) => + new Error(`${url} is ${n} bytes, more than maxDownload (${limit}); ` + + "the server does not support range requests, so the whole file would have to be downloaded"); + const declared = resp.headers.get("Content-Length"); + if (declared !== null && Number(declared) > limit) { + await discard(resp); + throw tooBig(declared); + } + // A declared length was checked above: read the body at once. (Only an + // undeclared length is streamed, to stop at the limit; stream reads + // also stalled in the headless Chromium test under --virtual-time-budget.) + if (!resp.body || declared !== null) { + const all = new Uint8Array(await resp.arrayBuffer()); + if (all.length > limit) throw tooBig(all.length); + return all; + } + const reader = resp.body.getReader(); + const parts = []; + let n = 0; + for (;;) { + const { done, value } = await reader.read(); + if (done) break; + n += value.length; + if (n > limit) { + await reader.cancel(); + throw tooBig(`over ${limit}`); + } + parts.push(value); + } + const all = new Uint8Array(n); + let at = 0; + for (const p of parts) { + all.set(p, at); + at += p.length; + } + return all; +} + +/** + * Ask for the file's first `firstLen` bytes. A server that honours the + * range (206) gives `{ length, first, validator, requests }`; one that + * answers 200 sends the whole file, which is kept (`{ whole, requests }`) + * when `opts.fallback` is "download" (the default) and the file is at most + * `opts.maxDownload` bytes, and is an error otherwise. + */ +export async function probe(url, firstLen, opts) { + const f = fetcher(opts); + const resp = await f(url, init(opts, { Range: `bytes=0-${firstLen - 1}` })); + if (resp.status === 206) { + const cr = contentRange(resp, url); + if (cr && cr.start !== 0) { + await discard(resp); + throw new Error(`${url}: asked for bytes from 0, the server sent bytes from ${cr.start}`); + } + const first = new Uint8Array(await resp.arrayBuffer()); + let length = cr?.total ?? null; + let requests = 1; + if (length === null) { + // Content-Range is not readable here: a cross-origin server that does + // not list it in Access-Control-Expose-Headers. Content-Length of a + // HEAD request is always readable. + const head = await f(url, init(opts, {}, "HEAD")); + requests++; + const cl = head.headers.get("Content-Length"); + if (!head.ok || cl === null) { + throw new Error(`${url}: cannot learn the file's size (a cross-origin server must send ` + + "Access-Control-Expose-Headers: Content-Range, or answer HEAD with Content-Length)"); + } + length = Number(cl); + } + if (!Number.isSafeInteger(length) || first.length !== Math.min(firstLen, length)) { + throw new Error(`${url}: asked for the first ${firstLen} bytes of ${length}, got ${first.length}`); + } + return { length, first, validator: validatorOf(resp), requests }; + } + if (resp.status === 200) { + if ((opts?.fallback ?? "download") !== "download") { + await discard(resp); + throw new Error(`${url}: the server does not support HTTP range requests (it answered 200 ` + + "to a Range request); open it with { fallback: \"download\" } to download the whole file"); + } + const whole = await readAll(resp, opts?.maxDownload ?? DEFAULT_MAX_DOWNLOAD, url); + return { whole, requests: 1 }; + } + await discard(resp); + throw new Error(`${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim()); +} + +/** + * Fetch `ranges` ([start0, end0, start1, end1, ...], ends exclusive) of a + * file opened by `probe`, at most `opts.parallel` (default 6) at a time. + * Every answer must be a 206 with exactly the bytes asked for, from the same + * file (validator and length). + */ +export async function fetchRanges(url, ranges, opts, validator, length) { + const f = fetcher(opts); + const n = ranges.length / 2; + const out = new Array(n); + let next = 0; + async function worker() { + while (next < n) { + const i = next++; + const start = ranges[2 * i]; + const end = ranges[2 * i + 1]; + const resp = await f(url, init(opts, { Range: `bytes=${start}-${end - 1}` })); + if (resp.status !== 206) { + await discard(resp); + throw new Error(resp.status === 200 + ? `${url}: the server stopped honouring range requests` + : `${url}: HTTP ${resp.status} ${resp.statusText ?? ""}`.trim()); + } + const cr = contentRange(resp, url); + const v = validatorOf(resp); + if ((validator !== null && v !== null && v !== validator) || + (cr?.total != null && cr.total !== length)) { + await discard(resp); + throw new Error(`${url} changed on the server since it was opened`); + } + if (cr && (cr.start !== start || cr.end !== end)) { + await discard(resp); + throw new Error(`${url}: asked for bytes ${start}-${end - 1}, the server sent ${cr.start}-${cr.end - 1}`); + } + const body = new Uint8Array(await resp.arrayBuffer()); + if (body.length !== end - start) { + throw new Error(`${url}: asked for ${end - start} bytes at offset ${start}, got ${body.length}`); + } + out[i] = body; + } + } + const workers = Math.max(1, Math.min(opts?.parallel ?? DEFAULT_PARALLEL, n)); + await Promise.all(Array.from({ length: workers }, worker)); + return out; +} diff --git a/crates/clawhdf5-wasm/src/lib.rs b/crates/clawhdf5-wasm/src/lib.rs index 2fed048..a421573 100644 --- a/crates/clawhdf5-wasm/src/lib.rs +++ b/crates/clawhdf5-wasm/src/lib.rs @@ -1,7 +1,7 @@ //! clawhdf5's HDF5 reader for JavaScript, via `wasm-bindgen`. //! //! ```js -//! import init, { open } from "./pkg/clawhdf5_wasm.js"; +//! import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js"; //! await init(); //! const file = open(new Uint8Array(await blob.arrayBuffer())); //! file.list("/"); // [{ name, kind: "group" | "dataset" }] @@ -10,6 +10,12 @@ //! file.read("/x"); // { shape, dtype, data: Float64Array | ... | string[] } //! file.readHyperslab("/x", [0, 0], [10, 10]); // stride, block optional //! file.free(); +//! +//! // A file on a web server, read by HTTP range requests as needed: the +//! // same methods, returning promises. +//! const remote = await openUrl("https://example.org/data.h5"); +//! await remote.list("/"); +//! remote.stats(); // { requests, bytesFetched, size, ... } //! ``` //! //! Numeric data comes back in the typed array of the stored width @@ -18,16 +24,23 @@ //! Anything else is a thrown `Error` naming the datatype. Only the reader is //! exposed: nothing here writes files. //! -//! The logic lives in [`core`], which is plain Rust and tested natively. +//! The logic lives in [`core`] and [`lazy`], which are plain Rust and tested +//! natively; `js/remote.js` does the HTTP. pub mod core; pub mod lazy; -use clawhdf5::AttrValue; -use js_sys::{Array, Object, Reflect}; -use wasm_bindgen::prelude::*; +use std::ops::Range; +use std::rc::Rc; +use std::sync::Arc; -use crate::core::{Data, Hyperslab, Reader}; +use clawhdf5::AttrValue; +use js_sys::{Array, Object, Promise, Reflect, Uint8Array}; +use wasm_bindgen::prelude::*; +use wasm_bindgen_futures::future_to_promise; + +use crate::core::{Attr, Child, Data, DatasetInfo, Hyperslab, Reader}; +use crate::lazy::{LazyConfig, LazyStorage, Step}; /// JavaScript numbers are exact up to 2^53. const MAX_SAFE_INTEGER: f64 = 9_007_199_254_740_991.0; @@ -36,11 +49,30 @@ fn js_err(msg: String) -> JsError { JsError::new(&msg) } +/// A JavaScript exception (from `fetch`, or `js/remote.js`) as a `JsError` +/// with its message. +fn js_exception(e: JsValue) -> JsError { + let msg = e + .dyn_ref::() + .map(|e| String::from(e.message())) + .or_else(|| e.as_string()) + .unwrap_or_else(|| format!("{e:?}")); + JsError::new(&msg) +} + fn set(obj: &Object, key: &str, value: impl Into) { // Defining a property on a fresh plain object cannot fail. Reflect::set(obj, &JsValue::from_str(key), &value.into()).unwrap_throw(); } +fn get(obj: &JsValue, key: &str) -> JsValue { + if obj.is_object() { + Reflect::get(obj, &JsValue::from_str(key)).unwrap_or(JsValue::UNDEFINED) + } else { + JsValue::UNDEFINED + } +} + fn shape_to_js(shape: &[u64]) -> Array { shape.iter().map(|&d| JsValue::from_f64(d as f64)).collect() } @@ -59,6 +91,20 @@ fn indices_from_js(what: &str, v: &[f64]) -> Result, JsError> { .collect() } +fn slab_from_js( + start: &[f64], + count: &[f64], + stride: Option>, + block: Option>, +) -> Result { + Ok(Hyperslab { + start: indices_from_js("start", start)?, + count: indices_from_js("count", count)?, + stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?, + block: block.map(|b| indices_from_js("block", &b)).transpose()?, + }) +} + fn data_to_js(data: Data) -> JsValue { match data { Data::F32(v) => js_sys::Float32Array::from(&v[..]).into(), @@ -111,6 +157,72 @@ fn attr_to_js(value: AttrValue) -> (JsValue, Option) { } } +fn list_to_js(children: Vec) -> Array { + children + .into_iter() + .map(|c| { + let o = Object::new(); + set(&o, "name", c.name); + set(&o, "kind", c.kind.as_str()); + JsValue::from(o) + }) + .collect() +} + +fn info_to_js(i: DatasetInfo) -> Object { + let o = Object::new(); + set(&o, "shape", shape_to_js(&i.shape)); + let max: JsValue = match i.maxshape { + None => JsValue::NULL, + Some(dims) => dims + .into_iter() + .map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64))) + .collect::() + .into(), + }; + set(&o, "maxshape", max); + set(&o, "dtype", i.dtype); + set(&o, "elementShape", shape_to_js(&i.element_shape)); + o +} + +fn attrs_to_js(attrs: Vec) -> Array { + attrs + .into_iter() + .map(|a| { + let o = Object::new(); + set(&o, "name", a.name); + let (value, dtype) = attr_to_js(a.value); + set(&o, "value", value); + set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from)); + JsValue::from(o) + }) + .collect() +} + +fn errors_to_js(errors: Vec) -> Array { + errors.into_iter().map(JsValue::from).collect() +} + +/// A read's values with the dataset's datatype. +fn values_to_js((dtype, a): (String, core::Array)) -> Object { + let o = Object::new(); + set(&o, "shape", shape_to_js(&a.shape)); + set(&o, "dtype", dtype); + set(&o, "data", data_to_js(a.data)); + o +} + +/// Read the dataset at `path` (whole, or `slab`) with its datatype. +fn read_values( + r: &Reader, + path: &str, + slab: Option<&Hyperslab>, +) -> core::Result<(String, core::Array)> { + let dtype = r.info(path)?.dtype; + Ok((dtype, r.read(path, slab)?)) +} + /// An open HDF5 (or NetCDF-4) file. #[wasm_bindgen] pub struct H5File { @@ -146,38 +258,13 @@ impl H5File { /// The group's members: `[{ name, kind }]`, groups first. pub fn list(&self, path: &str) -> Result { - Ok(self - .inner - .list(path) - .map_err(js_err)? - .into_iter() - .map(|c| { - let o = Object::new(); - set(&o, "name", c.name); - set(&o, "kind", c.kind.as_str()); - JsValue::from(o) - }) - .collect()) + Ok(list_to_js(self.inner.list(path).map_err(js_err)?)) } /// `{ shape, maxshape, dtype, elementShape }`. `maxshape` is `null` /// when not recorded, with `null` for each unlimited dimension. pub fn info(&self, path: &str) -> Result { - let i = self.inner.info(path).map_err(js_err)?; - let o = Object::new(); - set(&o, "shape", shape_to_js(&i.shape)); - let max: JsValue = match i.maxshape { - None => JsValue::NULL, - Some(dims) => dims - .into_iter() - .map(|d| d.map_or(JsValue::NULL, |d| JsValue::from_f64(d as f64))) - .collect::() - .into(), - }; - set(&o, "maxshape", max); - set(&o, "dtype", i.dtype); - set(&o, "elementShape", shape_to_js(&i.element_shape)); - Ok(o) + Ok(info_to_js(self.inner.info(path).map_err(js_err)?)) } /// `[{ name, value, dtype }]`, sorted by name. Scalars are `number` @@ -186,31 +273,21 @@ impl H5File { /// and its `dtype`; one that could not be read at all is reported by /// [`attrErrors`](Self::attr_errors). pub fn attrs(&self, path: &str) -> Result { - let (attrs, _) = self.inner.attrs(path).map_err(js_err)?; - Ok(attrs - .into_iter() - .map(|a| { - let o = Object::new(); - set(&o, "name", a.name); - let (value, dtype) = attr_to_js(a.value); - set(&o, "value", value); - set(&o, "dtype", dtype.map_or(JsValue::NULL, JsValue::from)); - JsValue::from(o) - }) - .collect()) + Ok(attrs_to_js(self.inner.attrs(path).map_err(js_err)?.0)) } /// Messages for attributes that could not be read. #[wasm_bindgen(js_name = attrErrors)] pub fn attr_errors(&self, path: &str) -> Result { - let (_, errors) = self.inner.attrs(path).map_err(js_err)?; - Ok(errors.into_iter().map(JsValue::from).collect()) + Ok(errors_to_js(self.inner.attrs(path).map_err(js_err)?.1)) } /// The whole dataset: `{ shape, dtype, data }`, `data` in row-major /// order. pub fn read(&self, path: &str) -> Result { - self.read_impl(path, None) + Ok(values_to_js( + read_values(&self.inner, path, None).map_err(js_err)?, + )) } /// A regular hyperslab (`H5Sselect_hyperslab`): `stride` and `block` @@ -224,22 +301,315 @@ impl H5File { stride: Option>, block: Option>, ) -> Result { - let slab = Hyperslab { - start: indices_from_js("start", &start)?, - count: indices_from_js("count", &count)?, - stride: stride.map(|s| indices_from_js("stride", &s)).transpose()?, - block: block.map(|b| indices_from_js("block", &b)).transpose()?, - }; - self.read_impl(path, Some(&slab)) - } - - fn read_impl(&self, path: &str, slab: Option<&Hyperslab>) -> Result { - let dtype = self.inner.info(path).map_err(js_err)?.dtype; - let a = self.inner.read(path, slab).map_err(js_err)?; + let slab = slab_from_js(&start, &count, stride, block)?; + Ok(values_to_js( + read_values(&self.inner, path, Some(&slab)).map_err(js_err)?, + )) + } +} + +// --------------------------------------------------------------------------- +// Remote files: openUrl. + +#[wasm_bindgen(module = "/js/remote.js")] +extern "C" { + #[wasm_bindgen(catch)] + async fn probe(url: &str, first_len: f64, opts: &JsValue) -> Result; + + #[wasm_bindgen(catch, js_name = fetchRanges)] + async fn fetch_ranges( + url: &str, + ranges: Vec, + opts: &JsValue, + validator: &JsValue, + length: f64, + ) -> Result; +} + +/// Where a remote file's bytes come from. +struct Http { + url: String, + opts: JsValue, + /// ETag or Last-Modified at open (`null` if the server sent neither). + validator: JsValue, + length: u64, + /// Requests the probe made (1, or 2 with a HEAD for the length). + probe_requests: u64, +} + +enum Source { + /// Read by range requests through a restartable cache. + Lazy { + http: Http, + storage: Arc, + reader: Reader, + }, + /// The server ignored `Range`: the whole file, downloaded at open. + Whole { + reader: Reader, + size: u64, + requests: u64, + }, +} + +impl Http { + /// Fetch `ranges` and hand them to `storage`. + async fn fetch(&self, storage: &LazyStorage, ranges: &[Range]) -> Result<(), JsError> { + let flat: Vec = ranges + .iter() + .flat_map(|r| [r.start as f64, r.end as f64]) + .collect(); + let got = fetch_ranges( + &self.url, + flat, + &self.opts, + &self.validator, + self.length as f64, + ) + .await + .map_err(js_exception)?; + let got = Array::from(&got); + if got.length() as usize != ranges.len() { + return Err(js_err(format!( + "fetchRanges returned {} ranges for {}", + got.length(), + ranges.len() + ))); + } + for (r, bytes) in ranges.iter().zip(got.iter()) { + let bytes = Uint8Array::new(&bytes).to_vec(); + storage.supply_range(r, &bytes).map_err(js_err)?; + } + Ok(()) + } + + /// Run `f` over `storage` until it has every byte it reads. + async fn drive( + &self, + storage: &LazyStorage, + mut f: impl FnMut() -> T, + ) -> Result { + let _op = storage.operation(); + loop { + match storage.attempt(&mut f) { + Step::Done(v) => return Ok(v), + Step::Need(ranges) => self.fetch(storage, &ranges).await?, + } + } + } +} + +impl Source { + async fn run(&self, op: impl Fn(&Reader) -> core::Result) -> Result { + match self { + Source::Whole { reader, .. } => op(reader).map_err(js_err), + Source::Lazy { + http, + storage, + reader, + } => http.drive(storage, || op(reader)).await?.map_err(js_err), + } + } +} + +/// A non-negative integer option, or `None` when not given. +fn int_opt(opts: &JsValue, key: &str, min: f64, max: f64) -> Result, JsError> { + let v = get(opts, key); + if v.is_undefined() || v.is_null() { + return Ok(None); + } + match v.as_f64() { + Some(x) if x.fract() == 0.0 && (min..=max).contains(&x) => Ok(Some(x as u64)), + _ => Err(js_err(format!( + "openUrl: {key} must be an integer from {min} to {max}" + ))), + } +} + +fn config_from(opts: &JsValue) -> Result { + let mut c = LazyConfig::default(); + if let Some(b) = int_opt(opts, "blockSize", 512.0, (64u64 << 20) as f64)? { + c.block_size = b; + } + if let Some(n) = int_opt(opts, "cacheSize", 0.0, MAX_SAFE_INTEGER)? { + c.capacity = n; + } + Ok(c) +} + +/// Open the HDF5 file at `url` without downloading it: its bytes are +/// fetched with HTTP `Range` requests as the methods of the returned +/// [`RemoteFile`] need them, through a block cache. +/// +/// `opts` (all optional): +/// - `blockSize` — bytes per request block, 512 to 64 MiB (default 1 MiB); +/// - `cacheSize` — bytes of blocks kept between calls (default 64 MiB); +/// - `fallback` — `"download"` (default) reads the whole file when the +/// server ignores `Range` (answers 200), up to `maxDownload` bytes +/// (default 512 MiB); `"error"` refuses such a server; +/// - `headers`, `credentials` — passed to every `fetch`; +/// - `parallel` — range requests in flight at once (default 6); +/// - `fetch` — a `fetch`-compatible function to use instead of the global. +/// +/// Cross-origin servers must allow CORS and expose `Content-Range` (or +/// answer `HEAD` with `Content-Length`). +#[wasm_bindgen(js_name = openUrl)] +pub async fn open_url(url: String, opts: JsValue) -> Result { + let config = config_from(&opts)?; + let p = probe(&url, config.block_size as f64, &opts) + .await + .map_err(js_exception)?; + let requests = get(&p, "requests").as_f64().unwrap_or(1.0) as u64; + let whole = get(&p, "whole"); + if !whole.is_undefined() { + let bytes = Uint8Array::new(&whole).to_vec(); + let size = bytes.len() as u64; + let reader = Reader::open(bytes).map_err(js_err)?; + return Ok(RemoteFile { + inner: Rc::new(Source::Whole { + reader, + size, + requests, + }), + }); + } + let length = get(&p, "length") + .as_f64() + .filter(|x| x.fract() == 0.0 && (0.0..=MAX_SAFE_INTEGER).contains(x)) + .ok_or_else(|| js_err(format!("{url}: the server gave no usable file size")))?; + let http = Http { + url, + opts, + validator: get(&p, "validator"), + length: length as u64, + probe_requests: requests, + }; + let storage = Arc::new(LazyStorage::new(http.length, config)); + let first = Uint8Array::new(&get(&p, "first")).to_vec(); + storage.supply(0, &first).map_err(js_err)?; + let s = storage.clone(); + let reader = http + .drive(&storage, || Reader::open_storage(s.clone())) + .await? + .map_err(js_err)?; + Ok(RemoteFile { + inner: Rc::new(Source::Lazy { + http, + storage, + reader, + }), + }) +} + +/// A file opened with [`openUrl`](open_url): the methods of [`H5File`], +/// each returning a `Promise` (it may have to fetch bytes first). +#[wasm_bindgen] +pub struct RemoteFile { + inner: Rc, +} + +impl RemoteFile { + /// Run `op` (fetching what it needs) and convert its result. + fn call( + &self, + op: impl Fn(&Reader) -> core::Result + 'static, + to_js: impl FnOnce(T) -> JsValue + 'static, + ) -> Promise { + let inner = self.inner.clone(); + future_to_promise(async move { + let v = inner.run(op).await.map_err(JsValue::from)?; + Ok(to_js(v)) + }) + } +} + +#[wasm_bindgen] +impl RemoteFile { + /// `"group"` or `"dataset"`. + #[wasm_bindgen(unchecked_return_type = "Promise")] + pub fn kind(&self, path: String) -> Promise { + self.call(move |r| r.kind(&path), |k| k.as_str().into()) + } + + /// The group's members: `[{ name, kind }]`, groups first. + #[wasm_bindgen(unchecked_return_type = "Promise>")] + pub fn list(&self, path: String) -> Promise { + self.call(move |r| r.list(&path), |c| list_to_js(c).into()) + } + + /// `{ shape, maxshape, dtype, elementShape }`, as [`H5File::info`]. + #[wasm_bindgen(unchecked_return_type = "Promise")] + pub fn info(&self, path: String) -> Promise { + self.call(move |r| r.info(&path), |i| info_to_js(i).into()) + } + + /// `[{ name, value, dtype }]`, as [`H5File::attrs`]. + #[wasm_bindgen(unchecked_return_type = "Promise>")] + pub fn attrs(&self, path: String) -> Promise { + self.call(move |r| r.attrs(&path), |a| attrs_to_js(a.0).into()) + } + + /// Messages for attributes that could not be read. + #[wasm_bindgen(js_name = attrErrors, unchecked_return_type = "Promise>")] + pub fn attr_errors(&self, path: String) -> Promise { + self.call(move |r| r.attrs(&path), |a| errors_to_js(a.1).into()) + } + + /// The whole dataset: `{ shape, dtype, data }`, as [`H5File::read`]. + #[wasm_bindgen(unchecked_return_type = "Promise")] + pub fn read(&self, path: String) -> Promise { + self.call( + move |r| read_values(r, &path, None), + |v| values_to_js(v).into(), + ) + } + + /// A regular hyperslab, as [`H5File::read_hyperslab`]. Only the chunks + /// (or the contiguous runs) the selection touches are fetched. + #[wasm_bindgen(js_name = readHyperslab, unchecked_return_type = "Promise")] + pub fn read_hyperslab( + &self, + path: String, + start: Vec, + count: Vec, + stride: Option>, + block: Option>, + ) -> Result { + let slab = slab_from_js(&start, &count, stride, block)?; + Ok(self.call( + move |r| read_values(r, &path, Some(&slab)), + |v| values_to_js(v).into(), + )) + } + + /// What reading this file has cost so far: `{ lazy, size, requests, + /// bytesFetched, cachedBytes, passes }`. `lazy` is false when the + /// server ignored `Range` and the file was downloaded whole. + pub fn stats(&self) -> Object { let o = Object::new(); - set(&o, "shape", shape_to_js(&a.shape)); - set(&o, "dtype", dtype); - set(&o, "data", data_to_js(a.data)); - Ok(o) + match &*self.inner { + Source::Lazy { http, storage, .. } => { + let st = storage.stats(); + set(&o, "lazy", true); + set(&o, "size", http.length as f64); + set( + &o, + "requests", + (st.requests.saturating_sub(1) + http.probe_requests) as f64, + ); + set(&o, "bytesFetched", st.bytes_fetched as f64); + set(&o, "cachedBytes", st.cached_bytes as f64); + set(&o, "passes", st.passes as f64); + } + Source::Whole { size, requests, .. } => { + set(&o, "lazy", false); + set(&o, "size", *size as f64); + set(&o, "requests", *requests as f64); + set(&o, "bytesFetched", *size as f64); + set(&o, "cachedBytes", *size as f64); + set(&o, "passes", 0.0); + } + } + o } } diff --git a/examples/wasm-viewer/test/make_fixture.py b/examples/wasm-viewer/test/make_fixture.py index 8c7722c..1e79065 100644 --- a/examples/wasm-viewer/test/make_fixture.py +++ b/examples/wasm-viewer/test/make_fixture.py @@ -14,6 +14,7 @@ encoded as strings so JSON.parse keeps 64-bit values exact. """ import json +import os import sys import warnings from pathlib import Path @@ -201,3 +202,39 @@ def slab_for(obj): json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)}, open(out / "expected.json", "w"), indent=1, ensure_ascii=False) + + +def write_big(path, megabytes): + """A large file for the range-request tests (`openUrl`): `/big`, about + `megabytes` MB of float64 in 1 MiB chunks, written after a small + dataset and a group, so listing and reading `/small` touch a few blocks + of the file and a window of `/big` one chunk. Returns what h5py reads + back.""" + n = megabytes * 1_000_000 // 8 + chunk = 1 << 17 + with h5py.File(path, "w") as f: + f.attrs["note"] = "large file for range reads" + f.create_dataset("small", data=np.array([1.5, -2.0, 3.25])) + g = f.create_group("meta") + g.attrs["units"] = "m" + g.create_dataset("ids", data=np.arange(10, dtype=" 0: + json.dump(write_big(out / "big.h5", big_mb), open(out / "big.json", "w"), indent=1) diff --git a/examples/wasm-viewer/test/run.sh b/examples/wasm-viewer/test/run.sh index c1cdae0..68c4677 100755 --- a/examples/wasm-viewer/test/run.sh +++ b/examples/wasm-viewer/test/run.sh @@ -1,11 +1,17 @@ #!/usr/bin/env bash # Build the wasm package (../build.sh), test it under Node against files -# written by h5py and netCDF4 (make_fixture.py), then load the viewer page -# in headless Chromium if one is found (browser.sh). +# written by h5py and netCDF4 (make_fixture.py) — opened from bytes, and +# opened by URL from a local range-capable HTTP server (serve.py) — then +# load the viewer page in headless Chromium if one is found (browser.sh). # # Needs node, the wasm-bindgen CLI (see ../build.sh) and a Python with h5py, # netCDF4 and numpy: CLAWHDF5_PYTHON names it (default python3). Without that # Python the test is skipped, unless CLAWHDF5_REQUIRE_INTEROP=1. +# +# WASM_BIG_MB (default 200) sizes the large file of the range-request +# budget test (0 leaves it out); it is written under TMPDIR. +# CLAWHDF5_WASM_CORPUS=DIR also compares every HDF5 file under DIR (up to +# 16 MiB) read over HTTP with the same file read from bytes. set -euo pipefail HERE="$(cd "$(dirname "$0")" && pwd)" @@ -23,9 +29,24 @@ fi bash "$HERE/../build.sh" fix="$(mktemp -d)" -trap 'rm -rf "$fix"' EXIT -"$PY" "$HERE/make_fixture.py" "$fix" -node "$HERE/test.mjs" "$HERE/../pkg" "$fix" +server="" +cleanup() { + [ -n "$server" ] && kill "$server" 2>/dev/null || true + rm -rf "$fix" +} +trap cleanup EXIT +WASM_BIG_MB="${WASM_BIG_MB:-200}" "$PY" "$HERE/make_fixture.py" "$fix" + +# The fixtures (and the corpus) over HTTP with range support. +roots=(--root "fix=$fix") +[ -n "${CLAWHDF5_WASM_CORPUS:-}" ] && roots+=(--root "corpus=$CLAWHDF5_WASM_CORPUS") +"$PY" "$HERE/serve.py" "${roots[@]}" > "$fix/port" & +server=$! +for _ in $(seq 50); do + [ -s "$fix/port" ] && break + sleep 0.1 +done +node "$HERE/test.mjs" "$HERE/../pkg" "$fix" "http://127.0.0.1:$(head -1 "$fix/port")" # The page itself, in headless Chromium when one is available. status=0 diff --git a/examples/wasm-viewer/test/serve.py b/examples/wasm-viewer/test/serve.py new file mode 100644 index 0000000..f4a6e09 --- /dev/null +++ b/examples/wasm-viewer/test/serve.py @@ -0,0 +1,201 @@ +"""A static HTTP server for the wasm tests, with HTTP Range support and +request counting. + + python serve.py [--root PREFIX=DIR ...] + +Serves each DIR under URL PREFIX (the first match wins; PREFIX "" is the +site root), prints the port on its first line of stdout, and runs until +killed. No symlinks or copies are made: files are read where they are. + +- `Range: bytes=a-b`, `bytes=a-` and `bytes=-n` get 206 with Content-Range, + an unsatisfiable range 416; every file answer carries an ETag, and CORS + headers exposing Content-Range, so a page on another origin can use it. +- Under `/norange/...` the same files are served but Range is ignored + (200 with the whole file), as by a server without range support. +- `GET /__stats` returns `{"requests": n, "bytes": n, "log": [...]}` for + file requests since the last `GET /__reset`, which zeroes them. +""" + +import argparse +import hashlib +import json +import os +import posixpath +import sys +import threading +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +from urllib.parse import unquote, urlsplit + +TYPES = { + ".html": "text/html; charset=utf-8", + ".js": "text/javascript; charset=utf-8", + ".mjs": "text/javascript; charset=utf-8", + ".wasm": "application/wasm", + ".json": "application/json", + ".ts": "text/plain; charset=utf-8", +} + +lock = threading.Lock() +stats = {"requests": 0, "bytes": 0, "log": []} + + +def resolve(roots, path): + """The file for URL `path`, or None. `..` never leaves a root.""" + parts = [p for p in posixpath.normpath(unquote(path)).split("/") if p] + if any(p in (".", "..") for p in parts): + return None + for prefix, root in roots: + pre = [p for p in prefix.split("/") if p] + if parts[: len(pre)] == pre: + rest = parts[len(pre):] or ["index.html"] + f = os.path.join(root, *rest) + if os.path.isfile(f): + return f + return None + + +def parse_range(header, size): + """(start, end exclusive) for a single `bytes=` range, "bad" when + unsatisfiable, None when absent or unparsable (served whole).""" + if not header or not header.startswith("bytes=") or "," in header: + return None + a, _, b = header[len("bytes="):].strip().partition("-") + try: + if a == "": + n = int(b) + return (max(0, size - n), size) if n > 0 and size > 0 else "bad" + start = int(a) + end = int(b) + 1 if b else size + except ValueError: + return None + if start >= size or end <= start: + return "bad" + return start, min(end, size) + + +def make_handler(roots): + class Handler(BaseHTTPRequestHandler): + protocol_version = "HTTP/1.1" + + def log_message(self, *args): + pass + + def cors(self): + self.send_header("Access-Control-Allow-Origin", "*") + self.send_header("Access-Control-Expose-Headers", + "Content-Range, Content-Length, ETag, Accept-Ranges") + + def do_OPTIONS(self): + self.send_response(204) + self.cors() + self.send_header("Access-Control-Allow-Headers", "Range") + self.send_header("Content-Length", "0") + self.end_headers() + + def do_HEAD(self): + self.serve(head=True) + + def do_GET(self): + self.serve(head=False) + + def json(self, obj): + body = json.dumps(obj).encode() + self.send_response(200) + self.cors() + self.send_header("Content-Type", "application/json") + self.send_header("Content-Length", str(len(body))) + self.send_header("Cache-Control", "no-store") + self.end_headers() + self.wfile.write(body) + + def serve(self, head): + path = urlsplit(self.path).path + if path == "/__stats": + with lock: + return self.json(stats) + if path == "/__reset": + with lock: + stats.update(requests=0, bytes=0, log=[]) + return self.json({}) + ranges = True + if path.startswith("/norange/"): + ranges = False + path = path[len("/norange"):] + f = resolve(roots, path) + if f is None: + self.send_response(404) + self.cors() + self.send_header("Content-Length", "0") + self.end_headers() + return + size = os.path.getsize(f) + st = os.stat(f) + etag = '"%s"' % hashlib.sha1( + f"{f}:{size}:{st.st_mtime_ns}".encode()).hexdigest()[:16] + r = parse_range(self.headers.get("Range"), size) if ranges else None + if r == "bad": + self.send_response(416) + self.cors() + self.send_header("Content-Range", f"bytes */{size}") + self.send_header("Content-Length", "0") + self.end_headers() + return + start, end = r if r else (0, size) + self.send_response(206 if r else 200) + self.cors() + ext = os.path.splitext(f)[1] + self.send_header("Content-Type", TYPES.get(ext, "application/octet-stream")) + self.send_header("Content-Length", str(end - start)) + self.send_header("ETag", etag) + self.send_header("Cache-Control", "no-store") + if ranges: + self.send_header("Accept-Ranges", "bytes") + if r: + self.send_header("Content-Range", f"bytes {start}-{end - 1}/{size}") + self.end_headers() + if not head: + with lock: + stats["requests"] += 1 + stats["bytes"] += end - start + stats["log"].append([path, start, end, 206 if r else 200]) + with open(f, "rb") as fh: + fh.seek(start) + left = end - start + try: + while left: + buf = fh.read(min(left, 1 << 20)) + if not buf: + break + self.wfile.write(buf) + left -= len(buf) + except (BrokenPipeError, ConnectionResetError): + pass + + return Handler + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--root", action="append", default=[], + help="PREFIX=DIR: serve DIR under URL PREFIX") + args = ap.parse_args() + roots = [] + for spec in args.root: + prefix, _, d = spec.partition("=") + roots.append((prefix, os.path.abspath(d))) + class Server(ThreadingHTTPServer): + def handle_error(self, request, client_address): + # A client that drops a connection (a cancelled download) is + # not an error of the server. + if not isinstance(sys.exc_info()[1], (ConnectionError, TimeoutError)): + super().handle_error(request, client_address) + + httpd = Server(("127.0.0.1", 0), make_handler(roots)) + httpd.daemon_threads = True + print(httpd.server_address[1], flush=True) + sys.stdout.close() + httpd.serve_forever() + + +if __name__ == "__main__": + main() diff --git a/examples/wasm-viewer/test/test.mjs b/examples/wasm-viewer/test/test.mjs index 80ca30a..7a264dc 100644 --- a/examples/wasm-viewer/test/test.mjs +++ b/examples/wasm-viewer/test/test.mjs @@ -1,20 +1,31 @@ // Node test of the built wasm package (the exact pkg/ the viewer page loads) // and the viewer's DOM-free helpers. Run by test/run.sh: -// node test.mjs PKG_DIR FIXTURE_DIR +// node test.mjs PKG_DIR FIXTURE_DIR [SERVER_URL] // FIXTURE_DIR holds fixture.h5, fixture.nc and expected.json from -// make_fixture.py (values as libhdf5 reads them back). +// make_fixture.py (values as libhdf5 reads them back), and big.h5/big.json +// when it was run with WASM_BIG_MB. SERVER_URL is test/serve.py serving +// FIXTURE_DIR under /fix (and $CLAWHDF5_WASM_CORPUS under /corpus): with +// it, every check is repeated on files opened with openUrl (HTTP range +// requests), and the request budget, the full-download fallback and the +// error paths of openUrl are tested. import assert from "node:assert/strict"; -import { readFileSync } from "node:fs"; -import { join } from "node:path"; +import { existsSync, readFileSync, readdirSync, statSync } from "node:fs"; +import { join, relative } from "node:path"; import { pathToFileURL } from "node:url"; -const [pkgDir, fixDir] = process.argv.slice(2); +const [pkgDir, fixDir, base] = process.argv.slice(2); const pkg = await import(pathToFileURL(join(pkgDir, "clawhdf5_wasm.js"))); pkg.initSync({ module: readFileSync(join(pkgDir, "clawhdf5_wasm_bg.wasm")) }); const lib = await import(pathToFileURL(join(import.meta.dirname, "..", "viewer-lib.js"))); let checks = 0; const eq = (a, b, msg) => { assert.deepEqual(a, b, msg); checks++; }; +// A call that must fail: a thrown Error (open) or a rejected promise +// (openUrl), whose message matches `re`. +const fails = async (fn, re, msg) => { + await assert.rejects(async () => fn(), (e) => e instanceof Error && re.test(e.message), msg); + checks++; +}; const ARRAY_TYPES = { f32: Float32Array, f64: Float64Array, i8: Int8Array, i16: Int16Array, i32: Int32Array, @@ -54,13 +65,13 @@ function checkAttr(ctx, a, want) { assert.fail(`${ctx}: unknown expectation ${JSON.stringify(want)}`); } -const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8")); -for (const [name, exp] of Object.entries(expected)) { - const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name)))); - +// Every listing, dataset, error and attribute of `file` against what +// libhdf5 reads (expected.json). `file` is an H5File (synchronous methods) +// or a RemoteFile (promises): every call is awaited. +async function checkFile(name, exp, file) { for (const [path, want] of Object.entries(exp.lists)) { - eq(file.kind(path), "group", `${name}:${path} kind`); - const list = file.list(path); + eq(await file.kind(path), "group", `${name}:${path} kind`); + const list = await file.list(path); for (const [kind, key] of [["group", "groups"], ["dataset", "datasets"]]) { eq(list.filter((c) => c.kind === kind).map((c) => c.name).sort(), want[key], `${name}:${path} ${key}`); } @@ -68,58 +79,67 @@ for (const [name, exp] of Object.entries(expected)) { for (const [path, want] of Object.entries(exp.datasets)) { const ctx = `${name}:${path}`; - eq(file.kind(path), "dataset", `${ctx} kind`); + eq(await file.kind(path), "dataset", `${ctx} kind`); if (want.unavailable) { // The wasm build has no zstd (it links C): a clear error, no data. - assert.throws(() => file.read(path), (e) => e.message.includes(want.unavailable), ctx); - checks++; + await fails(() => file.read(path), new RegExp(want.unavailable), ctx); continue; } - const info = file.info(path); + const info = await file.info(path); eq([...info.shape, ...info.elementShape], want.shape, `${ctx} info shape`); - const r = file.read(path); + const r = await file.read(path); eq(r.shape, want.shape, `${ctx} shape`); eq(r.dtype, info.dtype, `${ctx} dtype`); assert.ok(r.data instanceof ARRAY_TYPES[want.kind], `${ctx}: ${r.data.constructor.name} for ${want.kind}`); eq(values(want.kind, r.data), want.values, ctx); if (want.slab) { const s = want.slab; - const part = file.readHyperslab(path, s.start, s.count, s.stride); + const part = await file.readHyperslab(path, s.start, s.count, s.stride); eq(part.shape, s.shape, `${ctx} slab shape`); eq(values(want.kind, part.data), s.values, `${ctx} slab`); } } for (const [path, what] of Object.entries(exp.errors)) { - assert.throws(() => file.read(path), (e) => e instanceof Error && e.message.includes(what), `${name}:${path}`); - checks++; + await fails(() => file.read(path), new RegExp(what), `${name}:${path}`); } for (const [path, want] of Object.entries(exp.attrs)) { - const attrs = file.attrs(path); - eq(file.attrErrors(path), [], `${name}:${path} attr errors`); + const attrs = await file.attrs(path); + eq(await file.attrErrors(path), [], `${name}:${path} attr errors`); const seen = attrs.filter((a) => !a.name.startsWith("_") && !exp.skip_attrs.includes(a.name)); eq(seen.map((a) => a.name).sort(), Object.keys(want).sort(), `${name}:${path} attr names`); for (const a of seen) checkAttr(`${name}:${path}@${a.name}`, a, want[a.name]); } +} + +// The error paths of the reader, for an H5File or a RemoteFile of +// fixture.h5. +async function checkErrors(h5) { + await fails(() => h5.read("/nope"), /./); + await fails(() => h5.list("/grid"), /not a group/); + await fails(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/); + await fails(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/); + await fails(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/); + await fails(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/); + // Big integers stay exact. + eq((await h5.read("/u64")).data[0], 18446744073709551615n, "u64 max"); +} + +const expected = JSON.parse(readFileSync(join(fixDir, "expected.json"), "utf8")); +for (const [name, exp] of Object.entries(expected)) { + const file = pkg.open(new Uint8Array(readFileSync(join(fixDir, name)))); + await checkFile(name, exp, file); file.free(); } // Errors reach JavaScript as thrown Errors, never as data. const h5 = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.h5")))); -const throwsMsg = (fn, re) => { assert.throws(fn, (e) => e instanceof Error && re.test(e.message)); checks++; }; -throwsMsg(() => pkg.open(new Uint8Array(64)), /./); -throwsMsg(() => h5.read("/nope"), /./); -throwsMsg(() => h5.list("/grid"), /not a group/); -throwsMsg(() => h5.readHyperslab("/grid", [0], [1]), /dimensions/); -throwsMsg(() => h5.readHyperslab("/grid", [5, 0], [2, 1]), /exceeds/); -throwsMsg(() => h5.readHyperslab("/grid", [-1, 0], [1, 1]), /non-negative integers/); -throwsMsg(() => h5.readHyperslab("/grid", [0.5, 0], [1, 1]), /non-negative integers/); +await fails(() => pkg.open(new Uint8Array(64)), /./); +await checkErrors(h5); // Info for a dataset with an unlimited dimension (netCDF "time"). const nc = pkg.open(new Uint8Array(readFileSync(join(fixDir, "fixture.nc")))); eq(nc.info("/time").maxshape, [null], "unlimited dimension is null"); -// Big integers stay exact. -eq(h5.read("/u64").data[0], 18446744073709551615n, "u64 max"); eq(typeof pkg.version(), "string", "version"); // Viewer helpers. @@ -137,7 +157,261 @@ eq(lib.toRows(pairs.data, 2, 1, lib.perElement([2])), [["[2, 3]"], ["[4, 5]"]], eq(lib.formatValue(0.1 + 0.2), "0.3", "float formatting"); eq(lib.formatValue(2n ** 64n - 1n), "18446744073709551615", "bigint formatting"); eq(lib.formatValue("x"), '"x"', "string formatting"); +eq(lib.formatBytes(0), "0 B", "bytes"); +eq(lib.formatBytes(1536), "1.5 KiB", "KiB"); +eq(lib.formatBytes(200 * 1024 * 1024), "200 MiB", "MiB"); +eq(lib.formatStats({ lazy: true, requests: 3, bytesFetched: 2 << 20, size: 200 << 20 }), + "3 requests, 2 MiB of 200 MiB fetched (1.0%)", "stats line"); +eq(lib.formatStats({ lazy: false, requests: 1, bytesFetched: 1024, size: 1024 }), + "downloaded whole (1 KiB): the server does not support range requests", "stats line, no ranges"); h5.free(); nc.free(); - console.log(`wasm package: ${checks} checks passed`); + +if (base) await remoteTests(); + +async function serverStats() { + return (await fetch(`${base}/__stats`)).json(); +} + +async function remoteTests() { + const before = checks; + await fetch(`${base}/__reset`); + + // The same checks over HTTP range requests, at the default block size + // and at 512-byte blocks with a 4 KiB cache (almost every structure read + // a miss, evictions between calls). + for (const opts of [undefined, { blockSize: 512, cacheSize: 4096 }]) { + for (const [name, exp] of Object.entries(expected)) { + const f = await pkg.openUrl(`${base}/fix/${name}`, opts); + await checkFile(`${name} (openUrl ${JSON.stringify(opts ?? {})})`, exp, f); + const st = f.stats(); + eq(st.lazy, true, "read by ranges"); + eq(st.size, statSync(join(fixDir, name)).size, "size"); + f.free(); + } + } + const remote = await pkg.openUrl(`${base}/fix/fixture.h5`); + await checkErrors(remote); + eq((await (await pkg.openUrl(`${base}/fix/fixture.nc`)).info("/time")).maxshape, [null], "remote unlimited"); + + // What the page counts is what the server served. + await fetch(`${base}/__reset`); + const counted = await pkg.openUrl(`${base}/fix/fixture.nc`, { blockSize: 1024 }); + await counted.read("/temp"); + const server = await serverStats(); + eq(counted.stats().requests, server.requests, "requests counted"); + eq(counted.stats().bytesFetched, server.bytes, "bytes counted"); + + // A custom fetch is used for every request, with the caller's headers. + let calls = 0; + const seen = new Set(); + const viaCustom = await pkg.openUrl(`${base}/fix/fixture.h5`, { + blockSize: 4096, + headers: { "X-Test": "1" }, + fetch: (url, init) => { calls++; seen.add(init.headers["X-Test"]); return fetch(url, init); }, + }); + eq((await viaCustom.read("/sensors/temp")).data[0], 21.5, "custom fetch values"); + eq(calls, viaCustom.stats().requests, "custom fetch calls"); + eq([...seen], ["1"], "headers passed"); + + // Calls in flight at once share the cache (a 1 KiB budget: nothing is + // evicted while any of them runs) and each gets its own answer. + const both = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, cacheSize: 1024 }); + const [g, t, l, a] = await Promise.all([ + both.read("/grid"), both.read("/sensors/temp"), both.list("/sensors"), both.attrs("/"), + ]); + const exp = expected["fixture.h5"]; + eq(Array.from(g.data), exp.datasets["/grid"].values, "concurrent /grid"); + eq(Array.from(t.data), exp.datasets["/sensors/temp"].values, "concurrent /sensors/temp"); + eq(l.map((c) => c.name).sort(), [...exp.lists["/sensors"].groups, ...exp.lists["/sensors"].datasets].sort(), "concurrent list"); + eq(a.length > 0, true, "concurrent attrs"); + assert.ok(both.stats().cachedBytes <= 1024, "trimmed to the budget when idle"); + + // Listing and reading small things of a large file fetches a few blocks, + // not the file. + if (existsSync(join(fixDir, "big.json"))) { + const big = JSON.parse(readFileSync(join(fixDir, "big.json"), "utf8")); + await fetch(`${base}/__reset`); + const f = await pkg.openUrl(`${base}/fix/big.h5`); + const list = await f.list("/"); + eq(list.filter((c) => c.kind === "group").map((c) => c.name), big.list.groups, "big: groups"); + eq(list.filter((c) => c.kind === "dataset").map((c) => c.name).sort(), big.list.datasets, "big: datasets"); + eq(Array.from((await f.read("/small")).data), big.small, "big: /small"); + eq(Array.from((await f.read("/meta/ids")).data, String), big.ids, "big: /meta/ids"); + eq((await f.info("/big")).shape, big.big_shape, "big: shape"); + eq((await f.attrs("/meta"))[0].value, "m", "big: group attribute"); + const win = await f.readHyperslab("/big", [big.window.start], [big.window.count]); + eq(Array.from(win.data), big.window.values, "big: window"); + const st = f.stats(); + const server = await serverStats(); + eq(st.requests, server.requests, "big: requests counted"); + eq(st.bytesFetched, server.bytes, "big: bytes counted"); + console.log(`big.h5 (${big.size} bytes): listed, 3 small reads and a window in ` + + `${server.requests} requests, ${server.bytes} bytes (${(100 * server.bytes / big.size).toFixed(2)}%)`); + assert.ok(server.requests <= 8, `big: ${server.requests} requests`); + assert.ok(server.bytes * 20 < big.size, `big: ${server.bytes} bytes fetched`); + checks += 2; + } + + // A server without range support: downloaded whole (the default), or + // refused. + const whole = await pkg.openUrl(`${base}/norange/fix/fixture.h5`); + await checkFile("fixture.h5 (no range support)", expected["fixture.h5"], whole); + eq(whole.stats().lazy, false, "downloaded whole"); + eq(whole.stats().requests, 1, "one request"); + await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fallback: "error" }), + /does not support HTTP range requests/, "fallback: error"); + await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000 }), + /more than maxDownload/, "maxDownload"); + // Without a Content-Length the body is streamed, and stopped at the limit. + const undeclared = async (url, init) => { + const r = await fetch(url, init); + return new Response(r.body, { status: r.status }); + }; + await fails(() => pkg.openUrl(`${base}/norange/fix/fixture.h5`, { maxDownload: 1000, fetch: undeclared }), + /more than maxDownload/, "maxDownload, streamed"); + const streamed = await pkg.openUrl(`${base}/norange/fix/fixture.h5`, { fetch: undeclared }); + eq(Array.from((await streamed.read("/sensors/temp")).data), [21.5, 22, 22.25], "streamed download"); + + // Errors: HTTP status, not HDF5, bad options, a file that changes, a + // server that answers with the wrong bytes. + await fails(() => pkg.openUrl(`${base}/fix/missing.h5`), /HTTP 404/, "404"); + await fails(() => pkg.openUrl(`${base}/fix/expected.json`), /./, "not HDF5"); + await fails(() => pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 100 }), /blockSize/, "blockSize"); + const tamper = (edit) => async (url, init) => { + const r = await fetch(url, init); + return init.headers.Range === "bytes=0-511" ? r : edit(r); + }; + const withHeaders = async (r, headers) => { + const h = new Headers(r.headers); + for (const [k, v] of Object.entries(headers)) h.set(k, v); + return new Response(await r.arrayBuffer(), { status: r.status, headers: h }); + }; + await fails(async () => { + const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, fetch: tamper((r) => withHeaders(r, { ETag: '"other"' })) }); + await f.read("/grid"); + }, /changed on the server/, "changed file"); + await fails(async () => { + const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { + blockSize: 512, + fetch: tamper(async (r) => new Response((await r.arrayBuffer()).slice(1), { status: 206 })), + }); + await f.read("/grid"); + }, /got \d+/, "short answer"); + await fails(async () => { + const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { + blockSize: 512, + fetch: tamper(async (r) => new Response(await r.arrayBuffer(), { status: 200 })), + }); + await f.read("/grid"); + }, /stopped honouring range requests/, "200 mid-file"); + await fails(async () => { + const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { + blockSize: 512, + fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })), + }); + await f.read("/grid"); + }, /the server sent 0-511/, "wrong range"); + await fails(async () => { + const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { + blockSize: 512, + fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes */25752" })), + }); + await f.read("/grid"); + }, /unusable Content-Range/, "unusable Content-Range"); + + // Corpus files: what the viewer can show of each is the same read by + // ranges as in memory (an error wherever it gives one). + const corpus = process.env.CLAWHDF5_WASM_CORPUS; + if (corpus) await corpusTests(corpus); + console.log(`openUrl: ${checks - before} checks passed`); +} + +function hdf5Files(dir, out) { + for (const e of readdirSync(dir, { withFileTypes: true })) { + const p = join(dir, e.name); + let st; + try { + st = statSync(p); + } catch { + continue; + } + if (st.isDirectory()) hdf5Files(p, out); + else if (st.size <= 16 << 20) { + const b = readFileSync(p); + const sig = [0x89, 0x48, 0x44, 0x46, 0x0d, 0x0a, 0x1a, 0x0a]; + for (let at = 0; at + 8 <= b.length; at = at === 0 ? 512 : at * 2) { + if (sig.every((x, i) => b[at + i] === x)) { out.push(p); break; } + } + } + } + return out; +} + +function show(v) { + return JSON.stringify(v, (_, x) => { + if (typeof x === "bigint") return `${x}n`; + if (ArrayBuffer.isView(x)) return Array.from(x, (y) => (typeof y === "bigint" ? `${y}n` : Number.isNaN(y) ? "NaN" : y)); + return x; + }); +} + +async function transcript(f) { + const out = []; + const call = async (what, fn) => { + try { + out.push(`${what}: ${show(await fn())}`); + } catch { + out.push(`${what}: Err`); + } + }; + const todo = [["/", 0]]; + while (todo.length && out.length < 3000) { + const [path, depth] = todo.pop(); + let kind = null; + await call(`${path} kind`, async () => (kind = await f.kind(path))); + await call(`${path} attrs`, () => f.attrs(path)); + if (kind === "group") { + let list = []; + await call(`${path} list`, async () => (list = await f.list(path))); + if (depth < 12) for (const c of list.reverse()) todo.push([lib.joinPath(path, c.name), depth + 1]); + } else if (kind === "dataset") { + let info = null; + await call(`${path} info`, async () => (info = await f.info(path))); + if (info && [...info.shape, ...info.elementShape].reduce((a, b) => a * b, 1) <= 1 << 20) { + await call(`${path} read`, () => f.read(path)); + } + } + } + return out; +} + +async function corpusTests(corpus) { + const files = hdf5Files(corpus, []).sort(); + let opened = 0; + let bytes = 0; + let size = 0; + for (const p of files) { + const rel = relative(corpus, p).split("/").map(encodeURIComponent).join("/"); + let local; + try { + local = pkg.open(new Uint8Array(readFileSync(p))); + } catch { + await fails(() => pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 }), /./, `${rel}: opens in neither`); + continue; + } + const remote = await pkg.openUrl(`${base}/corpus/${rel}`, { blockSize: 65536 }); + const want = await transcript(local); + const got = await transcript(remote); + // Errors are compared as errors: a malformed file can fail at another + // check, with another message, when read by ranges. + eq(got, want, `${rel}: transcript`); + opened++; + bytes += remote.stats().bytesFetched; + size += remote.stats().size; + local.free(); + remote.free(); + } + console.log(`corpus: ${opened} of ${files.length} files agree over HTTP (${bytes} of ${size} bytes fetched)`); +} diff --git a/examples/wasm-viewer/viewer-lib.js b/examples/wasm-viewer/viewer-lib.js index 72c5bf3..fece1ab 100644 --- a/examples/wasm-viewer/viewer-lib.js +++ b/examples/wasm-viewer/viewer-lib.js @@ -75,3 +75,24 @@ export function toRows(data, rows, cols, per = 1) { export function perElement(elementShape) { return elementShape.reduce((a, b) => a * b, 1); } + +/** A byte count for people: `1.5 KiB`, `200 MiB`. */ +export function formatBytes(n) { + const units = ["B", "KiB", "MiB", "GiB", "TiB"]; + let i = 0; + let v = n; + while (v >= 1024 && i < units.length - 1) { + v /= 1024; + i++; + } + const s = i === 0 ? String(v) : v < 10 ? String(Number(v.toFixed(1))) : String(Math.round(v)); + return `${s} ${units[i]}`; +} + +/** What reading a remote file has cost, from `RemoteFile.stats()`. */ +export function formatStats({ lazy, requests, bytesFetched, size }) { + if (!lazy) return `downloaded whole (${formatBytes(size)}): the server does not support range requests`; + const pct = size > 0 ? (100 * bytesFetched) / size : 0; + return `${requests} request${requests === 1 ? "" : "s"}, ${formatBytes(bytesFetched)} of ` + + `${formatBytes(size)} fetched (${pct.toFixed(1)}%)`; +} -- 2.54.0 From b8c85f7627a612c3e376fac76ffad1f828e746bd Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 06:44:51 -0500 Subject: [PATCH 06/30] wasm-viewer: open by URL, lazily, with a request counter The page gets a URL box; ?file= now opens with openUrl instead of downloading the file, and the header shows the requests made and bytes fetched so far ("5 requests, 5 MiB of 191 MiB fetched (2.6%)"). Every file call is awaited, so local files (open) and remote ones share the code; a selection that finishes after another was made is not shown. browser.sh serves the page and fixtures with test/serve.py (no symlinks into the repository any more) and also checks a server without range support and, on the 200 MB file, that a small dataset and a window of the big one render with a single-digit percentage fetched. Co-Authored-By: Claude Opus 5.5 (1M context) --- examples/wasm-viewer/index.html | 138 ++++++++++++++++++--------- examples/wasm-viewer/test/browser.sh | 76 +++++++++------ 2 files changed, 140 insertions(+), 74 deletions(-) diff --git a/examples/wasm-viewer/index.html b/examples/wasm-viewer/index.html index d477e49..279c9b8 100644 --- a/examples/wasm-viewer/index.html +++ b/examples/wasm-viewer/index.html @@ -56,27 +56,39 @@ border: 1px solid var(--line); border-radius: 4px; } .error { color: var(--bad); font-family: var(--mono); white-space: pre-wrap; } .muted { color: var(--muted); } + form.url { display: flex; gap: 6px; margin-left: auto; flex: 1 1 320px; max-width: 560px; } + form.url input { flex: 1; min-width: 0; font: 13px var(--mono); padding: 4px 8px; background: var(--panel); + color: var(--ink); border: 1px solid var(--line); border-radius: 6px; } + .netstats { flex-basis: 100%; color: var(--muted); font-size: 12.5px; font-family: var(--mono); } + .netstats:empty { display: none; }

HDF5 Viewer

no file +
+ + +
+
-

Drop an HDF5 or NetCDF-4 file here, or use “Open file…”.

-

The file is read in this page by clawhdf5 compiled to WebAssembly; it is not uploaded anywhere.

+

Drop an HDF5 or NetCDF-4 file here, use “Open file…”, or give a URL.

+

The file is read in this page by clawhdf5 compiled to WebAssembly; a local file is not uploaded anywhere. + A file given by URL is not downloaded: only the byte ranges each view needs are fetched (HTTP range requests), + and the requests and bytes it has cost are shown at the top.