- Datatype::Complex serializes class 11 version 5 byte-identically to
libhdf5 2.2.0; containers holding it are written as version 5.
- DatasetBuilder::with_complex_f32/f64_data (h5py's {r, i} compound,
default) and with_native_complex_f32/f64_data (class 11, opt-in);
make_(native_)complex_f32/f64_type for attributes.
- Dataset::read_complex_f64/f32 read either form.
- Python create_dataset accepts complex64/complex128 (compound form).
- Parsing unchanged: class 11 still surfaces as {r, i}.
- Tests vs h5py 3.16 / libhdf5 2.0.0 and h5dump 2.2.0; docs.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every crate under crates/ now has a README (android, bench, cli, napi and
wasm had none), each saying what the crate is, its main types and
functions (names checked against the code), its cargo features with
defaults and which ones build C (checked with `cargo tree`), and links to
the top-level docs.
Corrections to the old stubs:
- clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs
clawhdf5-format as a dependency.
- clawhdf5-filters: deflate backends only, and no library crate depends
on it; the filter pipeline and every other codec are in -format.
- clawhdf5-gpu: vector distance compute, not I/O; not used by
HDF5Memory::search.
- clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO.
- clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and
search(q, k, ef).
- clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm
backends are reported but run the scalar kernels.
- clawhdf5-gpu: the old example called l2_distances, which does not
exist (l2_search).
- clawhdf5-agent: it described a "vector store" with "GPU acceleration";
it now covers HDF5Memory, search options, WAL, signing, the graph.
- crates.io/docs.rs badges removed and `cargo install <crate>` replaced:
nothing is published; depend on git.
- fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh.
- tools: the FileEditor interop tests that live in this crate.
- remote, py: license, other front ends, limits, File.mode/flush/chunks.
The Rust examples of the facade, format, filters, accel, ann, derive and
agent READMEs were compiled and run as tests (netcdf4, gpu and remote
compiled only) in a scratch crate; the CLI example was run.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).
`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.
Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
earliest, 1 MiB: 7/74/193.6 MB -> 6/5/5.2 MB
earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
-> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.
- The v1 and v2 B-tree collectors descend into every child of a node
after one fails (they only read the siblings before, so a sibling's
subtree came a pass later), then return the first error: results and
errors unchanged. The v2 walk stops once its record budget is spent,
so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
a group B-tree node's and a symbol table node's body (read once their
header gives a length, a round trip later when the body is in the
next block), an object header's first chunk and its continuation
chunks, the symbol table nodes a B-tree leaf names, a dense group's
name index header and the heap's root block (both read right after
the heap header). A listing also hints every child's object header as
its entry is read, even after a failure, and every direct block of a
dense group's heap (reading the indirect blocks, at most 4096 entries
and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
n entries end) is one helper used by the object reads and the hints.
Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
earliest, 1 MiB: 6/73/192.5 MB -> 4/68/192.5 MB
earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
latest, 1 MiB: 9/98/196.5 MB -> 5/86/196.5 MB
latest, 64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A parser often learns where the next structures are (a node's children,
a structure's body once its prefix gives its length) before it reads
them one at a time. Over openUrl's restartable reader a structure only
reached after a miss costs a pass, and a round trip, of its own.
- clawhdf5-format: `Storage::hint(offset, len)`, "about to be read by
this operation". Default: nothing (every backend that reads when
asked); `&T`, `Box`, `Arc` and the facade's `FileData` forward it
(shifted past a user block, clamped to the file).
- clawhdf5-wasm `LazyStorage` records hinted blocks it lacks. A pass
that misses nothing ignores them (a hint never adds a round trip); a
pass that misses also asks for them, in file order, while the pass
stays within what is left of the operation's `maxFetch` budget (a
hint never makes a call fail). At most 1 MiB (or a block) per hint
and 65536 blocks per pass are recorded, whatever a file makes a
parser hint. `Operation::attempt` follows hints; the plain
`LazyStorage::attempt` does not. `run_blocking` and the browser's
driver use the former. `LazyStats::hinted_blocks` counts them.
No parser hints yet: results, passes and requests are unchanged.
Tests: hinted blocks come with a miss and never alone (cached ones
skipped, a plain attempt ignores them), and stay within the fetch
budget, a hint past the end or longer than the file harmless.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- opts.headers is read as fetch reads it (a Headers, [name, value] pairs
or a plain object); it was spread as an object, which silently dropped
a Headers instance (a common way to pass Authorization). A caller's
Range is not sent.
- When one range request of a batch fails, the others in flight are
aborted (one AbortController per batch, its signal passed to fetch)
and no new ones start; the first failure is the error. The workers
used to go on issuing requests nobody waited for.
- parallel must be an integer from 1 to 1024 (openUrl) or a positive
integer (fetchRanges): a non-number gave NaN workers, so none ran and
fetchRanges returned nothing.
Tests (test.mjs): headers as object, Headers and pairs reach the fetch;
bad parallel values are option errors; fetchRanges with a 500 on the
third of 20 ranges at parallel 3 starts 3 requests and aborts the 2 in
flight. Before: the Headers case sent no header, 17 requests started
after the failure with none aborted, and parallel "x" returned.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.
- core::Reader::list reads every child's header before returning the
first error (the same error, in listing order, Group::groups/datasets
return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
and v2 collectors, the symbol table node loop and the dense-link loop
go on reading (not using) the remaining siblings, then return that
first error: results and errors are unchanged, only failing
traversals read more, and in memory that is free (storage::touch).
A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
(it was fetched again: 215 MB fetched from a 198 MB file).
Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
libver earliest, 64 KiB: 537/536 -> 8/531 (6 in flight)
libver latest, 1 MiB: 189/188 -> 9/98
libver latest, 64 KiB: 453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
remote.js read every 206 body (the probe and each range) with
resp.arrayBuffer() and checked its length afterwards, so a hostile
server could make the page buffer gigabytes before the check failed.
Every body is now piped through a TransformStream that errors as soon as
the count passes the limit, which cancels the body (and the request):
the requested range length for a 206, maxDownload for the 200 fallback.
A declared Content-Length past the limit is refused before reading. The
200 path now always streams (it read a declared length at once, because
a reader loop stalls on small bodies in headless Chromium under
--virtual-time-budget; a pipe does not).
Test (test.mjs): a probe and a range answered with a 64 MiB body read at
most the range + one 64 KiB piece; before, all 64 MiB were read ("asked
for the first 1048576 bytes of 2000000, got 67108864").
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in
LazyStorage::assemble (capacity_overflow), taking every open file on the
page with it, and a hostile server only had to claim a large length and
serve a heap collection of 2 GiB + 4 KiB to get there (after fetching
2 GiB). Reading a large u8 dataset whole aborted the same way when its
values were widened to 64 bits.
- LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at
most 1 GiB): a read longer than it fails at once, before anything is
fetched, and an operation whose passes would fetch more than it fails
before fetching (Operation::charge). assemble reserves fallibly.
- Reader::read refuses a read that would use more than 1 GiB while
decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result)
with an error naming readHyperslab, before reading.
- openUrl refuses a file of 4 GiB or more at open on wasm32: the format
code turns offsets into usize, so nothing past 4 GiB can be read there
(shown by a new test: data at 3 GiB reads, a 4 GiB file is refused).
maxDownload is bounded to 1 GiB.
Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8
dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at
3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an
error or reads, and that the module survives. Before: RuntimeError:
unreachable in Node; the native test read the huge dataset and fetched
2 GiB.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
openUrl(url, opts) returns a RemoteFile with the methods of H5File
(kind, list, info, attrs, attrErrors, read, readHyperslab), each a
promise, and stats(). It runs every call through the restartable
LazyStorage: a pass that misses reports the byte ranges, js/remote.js
fetches them with fetch() and Range headers (six at a time), and the
pass is re-run. This keeps the main thread free without a Worker or
synchronous XHR (h5wasm's lazy files need both), as the design doc
recommends; the cost is re-running a pass per wave of misses.
Every answer is checked: a 206 with exactly the bytes asked for, and
the same ETag/Last-Modified and length as at open, else an error (never
data). A server that ignores Range (200) is downloaded whole, up to
maxDownload (512 MiB), unless fallback: "error". Options: blockSize,
cacheSize, headers, credentials, parallel, fetch.
test/serve.py is a range-capable static server with request counting
(and /norange/ for a server without range support). test.mjs repeats
every fixture check on files opened by URL (1 MiB and 512 B blocks),
checks the request budget on a 200 MB h5py file (list, three small
reads and a window of the big dataset: 5 requests, 6 MiB), the
download fallback, and HTTP errors, changed files and wrong answers;
with CLAWHDF5_WASM_CORPUS every corpus file is compared with open(bytes).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
LazyStorage holds the blocks of a remote file fetched so far. An
operation runs as passes over it: a read that misses records the
missing blocks and fails, the pass's result is dropped whatever it is
(a parser may have caught the error and carried on), and the caller
fetches the reported ranges and re-runs the pass. No block is evicted
while an operation is in flight, so every pass that does not finish
asks for at least one new block and the operation ends. Blocks sit in
an LRU with a byte budget, trimmed between operations, bulk (raw data)
blocks first.
Reader::open_storage opens a file through any Storage, and variable-
length strings resolve through the file's storage instead of
File::as_bytes, which panics for a file not in memory.
tests/lazy.rs compares, file by file, what the viewer can show (kinds,
listings, attributes, info, whole reads and hyperslabs) read lazily
with the facade's range-storage path and the in-memory reader: files
built here at 512 B to 1 MiB blocks, the h5py/netCDF4 fixture, and with
CLAWHDF5_WASM_CORPUS the conformance corpus (656 files agree). A
listing plus a small read of a 48 MB file fetches 3 ranges.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
libhdf5 fails to read a VL element whose global heap address is
undefined (all 0xff), even at length 0 ("addr undefined"); we returned
"" (or an empty sequence) in every reader. Checked with h5py first:
libhdf5 writes a null element with address 0, which still reads as
empty, and h5py writes "" as a zero-size heap object at a real address,
so no file they write relies on the old behaviour. read_vl_bytes now
treats address 0 as null whatever the length, as VlResolver does.
Tests, each failing before: vl_data unit test (8- and 4-byte offsets,
lengths 0 and 1); clawhdf5 vl_data_interop
a_vl_element_at_the_undefined_heap_address_fails_like_h5py (also checks
where h5py writes ""); h5rs dump --json and check --data on the patched
`undef` dataset; clawhdf5-wasm vl_strings.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
h5rs (dump, ls, diff, check --data) kept its own lenient VL decoder:
a heap object longer than its element was cut to the element's length
(libhdf5 and h5py refuse it), a null string printed "" where h5dump
prints NULL, the stored element size was trusted, and every heap
collection was kept as an owned copy for the whole run. It now resolves
each element with VlResolver::element / string_element (new: one element
in place, borrowing from the file), and refuses a VL type whose stored
element size is not 4 + offset size + 4, as File does. H5::heap_object
and its cache are gone. h5diff compares a null VL string equal to an
empty one; so does h5rs diff.
clawhdf5-wasm already resolved VL strings with read_vl_strings; it now
uses VlResolver and checks the stored element size before reading, as
File::read_string does.
Tests (h5py writes the files, patched for "a\0b", a null element and
mis-sized heap objects, with 8- and 4-byte offsets):
- h5rs_interop dump_prints_vl_data_like_h5dump: byte-identical to h5dump;
- dump_json_vl_values_match_h5py: h5py's values, errors where h5py fails;
- check_data_flags_mis_sized_vl_heap_objects;
- clawhdf5-wasm tests/vl_strings.rs: wasm, File and h5py agree.
All four fail before. check --data over the 150 cve_hdf5 CVE and fuzzer
files now passes 15 (h5dump rejects 8 of them), was 16 and 9: the
stored-size check flags cve-2024-32608. h5rs-check-ok-files.sh --data:
0 of 422 flagged; h5rs-fuzz.sh: clean on 180 files.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
cargo test --workspace unifies clawhdf5-format/zstd on (another member
enables it), so the native interop test read the Zstd dataset that the
wasm build refuses. The fixture now records its values plus the error
the wasm build must give; the native test accepts either, the Node test
of the real wasm package still requires the error.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Drop a file (or pass ?file=<url>&path=<object>), browse the tree lazily,
see a dataset's type, shape, max shape and attributes, and page through
its values as 50x12 hyperslab windows (leading dims of 3-D+ data held
at chosen indices). build.sh produces pkg/ (not committed) with
wasm-bindgen --target web and checks the CLI matches the crate version.
test/run.sh builds it and runs test.mjs under Node against the h5py/
netCDF4 fixture (250 checks: every dataset whole and as a strided
hyperslab, listings, attributes, error paths, the page's DOM-free
helpers), then browser.sh renders the page in headless Chromium for
eight objects and checks the DOM. The fixture gains LZ4 (read) and Zstd
(refused: links C) datasets and a compound attribute (value null plus
its type). ci-test.sh runs it when node and wasm-bindgen exist; the CI
container has neither, so CI relies on the native h5py_interop test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
open(bytes) -> H5File with kind/list/info/attrs/attrErrors/read/
readHyperslab. Numeric data comes back in the typed array of the
stored width (Int16Array for i16, BigInt64Array for i64, Float32Array
for f32/f16, ...), strings and enum names as string arrays, array
datatypes flattened with their dims appended to the shape. Compound,
reference, opaque and VL-sequence datasets are refused with an error
naming the type; nothing is returned as reinterpreted bytes.
The logic is in a plain-Rust core module, tested natively: unit tests,
and h5py_interop, which compares every dataset, hyperslab, listing and
attribute of an h5py- and a netCDF4-written file with what libhdf5
reads back (generator shared with the Node test of the built package).
No mmap, no threads; lz4 is on, zstd/szip (C) are not. A
wasm-release profile (opt-level s, LTO) serves the browser build.
ci-test.sh lints the crate for wasm32 and checks it builds no C.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>