Commit Graph
279 Commits
Author SHA1 Message Date
osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00
osobhandClaude Opus 5.5 b55b24b7ba docs: crate READMEs describe each crate as it is today
Every crate under crates/ now has a README (android, bench, cli, napi and
wasm had none), each saying what the crate is, its main types and
functions (names checked against the code), its cargo features with
defaults and which ones build C (checked with `cargo tree`), and links to
the top-level docs.

Corrections to the old stubs:
- clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs
  clawhdf5-format as a dependency.
- clawhdf5-filters: deflate backends only, and no library crate depends
  on it; the filter pipeline and every other codec are in -format.
- clawhdf5-gpu: vector distance compute, not I/O; not used by
  HDF5Memory::search.
- clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO.
- clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and
  search(q, k, ef).
- clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm
  backends are reported but run the scalar kernels.
- clawhdf5-gpu: the old example called l2_distances, which does not
  exist (l2_search).
- clawhdf5-agent: it described a "vector store" with "GPU acceleration";
  it now covers HDF5Memory, search options, WAL, signing, the graph.
- crates.io/docs.rs badges removed and `cargo install <crate>` replaced:
  nothing is published; depend on git.
- fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh.
- tools: the FileEditor interop tests that live in this crate.
- remote, py: license, other front ends, limits, File.mode/flush/chunks.

The Rust examples of the facade, format, filters, accel, ann, derive and
agent READMEs were compiled and run as tests (netcdf4, gpu and remote
compiled only) in a scratch crate; the CLI example was run.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:30 -05:00
osobh bf5a163dcf Merge branch 'perf/wasm-listing-passes' into feat/listing-header-last-files
# Conflicts:
#	CHANGELOG.md
2026-09-27 23:16:27 -05:00
osobhandClaude Opus 5.5 761bdbf24f format: look a name up in a v1 group down its B-tree, as libhdf5 does
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).

`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.

Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
  earliest, 1 MiB:  7/74/193.6 MB -> 6/5/5.2 MB
  earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
  latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
  -> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
  commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:47:47 -05:00
osobhandClaude Opus 5.5 6f5d14fd62 format: group walks go on past a failed node and hint what they read next
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.

- The v1 and v2 B-tree collectors descend into every child of a node
  after one fails (they only read the siblings before, so a sibling's
  subtree came a pass later), then return the first error: results and
  errors unchanged. The v2 walk stops once its record budget is spent,
  so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
  a group B-tree node's and a symbol table node's body (read once their
  header gives a length, a round trip later when the body is in the
  next block), an object header's first chunk and its continuation
  chunks, the symbol table nodes a B-tree leaf names, a dense group's
  name index header and the heap's root block (both read right after
  the heap header). A listing also hints every child's object header as
  its entry is read, even after a failure, and every direct block of a
  dense group's heap (reading the indirect blocks, at most 4096 entries
  and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
  n entries end) is one helper used by the object reads and the hints.

Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
  earliest, 1 MiB:  6/73/192.5 MB -> 4/68/192.5 MB
  earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
  latest,   1 MiB:  9/98/196.5 MB -> 5/86/196.5 MB
  latest,   64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:47:44 -05:00
osobhandClaude Opus 5.5 ca81c3ebfa format, wasm: Storage::hint, fetched by the lazy reader with a pass's misses
A parser often learns where the next structures are (a node's children,
a structure's body once its prefix gives its length) before it reads
them one at a time. Over openUrl's restartable reader a structure only
reached after a miss costs a pass, and a round trip, of its own.

- clawhdf5-format: `Storage::hint(offset, len)`, "about to be read by
  this operation". Default: nothing (every backend that reads when
  asked); `&T`, `Box`, `Arc` and the facade's `FileData` forward it
  (shifted past a user block, clamped to the file).
- clawhdf5-wasm `LazyStorage` records hinted blocks it lacks. A pass
  that misses nothing ignores them (a hint never adds a round trip); a
  pass that misses also asks for them, in file order, while the pass
  stays within what is left of the operation's `maxFetch` budget (a
  hint never makes a call fail). At most 1 MiB (or a block) per hint
  and 65536 blocks per pass are recorded, whatever a file makes a
  parser hint. `Operation::attempt` follows hints; the plain
  `LazyStorage::attempt` does not. `run_blocking` and the browser's
  driver use the former. `LazyStats::hinted_blocks` counts them.

No parser hints yet: results, passes and requests are unchanged.
Tests: hinted blocks come with a miss and never alone (cached ones
skipped, a plain attempt ignores them), and stay within the fetch
budget, a hint past the end or longer than the file harmless.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:41:09 -05:00
osobhandClaude Opus 5.5 96086add99 format: inline the version-1 message loop again (ObjectHeader::parse back to 8f59b2e's speed)
4313917 kept parse_v1_messages out of line (#[inline(never)]) so the
generic parser would not carry the loop; since ef428d7 the slice path is
compiled once in this crate, and the call itself was the remaining cost
of the chunk queue: A/B builds of object_header_parse_x401 with only this
attribute changed put #[inline(never)] and no attribute at 24.5-24.9 us
and #[inline] at 23.6-24.0 us, with 8f59b2e at 23.6-24.1 us. Lazily
creating the chunk list only when a continuation is found (tried too)
measured no faster and was not kept.

Same code otherwise: every chunk-queue check (65,536 chunks, cycles, file
size budget, one chunk buffer at a time, libhdf5 order, overlap allowed)
is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:32:11 -05:00
osobhandClaude Opus 5.5 4ad80736f4 format: the chunk cache keeps the index's chunk order
ChunkCache::chunks_for returned the chunk index's values in hash-map
order, which differs between File instances (a new HashMap per open).
Readers that stop at the first failing chunk therefore named different
chunks on different opens of the same damaged file: whole-dataset
selections here, and the indexed reader behind the storage harness's
intermittent cve-2025-2310.h5 failure. The cache now also keeps the
chunks in the order the index lists them (fetched or built under one
lock) and returns that order, as the uncached readers use.

Regression: several_damaged_chunks_report_the_same_chunk_every_time
(four chunks inflating to different short lengths; before the fix two
opens reported chunk [40, 0] and chunk [320, 0]).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 08:03:11 -05:00
osobh e80b12521e Merge branch 'feat/p3-m4-wasm-lazy' into feat/p3-wasm-swmr-python
# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
2026-09-27 07:59:14 -05:00
osobh 5d42b4241c Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python
# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
2026-09-27 07:59:14 -05:00
osobhandClaude Opus 5.5 e553153e48 wasm, format: listing a group asks for all its missing blocks per pass
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.

- core::Reader::list reads every child's header before returning the
  first error (the same error, in listing order, Group::groups/datasets
  return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
  and v2 collectors, the symbol table node loop and the dense-link loop
  go on reading (not using) the remaining siblings, then return that
  first error: results and errors are unchanged, only failing
  traversals read more, and in memory that is free (storage::touch).
  A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
  (it was fetched again: 215 MB fetched from a 198 MB file).

Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
  libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
  libver earliest, 64 KiB:       537/536 -> 8/531 (6 in flight)
  libver latest,   1 MiB:        189/188 -> 9/98
  libver latest,   64 KiB:       453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:48:49 -05:00
osobhandClaude Opus 5.5 bdf584abb2 format: chunk grids with a zero extent have no chunks (no division by zero)
A Fixed or Extensible Array index whose maximum extent (the current one
when no maximum is recorded) is 0 along a dimension has a zero stride for
every dimension before it; ChunkGrid::offsets divided by it. The unfixed
editor made such files by resizing a clawhdf5-written dataset to a zero
extent: `h5rs check` panicked and the next resize raised an internal
error (12 of the reviewer's random-edit seeds 10..39). Such an index has
no slot for any chunk of the dataset; offsets now returns None.

Tests: chunk_grid::zero_extent_has_no_chunks; edit_interop's
zero_extent_resizes_without_a_recorded_maximum on a file the unfixed
editor left (fixture) and on a 2.7.0-written file taken through zero
extents with `h5rs check --data` and h5dump at every step; test_edit.py
random edits on seeds 10..39 of a clawhdf5-written file.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:40:11 -05:00
osobhandClaude Opus 5.5 1f7651644b edit: record the maximum before resizing a chunked dataset that has none
FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.

libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).

Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:40:01 -05:00
osobhandClaude Opus 5.5 2e906ffe3a facade: retry every read path of a live file (strings, vlen, attributes)
read_string, read_string_bytes, read_string_selection, read_vlen and
read_vlen_selection now retry as a whole, the global-heap decoding after the
read included; so do File::decode_strings / decode_string_bytes /
decode_vlen, Group::datasets / groups / attrs / attr, Dataset::attrs / attr,
the typed full reads (read_f64 ...), the header messages behind shape() and
dtype() (a shared message is read from another header), and
verify_provenance. Before, a transient failure on those reached the caller.

Attribute reads leave out an attribute they cannot read (or return a
variable-length string one as AttrValue::Raw) instead of failing, which hid
a transient error as a missing or raw attribute: on a live file such an
error of a retried kind now runs the read again too, and after the last
attempt the last result is returned as before. The format crate gains
find_attribute_reporting_in, which returns the errors find_attribute_in
skips (dense name-index lookups dropped them). The zero-copy reads need the
file in memory, which a live file never is, so they have nothing to retry.

Test: a storage that fails one read with a checksum mismatch; for 14 read
paths over a new fixture (tests/fixtures/swmr_strings_attrs.h5, an h5py
copy with the SWMR-write flag: vlen strings, vlen int32, dense attributes),
each read the path makes fails once in turn and the path must return the
same result with exactly one retry. It fails on the previous commit
(root.attrs, read 0). Dataset::attr on dense attributes returned None
before find_attribute_reporting_in.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:38:12 -05:00
osobhandClaude Opus 5.5 e397376820 format: a SWMR-flagged superblock's data ends at the end of the file
A SWMR writer does not keep the superblock's end-of-file address up to
date: a copy h5py made of its own file mid-write records 715 in a
6 030-byte file. Superblock::data_end only ignored the recorded end when
it lay past the end of the file, so every open path bounded reads at
715: the file listed, but every chunked read failed ("unexpected EOF:
need 787 bytes, have 715") and h5rs check reported the chunk indexes
past the end of the file. libhdf5's SWMR reader skips the
end-of-allocation check for every read (H5FD_read); for a v3 superblock
with the SWMR-write flag the data now ends at the end of the file.

Test: tests/swmr_interop.rs over the mid-write copy (fixture), through
File::open, open_buffered and from_bytes, and against h5py's SWMR reader.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 06:31:15 -05:00
osobhandClaude Opus 5.5 ef428d756c format: slice entry points for the facade's hot *_in calls (local listing back to main's speed)
CI / test-arm64 (pull_request) Successful in 1m25s
CI / test (pull_request) Successful in 19m0s
Since the M2 merge the facade handed in-memory files to the generic
`*_in` parsers as `&[u8]` (`with_bytes!`), which instantiates them in
the facade crate, where the format crate's private helpers do not
inline without LTO: listing a 400-group v1 file through `File::open`
was 7-10% slower than main. `ObjectHeader::parse_in`,
`group_v2::{resolve_child_in, resolve_group_children_in,
resolve_path_any_in}` and `attribute::{extract_attributes_tolerant_in,
find_attribute_in}` now pass a storage with `as_contiguous()` to their
non-generic slice entry point, compiled once in the format crate; other
storages reach the same generic core as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 22:28:22 -05:00
osobhandClaude Opus 5.5 4313917b4d format: object headers without per-header allocations for their chunk list
Reading continuation chunks from a queue (7e5e920, a69c5be) allocated a
queue Vec and a BTreeSet of chunk starts for every header, and inlined
the per-chunk message loop into the generic parser: ObjectHeader::parse
over 401 version-1 headers went from 24.8 to 45.7 us.

ChunkSpans now keeps the first 8 chunks in an inline array (cycle check
by scan) and is also the read queue; only a header of more chunks
allocates (a boxed spill list and start set). The message loop of one
version-1 chunk is its own non-generic function. Same checks as before:
any number of chunks up to 65,536, cycles refused, chunks bounded by the
file size, one chunk buffer alive at a time, libhdf5 message order,
overlap allowed. The cycle test now also covers spilled chunk lists.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 20:30:16 -05:00
osobhandClaude Opus 5.5 f37e7ae326 format: bound ZFP chunk fetches at 4x, not n + n/4
The M2 fetch bound gave every codec n + n/4 + 4096 stored bytes, but
ZFP's fixed-rate mode stores up to 64 bits per value, doubling 4-byte
types: 60 of the 2205 zfp_interop datasets (rate 64, f4/i4) failed with
'compressed stream ends early' once M2 and ZFP were merged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 19:33:03 -05:00
osobh 7447dce121 Merge branch 'feat/p3-editor-coverage' into feat/p3-remote-editor
# Conflicts:
#	CHANGELOG.md
#	CLAUDE.md
#	docs/design/range-reads.md
2026-09-26 19:17:46 -05:00
osobh 93e2d5f365 Merge branch 'feat/p3-m3-remote' into feat/p3-remote-editor
# Conflicts:
#	crates/clawhdf5/tests/storage_equivalence.rs
2026-09-26 19:17:32 -05:00
osobhandClaude Opus 5.5 ea0508aaa5 format: no truncating u64 -> usize casts in gather_storage and ExtentBytes
check-32bit-casts.sh flagged two casts added by the previous commits; both
values are bounded (checked by gather_storage's first walk, and built from
a usize fetch length), so they go through addr::saturating_usize.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 19:05:45 -05:00
osobhandClaude Opus 5.5 a69c5be8b2 format: read object header chunks from a queue, one buffer at a time
The version-1 chunk walk nested continuation chunks depth-first and kept
every enclosing chunk's buffer alive, up to 65 536 chunks. With storage
that hands out owned buffers (CountingStorage, the Storage trait, remote
storage) a crafted chain of chunks nested in each other read and held the
square of the file's size (a 192 KB file read 768 MB).

Chunks are now read from a FIFO queue of (address, length) pairs in the
order their continuation messages are found, as H5O_protect does and as
the editor's header walker already did, each buffer released before the
next read. In both header versions a chunk starting at an address seen
before (cycle) is refused, and so are chunks adding up to more than the
file, which bounds a header's reads by the file's size. Overlap itself is
allowed: libhdf5 reads cve-2025-7067.h5, whose continuation chunk overlaps
chunk 0 (refusing overlap cost that conformance file).

Tests: the nested chain is refused having read at most the file (it read
n^2 bytes before); a 3000-chunk chain reads each chunk once; a chunk's
messages follow the whole previous chunk (they were inserted at the
continuation message); an overlapping continuation chunk is read.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:50:23 -05:00
osobhandClaude Opus 5.5 159e588550 filters: Fletcher-32 as libhdf5 computes it
Our checksum reduced its sums with `% 65535`; libhdf5's
H5_checksum_fletcher32 folds them with `(s & 0xffff) + (s >> 16)`, which
leaves 0xffff where the modulo leaves 0. On about one chunk in 32768
libhdf5 refused the chunks we wrote and we refused the chunks it wrote.
Every release since v2.1.0 is affected.

clawhdf5_format::checksum::fletcher32 is a port of H5_checksum_fletcher32
and the filter's only implementation. Verification also accepts the
byte-swapped form libhdf5 accepts (1.6.2 and earlier) and the `% 65535`
form earlier releases wrote, so their files stay readable.

The new interop test compares the checksum with libhdf5's own function
(ctypes) on every 1- and 2-byte input and 40 000 random and fold-heavy
inputs, and moves fold-case chunks between h5py and FileBuilder/FileEditor
in both directions; with the old filters.rs the three file tests fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:50:23 -05:00
osobhandClaude Opus 5.5 6185874f9c format, clawhdf5: cut every Storage read to the range asked for
ExtentBytes and read_exact_at/read_upto rejected short results but passed
longer-than-asked ones through, and FileData forwarded them too, so a
Storage that broke read_at's contract by returning extra bytes had them
decoded or returned as data (a contiguous dataset read gained 37 junk
bytes). gather_storage alone trimmed.

- storage::exact_len (new, pub): a read of len bytes as exactly len — cut
  when longer, an error when short. read_exact_at, read_upto and
  ExtentBytes (so chunk fetches and selection gathers) go through it.
- FileData cuts a backend's answer to what it asked for before laying the
  cache image over it.
- Tests: over a storage that appends 37 junk bytes to every read, every
  format-crate fixture reads exactly as from the slice
  (overlong_reads_are_cut_to_the_range_asked_for), and every facade
  fixture opens and reads through File::open_storage as through File::open
  (overlong_storage_reads_identically). Both failed before.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:43:05 -05:00
osobhandClaude Opus 5.5 b086dc3c2b format: read a contiguous selection's runs merged across small gaps
gather_storage merged only runs that touch, so a strided selection of a
contiguous dataset over a Storage became one range (and one owned Vec) per
element: a stride-2 read of 32M f32 through File::open_storage made
16,777,232 read_at calls, took 2.0 s and peaked at 2.09 GB.

The selection is now walked twice. The first walk checks the runs and plans
spans: runs in increasing order at most 4 KiB apart (GATHER_GAP_BYTES) are
read as one span up to 8 MiB (GATHER_SPAN_BYTES; a longer run is split), so
nothing is stored per run. The spans are fetched in RAW_BATCH_BYTES batches
while the second walk copies each run out of its span. Same checks and
errors as before.

The same read is now 32 reads and 0.31 s (File::open: 0.08 s).
contiguous_read_interop: every h5py-checked selection is also read through
File::open_storage and must give libhdf5's bytes; a new test bounds the
range reads of strided, blocked, column and point selections (stride 2: at
most 1 data read; 563,200 before).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:34:31 -05:00
osobhandClaude Opus 5.5 7d629f49e3 format: bound and batch every chunk fetch over Storage
Only the full read split its chunk fetches into 64 MiB batches. The
selection path, the indexed read and the parallel_read decoders fetched
every chunk's stored bytes in one read_ranges call, each extent bounded only
by the file length, so a crafted chunk index pointing many chunks at one
large extent made File::open_storage hold chunks x extent bytes (3.3 GB from
a 16.8 MB file) before the first decode error.

- storage::for_each_extent_batch is now the one way raw-data reads fetch
  chunk bytes: batches of at most RAW_BATCH_BYTES (now pub), each decoded
  before the next is fetched. Used by the full, cached, indexed, selection
  and parallel_read paths; the sweep read uses read_extent per chunk.
- ExtentReq carries each chunk's claimed extent (bounds-checked as before,
  same errors) and the prefix actually fetched:
  filters::stored_chunk_limit — the chunk size if unfiltered, else each
  applied filter's worst-case growth (n + n/4 + 4096 per codec; unbounded
  only for an application-registered codec). The in-memory path cuts the
  slice it decodes the same way, so both paths still agree.
- tests/raw_fetch_bounds.rs: a crafted chunked_large.h5 (ten chunks all
  claiming 20 MiB at one padding blob) read through every path over a
  storage that records the largest single fetch; and 160 MiB of legitimate
  unfiltered chunks fetched batch by batch. Before: one 80 MiB fetch
  (selection) and one 160 MiB fetch; after: within the budget.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:30:30 -05:00
osobhandClaude Opus 5.5 a4f586e657 format: VlResolver::element_in and string_element_in over any Storage
VlResolver::element and string_element return slices of the whole file,
so they exist only for a resolver over &[u8]. Their *_in forms work for
any Storage (a remote file): the element's bytes borrowed from the
resolver's cache of heap collections, with the same null-element, NUL
and size checks. h5rs decodes variable-length values with them.

Test: over a read_at-only storage they give what element/string_element
give over the slice, for a string with an embedded NUL, a null element
and an element whose heap object has the wrong size.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 17:20:50 -05:00
osobhandClaude Opus 5.5 7e5e920c72 format: read object headers with long continuation chains
libhdf5 follows any number of continuation chunks, and a header that is
full gains one per message added (each new chunk holding the next
continuation message), so a version-1 header with a few dozen attributes
added one at a time is a chain dozens of chunks long. The reader recursed
once per chunk and refused a chain deeper than 32 (NestingDepthExceeded):
h5py read such files, we did not. Version-2 headers stopped at 256
continuation chunks.

Version-1 chunks are now followed with an explicit stack (the same
depth-first message order as before), version-2 ones as before; both
refuse a chunk address seen twice (a cycle, what the limits guarded
against) and more than 65 536 chunks.

Regression: long_v1_continuation_chains_are_read (a 200-chunk chain),
v1_continuation_cycles_are_refused; the dense-attribute interop test's
'earliest' case produces such a chain.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 16:56:09 -05:00
osobhandClaude Opus 5.5 1c3ef98828 clawhdf5: File::open_storage reads any Storage through the full API
File::open_storage(Arc<dyn Storage + Send + Sync>) opens a file served by
any backend: groups, datasets, attributes, read_*, selections, VL data
and virtual datasets (external sources through the new
File::set_vds_resolver) all read through Storage::read_at/read_ranges.
The file's view (user block skipped, bounded by the recorded end of
file) is itself a Storage; File::open and File::from_bytes keep their
mmap and in-memory paths, now as that view's as_contiguous() fast path.
A storage-backed file's metadata cache image is laid over each read it
covers (new CacheImage::entries), as libhdf5 loads it.

The typed readers keep their fast paths over any storage: a contiguous
dataset is read in one piece and converted (read_f64 and friends), and a
contiguous native selection reads only its runs
(data_read::read_selection_native_in). Zero-copy methods answer
ContiguousStorageRequired when the bytes are not in memory, and
File::as_bytes panics there (File::contiguous_bytes is the fallible
form).

Test: tests/storage_equivalence.rs reads every fixture, and with
CLAWHDF5_STORAGE_CORPUS every conformance-corpus file, through File::open
and through open_storage over a read_at-only CountingStorage — tree,
attributes, and every dataset's values several ways — and requires
identical transcripts; it prints the read_at calls and bytes a one-pass
read costs per file.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 16:41:47 -05:00
osobhandClaude Opus 5.5 17201e279d format: implicit chunk index addresses over the maximum chunk grid
libhdf5 allocates an implicit index's chunks for the whole maximum
extent and places chunk `scaled` at its row-major position in the
maximum chunk grid (H5D__none_idx_get_addr, max_down_chunks). The reader
used the current grid, so a dataset below its maximum shape with more
chunk columns at its maximum read other chunks' values from the second
chunk row on (h5py early allocation, fixed maxshape).

generate_implicit_chunks_in_grid takes the maximum dimensions;
generate_implicit_chunks keeps its signature (grid = current extent).
Regression: implicit_chunks_use_the_maximum_grid here, and
implicit_index_below_its_maximum_reads_like_libhdf5 (h5py file) with the
editor tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 16:35:06 -05:00
osobhandClaude Opus 5.5 3fa5ed1dda format: raw data, VDS and VL data over Storage
Every raw-data path has a generic *_in core, with the &[u8] functions as
thin wrappers: data_read (read_raw_data*, read_raw_data_selection,
read_chunked_native), chunked_read (the v1 B-tree chunk index, list_chunks,
the full, cached, sweep and indexed reads), parallel_read, partial_read,
fill_value (read_full_with_fill, apply_to_unallocated_chunks; and
dataset_fill_value_from_storage is now generic), vds (the virtual file
through Storage, external sources still through the resolver),
vl_data (VlResolver<'a, S = [u8]>, read_vl_strings_in, read_vl_bytes_in),
AttributeMessage::read_vl_strings_in and provenance::verify_dataset_in.

With the whole file in memory nothing changes: chunks and contiguous data
are sliced from it as before. Otherwise a chunked read lists its chunks,
fetches their stored bytes with one Storage::read_ranges call per 64 MiB
batch (chunks the cache already holds are not fetched), then decodes as
today; a selection fetches only the chunks it overlaps, and a contiguous
selection only its runs. Each extent's bounds error is the one the slice
code gave, reported when that extent is reached, so errors keep their
order.

Tests: the equivalence harness now reads every dataset's values (whole,
fill-aware, cached, indexed, three selections, VDS, VL strings and
sequences) through the read_at-only storage and requires the slice
results (all 653 corpus files agree); a misbehaving storage (a failing
Nth read, short reads) only ever yields errors or the right values; and
chunked reads are checked to use one read_ranges call.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 16:28:01 -05:00
osobhandClaude Opus 5.5 42894bf93b format: v2 B-trees, dense groups and group listings over Storage
BTreeV2Header::parse_in, collect_btree_v2_records_in and
find_btree_v2_records_in read one bounded window per node (its size is
known from the parent before the node is read; a count stretched past
node_size is checked against the end of the file first), with the
whole-file bounds errors unchanged. With them, dense attributes, a SOHM
B-tree index and huge fractal-heap objects no longer answer
ContiguousStorageRequired, and group_v1/group_v2 listings, lookups and
path resolution get *_in cores (resolve_group_children_in,
resolve_child_in, resolve_path_any_in, ...). The &[u8] functions are
thin wrappers, as in M1.

The equivalence harness now fails on any ContiguousStorageRequired and
compares v2 B-tree headers, records and descents, group listings, child
lookups and paths; a unit test compares a two-level tree through a
read_at-only storage truncated at every length and with every node byte
flipped.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 16:13:01 -05:00
osobh 8fadb9f424 Merge branch 'feat/p3-in-place-modify' into feat/p3-range-zfp-edit
# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-py/src/lib.rs
#	crates/clawhdf5/src/error.rs
2026-09-26 14:52:55 -05:00
osobh c233fbca6e Merge branch 'feat/p3-zfp' into feat/p3-range-zfp-edit
# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-format/Cargo.toml
2026-09-26 14:52:22 -05:00
osobh 437e81cfff Merge branch 'feat/p3-storage-trait' into feat/p3-range-zfp-edit
# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-format/src/attribute.rs
#	crates/clawhdf5-format/src/btree_v1.rs
#	crates/clawhdf5-format/src/data_layout.rs
#	crates/clawhdf5-format/src/extensible_array.rs
#	crates/clawhdf5-format/src/fixed_array.rs
#	crates/clawhdf5-format/src/fractal_heap.rs
#	crates/clawhdf5-format/src/local_heap.rs
#	crates/clawhdf5-format/src/shared_message.rs
2026-09-26 14:51:51 -05:00
osobhandClaude Opus 5.5 0e8522cfad clawhdf5-format: the writer skips optional filters that fail, as libhdf5 does
FileBuilder stored every chunk of an LZF or Blosc dataset through the
filter with filter mask 0. libhdf5 counts LZF and Blosc output no smaller
than the chunk as a failure of the optional filter and stores the chunk raw
with the filter's mask bit set. For an LZF chunk whose stream was exactly
the chunk's size, the first libhdf5 rewrite stored raw data at the same
size and kept our stale mask 0 in the index, and h5py could no longer read
the dataset.

precompress_chunks now runs chunks through compress_chunk_masked (as
FileEditor does since f7e2ab1), sequentially and on the parallel path, and
build_chunked_data_from_precompressed records each chunk's real mask in
every index the writer builds: single chunk (layout field), Fixed Array and
Extensible Array filtered elements, and version-2 B-tree type 11 records
(create_datasets_parallel goes through the same path). The writer builds
no version-1 B-tree or implicit index. PrecompressedChunks::chunks gains
the mask. Files whose chunks all compress are byte-identical.

Latent only in the unreleased LZF/Blosc writer (added 2026-09-26); no
tagged release writes either filter.

Tests:
- plugin_filters_interop skipped_optional_filters_are_masked_as_libhdf5_masks_them:
  LZF, shuffle+LZF+fletcher32 and Blosc over random, compressible and
  alternating chunks in every index; masks equal an h5py-written twin's;
  h5py r+ rewrites and extends them; h5py, h5dump and our reader read
  every value. Before: 20 of 24 datasets had masks other than h5py's, and
  with that check disabled h5py failed to read the rewritten datasets
  ("filter returned failure during read").
- plugin_filters_interop files_whose_chunks_all_compress_are_unchanged:
  pins the pre-fix bytes of five all-compressing files.
- chunked_write skipped_lzf_chunks_are_masked_in_every_index (fails before:
  mask 0, want 2).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:45:51 -05:00
osobhandClaude Opus 5.5 76c97f6c94 format: bound Storage reads that hostile size fields could stretch
On a backend without the file in memory, a structure read whose length
comes from untrusted header fields was clamped only by the end of the
file, so a crafted size made one read (and copy) of up to the rest of
the file. Each such read now covers what the parser actually uses:

- local heap names: read in growing pieces (64 bytes first, then 4x)
  up to the end of the data segment, instead of the rest of the segment
  per name (quadratic for a big symbol-table group);
- fractal heap indirect blocks: the doubling-table geometry locates the
  entry covering the object, and the first read ends at that entry; only
  if it is unallocated does the walk read the rest of the block (it
  visits every entry then). One walk implementation serves both;
- paged fixed/extensible array data blocks over 1 MiB: the prefix and
  page bitmap, then each page in use on its own (smaller blocks are
  still one read);
- blocks under one checksum (non-paged array data blocks, extensible
  array index and super blocks): the bounds check that comes first (the
  checksum's; the page bitmap's for a super block) is made against the
  file length before reading (Window::check_extent), so a block claimed
  past the end of the file costs no read. With the checksum feature off
  the parser has no such first check and the old read stands.

Other windows were already bounded (the superblock and object header
prefixes, the fractal heap header by a u16, SOHM tables by u8/u16
counts) or are exact reads checked against the file length first.
In memory nothing changes: the pieces are borrowed slices.

Tests: CountingStorage over a crafted heap (16 MiB file, width and rows
0xFFFF: under 1 KiB read, 16.7 MB before), a heap segment claiming 64 MiB
(one 64-byte read per short name), long names at every piece boundary,
a fixed array block claimed past the end of a 16 MiB file (under 64
bytes read), and in the equivalence harness an h5py file with a 2.4 MB
fixed array block and a >1 MiB extensible array block, whole and cut at
97 points: every chunk index agrees with the slice read and the largest
takes 205 KB (2.4 MB and 1.2 MB when read whole).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:31:46 -05:00
osobhandClaude Opus 5.5 e5359354b7 format: equivalence harness only accepts the known whole-file fallbacks
The harness counted ContiguousStorageRequired from any parser as an
allowed difference, so a converted module that wrongly fell back to the
whole file would still pass. It now accepts the error only from the
three sites that are not converted yet (dense attribute storage, a SOHM
B-tree index, huge fractal-heap objects, all found through a v2 B-tree)
and only in the checks that can reach them; anything else fails with
the check and the site named.

Checked by making LocalHeap::parse_in return the error first: the
fixture and h5py runs fail ("local heap fell back to the whole file").

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:22:42 -05:00
osobhandClaude Opus 5.5 052098bf36 format: monomorphise the Storage parsers so local files stay as fast
Every `*_in` core and the read helpers take `file: &S` with
`S: Storage + ?Sized` instead of `&dyn Storage`, and the `&[u8]`
wrappers pass the slice itself, so they compile to a `[u8]` instance:
`as_contiguous()` inlines to `Some(self)` and each structure read is the
slice code's bounds check again, with no indirect call. `&dyn Storage`
still works (`S = dyn Storage`); there is one parser implementation.

Also, so the structure reads cost no more than the slice checks did:
- ObjectHeader::parse_in reads the prefix once (signature included)
  instead of the signature and then the prefix: two reads for a
  one-chunk header instead of three on a range backend;
- the symbol-table node and group B-tree (v1) loops walk their entries
  with chunks_exact over the bytes read, and the node's redundant second
  bounds check is gone (the entries' read is the check, same error);
- a version-1 header's message list is sized from its (capped) count.
Same results and errors; the unit and equivalence tests are unchanged.

New Criterion bench `clawhdf5/benches/local_metadata_bench.rs` over a
400-group version-1 file written by h5py (new fixture
`v1_groups_400.h5`): ObjectHeader::parse, symbol-table nodes, the group
B-tree walk and a facade listing, using only APIs that exist at f2ff2c4
so it builds there for an A/B.

Provisional A/B against f2ff2c4 (busy machine, not for docs): both
builds linked into one binary and timed in alternation, 200 rounds;
median ratio new/old: facade listing -0.5% to -3.5% (was +14%),
ObjectHeader::parse +1% to +2% (was +25%), symbol-table nodes -18%,
group B-tree walk -18%, local-heap names and resolve_group_children
within +-1.5%. An old-vs-old-copy run shows +-2% from code layout alone.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:22:37 -05:00
osobhandClaude Opus 5.5 04a7f6f6c7 read: of two links with one name, the first wins everywhere
A valid group has one link per name, but a damaged or hand-made one can
have two. resolve_child followed the first soft link of the name, the
listing skipped a dangling one and listed the name via a later link, and
path resolution followed the last symbolic link: three answers. All now
take the first link of the name (header message order in a compact group,
name index order in a dense one) and ignore the rest, even if the first
dangles. That is libhdf5's rule for compact groups (H5G__compact_lookup
stops at the first Link message); h5py opens nothing for a dangling first
link although a later one resolves. For a dense group libhdf5
binary-searches the index and may land on another of several exact
duplicates; documented on first_link_named. find_symbolic_link's v2 branch
was dead (only v1 groups reach it) and is now v1-only.

Test: an h5py compact group with soft links dup_A (dangling, or to /d) and
dup_B (the other), dup_B renamed to dup_A in the header and re-checksummed.
Lookup, path and listing through all three readers match h5py for both
orders. With the old group_v2.rs the path lookup returned 42 where h5py
opens nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:11:47 -05:00
osobhandClaude Opus 5.5 5b3d32b37d format: test the address overflow path on 64-bit hosts
addr::to_usize's error branch only ran where usize is narrower than u64,
and no such target runs tests in CI, so on x86_64 the test checked only
that every u64 fits. to_usize and saturating_usize are now the usize
instances of generic to_index/saturating_index; the test runs the same
code with u32 standing in for a 32-bit usize: values past u32::MAX
(including one an `as` cast would wrap to 0x1234) are Overflow, and the
saturating form clamps. A mutant that truncates instead fails the test;
the old addr.rs does not provide the helper the test needs.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:09:37 -05:00
osobhandClaude Opus 5.5 f7e2ab12f2 clawhdf5: FileEditor skips optional filters that fail, as libhdf5 does
The editor stored every chunk through the whole pipeline with filter mask 0.
For LZF that did not shrink a chunk, h5py instead stores it raw with the
filter's mask bit set. A chunk the editor stored LZF-encoded at exactly the
raw size was then rewritten raw by libhdf5 at the same size; libhdf5 does not
touch the index entry when the size is unchanged, so the stale mask 0 stayed
and h5py (and h5dump) could no longer read the dataset.

clawhdf5_format::filters::compress_chunk_masked runs the pipeline as
H5Z_pipeline does: an optional filter (H5Z_FLAG_OPTIONAL) that fails is
skipped and its bit set, a mandatory one fails the write, and LZF/Blosc
output no smaller than the input counts as failure, as in the reference
filters (their output buffer is the input's size). Deflate, LZ4, Zstd,
bitshuffle and bzip2 never fail on size in libhdf5 and are kept as before.

Test: edit_interop optional_filters_that_fail_are_skipped — the reviewer's
repro at every libver: the editor stores the chunk exactly as h5py does
(mask 1, size 5; shuffle+LZF+fletcher32 mask 2), h5py r+ rewrites and
extends the datasets, and h5py, h5dump and our reader read every value.
Fails on the previous editor (mask 0; h5dump cannot read /u8).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:09:34 -05:00
osobhandClaude Opus 5.5 b6cbd2319f format: checked chunk addresses on the parallel read path
Three `chunk_info.address as usize` casts behind the `parallel` feature
survived the conversion, because check-32bit-casts.sh linted only default
features plus plugin-filters. On a 32-bit target with rayon a chunk address
past 4 GiB still wrapped onto another part of the file. They go through
addr::to_usize now, and the lane index (h % n, always < n) through
saturating_usize.

The script now lints no default features, default features, and every
optional feature but szip (wasm32; the set with zstd, which does not build
for wasm32, on the host, where the lint reports the same casts). With the
old parallel_read.rs/lane_partition.rs it fails listing the four casts; the
old script passed them. CHANGELOG and the design note give the exact count
(119) and what is not covered.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:08:57 -05:00
osobhandClaude Opus 5.5 92c8285549 format: verify B-tree v2 internal node checksums
Only leaves and the header were checked. Harmless while every lookup read
the whole tree, but the indexed lookup prunes children by the keys stored
in internal nodes, so one corrupted byte there could route a name to the
wrong child and report it missing with no error. A BTIN whose lookup3
checksum does not match is now ChecksumMismatch on every read (lookups and
full traversals), as in libhdf5.

Test: one byte of the root BTIN of the 35 001-link h5py group's name index
changed -> lookups, paths and listings through File, MmapFile and LazyFile
all fail with ChecksumMismatch, and h5py refuses both. Before, lookups
returned Ok.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:07:01 -05:00
osobhandClaude Opus 5.5 d2b25f154f format: in-memory fast path in the Storage read helpers
read_exact_at and read_upto (and so Window::read) ask as_contiguous()
first and slice the file directly when the backend holds it in memory:
one dynamic call per structure read instead of two or three (len,
read_at, then len again for errors). Same results and errors.

Provisional (busy machine, not for docs): a listing that walks 400
symbol-table groups through the facade went from about 18% to about 14%
slower than before the Storage conversion; the extra cost is a few tens
of nanoseconds per structure read, which the facade's per-lookup
re-listing (range-reads.md M0) multiplies.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:41:26 -05:00
osobhandClaude Opus 5.5 3c89a31df0 clawhdf5: FileEditor modifies existing files in place
New clawhdf5::FileEditor opens an HDF5 file (h5py-written at any libver,
HDF5 2.0 format included, or clawhdf5-written) under an exclusive flock
and changes only what an edit touches:
- write_selection/write_all/write_values: compact, contiguous (also
  late-allocated) and chunked datasets, any selection. Chunks are decoded,
  updated and re-encoded; a filtered chunk that no longer fits moves to the
  end of the file unless it is the file's last structure, which grows in
  place. New chunks go into v1 B-tree, Extensible Array (paged data blocks
  included), Fixed Array and single-chunk indexes, created on first use.
- resize: grow chunked datasets up to maxshape.
- set_attr: add/replace compact attributes, in a NIL slot or a new
  continuation chunk.
Each edit is planned in an in-memory image and refused whole
(Error::Unsupported) when any part is unsupported (v2 B-tree / implicit
new chunks, shrinking, vlen/reference data, dense or order-tracked
attributes, cache images, paged/persistent free space). Commit writes and
syncs new space before patching existing bytes. Layout v5 (HDF5 2.0)
array indexes use 8-byte filtered chunk sizes, as libhdf5 does.

Error gains Unsupported/InvalidArgument/Locked and is #[non_exhaustive];
the Python bindings map them. build_attr_message is public.

Tests (h5py, h5dump, h5rs check --data after every round; h5py r+
afterwards): appends crossing EA super/data blocks and B-tree splits, the
same B-tree node counts and EA statistics as libhdf5 for the same writes
(in order, reversed and shuffled; paged blocks), every layout and chunk
index overwritten under random selections, attributes to continuation
chunks, random operations against a model, refused edits leave the file
byte-identical, locking.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:35:06 -05:00
osobhandClaude Opus 5.5 b41583113a format: no truncating u64 -> usize casts
Every `u64 as usize` cast in clawhdf5-format (115 on wasm32) now goes
through addr::to_usize for values read from the file — addresses, lengths,
counts, dimensions: FormatError::Overflow where the value does not fit
instead of wrapping onto another part of the file on a 32-bit target — or
addr::saturating_usize for counts bounded by something in memory (codec
progress counters, writer sizes), which fail a bounds check or allocation
rather than wrap. A chunk whose offset does not fit lies outside the
dataset and is skipped; partial reads treat such an offset as out of the
buffers. On 64-bit targets nothing changes.

scripts/check-32bit-casts.sh (run by ci-test.sh) lints the wasm32 build
with clippy's cast_possible_truncation and fails on any u64 -> usize
finding; before this commit it listed 115.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:33:24 -05:00
osobhandClaude Opus 5.5 02e89c1d2d read: look names up through the dense name indexes
Finding one link or attribute by name read every entry: Group::dataset and
Group::group (File, MmapFile, LazyFile) listed the whole group per call, and
path resolution scanned each group's links. Opening every child of a
35 001-link group by name decoded ~1.2e9 links.

Now a dense group's v2 B-tree name index (type 5, lookup3 hash of the
name) is descended to the records with the name's hash
(btree_v2::find_btree_v2_records reads only the nodes whose key interval
overlaps), and only those links are read and compared; all hash-equal
records are compared, so libhdf5's tie order does not matter. Dense
attributes the same through their type 8 index
(attribute::find_attribute_in_file, facade attr(name)); huge heap objects
through their ID-ordered index. group_v2::resolve_child returns what the
listing has under a name (soft links followed, dangling/external ones not
found). Group::entries and File::group_at hand out a listing's addresses.

The lookup-stats feature counts heap objects read. Tests: one lookup in
an h5py-written 35 001-link group with colliding hashes reads at most two
links (before: 35 001, failing), attribute lookups likewise (before: 3 000,
failing), every child opens through all three readers and matches h5py,
every link kind resolves as h5py resolves it in dense and compact groups,
300 huge attributes are found, and a range search matches a full scan at
every tree depth.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:33:19 -05:00
osobhandClaude Opus 5.5 476960f4b8 format: fix a clippy lint in the global heap test fixture
With Storage in scope, `data.len()` on a `&&[u8]` resolves to
Storage::len (already a u64), so `as u64` was a no-op cast; name the
slice method explicitly.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:11:46 -05:00
osobhandClaude Opus 5.5 24f0c71939 format: keep dataset_fill_value_in's &[u8] signature
Making dataset_fill_value_in generic over Storage broke callers that pass
an array (`include_bytes!`): a generic parameter does not unsize-coerce
`&[u8; N]`. It takes &[u8] again, as before this branch, and wraps the
new dataset_fill_value_from_storage(&dyn Storage, ..).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:11:46 -05:00