Commit Graph
109 Commits
Author SHA1 Message Date
osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00
osobhandClaude Opus 5.5 a48cb9f1a4 docs: known issue — NetCDF-4 unlimited dimensions report size 0
Found while checking the README refresh: clawhdf5-netcdf4 reads an
unlimited dimension's size from its dimension-scale dataset, which
netCDF-C never extends, so a dimension with 2 records reports 0.
Variable shapes and values are right. To be fixed separately.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:15:51 -05:00
osobhandClaude Opus 5.5 24cc9d14d8 docs: known issue for the flapping bad_nbit_parms_walk classification
The committed CONFORMANCE.md counts bad_nbit_parms_walk.h5 as an
our-error because its six h5py reads agreed in that run; a rerun the
same day confirmed the over-read and counted it ref-bug. The ok count
(602 of 697) is the same either way.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:07:20 -05:00
osobhandClaude Opus 5.5 798331ddbc docs: known issues split into open issues and a condensed fixed history
An "Open issues" table at the top links each open entry; fixed entries
move to "Fixed (history)", newest first, keeping the date, PR, affected
releases and what users must do. Open entries re-checked against main
(9b5803f): the remaining audit gaps are gathered into one entry, the
range-read and remote limits no longer contradict themselves (Python and
the browser open URLs; SWMR reading is File::open_swmr), and the
nondeterministic damaged-chunk error, fixed in PR #19, has its own
history entry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:04:38 -05:00
osobh bf5a163dcf Merge branch 'perf/wasm-listing-passes' into feat/listing-header-last-files
# Conflicts:
#	CHANGELOG.md
2026-09-27 23:16:27 -05:00
osobhandClaude Opus 5.5 5b45c60c9c docs: ObjectHeader::parse A/B after inlining the v1 message loop (at or below 8f59b2e)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 23:15:09 -05:00
osobhandClaude Opus 5.5 2e5b059530 docs: fewer round trips for remote files in the browser, counted
CHANGELOG (Unreleased): the v1 B-tree lookup, `Storage::hint`, the walks
that go on past a missing node, and the counts before and after on an
h5py file like the reviewer's (3000 datasets, 198 MB, earliest and
latest libver, 1 MiB and 64 KiB blocks), the corpus read lazily at
512 B and 64 KiB blocks, and the Node/Chromium suite.
known-issues (browser limits, round trips): the new counts, why the
passes cannot go lower (the chain of addresses), why merging nearby
requests does not help such a file, and that a second listing refetches
at 1 MiB blocks when the file's metadata blocks exceed `cacheSize`.
range-reads.md M4 status and the viewer README follow.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:48:47 -05:00
osobhandClaude Opus 5.5 05e1136027 docs: conformance report with the last non-ok files classified (602 of 697 ok, 0 our-error, 0 mismatch)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:40:03 -05:00
osobhandClaude Opus 5.5 5c44630ea2 conformance: compare h5py's big-endian VL values corrected, confirm libhdf5 over-reads per run
The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us, on a heuristic (any big-endian VL mismatch) and
a fixed list.

- ref.py checks that the installed h5py returns big-endian VL elements
  with the file's bytes under a little-endian dtype (writing and reading
  a vlen('>f4') in memory) and, if so, relabels them with the file's byte
  order before hashing, marking the object `ref_fix`. The values are now
  compared: attr_datatypes.hdf5 /@vlen_uint64 and tcomplex_be.h5
  /VariableLengthDatasetFloatComplex are identical to ours (h5dump 1.14.6
  prints the same (1, 2), (3, 4, 5), (42)).
- ref_bugs.py re-reads each object h5py reads only through a libhdf5 bug
  in six processes with different heaps (import order, MALLOC_PERTURB_).
  Values the file determines are the same every time; these three change
  (6, 6 and 3 distinct results), so they are over-read memory, not data
  clawhdf5 could match. compare.py classifies a file `ref-bug` only when
  every difference is such an object confirmed in the same run.
- report.py: the ref-bug class, the evidence table, the corrected
  objects; test_ref.py covers both (run in the nightly job).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:26:32 -05:00
osobhandClaude Opus 5.5 7179006aee docs: local read A/B re-run on an idle machine
CI / test-arm64 (pull_request) Successful in 1m25s
CI / test (pull_request) Successful in 15m57s
The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 19:43:07 -05:00
osobhandClaude Opus 5.5 d239292655 docs: local metadata and data reads after range-read M2/M3, rechecked
8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.

The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 15:35:59 -05:00
osobhandClaude Opus 5.5 175d3a1f50 docs: chunk-cache order fix in the changelog and known issues
CI / test-arm64 (pull_request) Successful in 1m26s
CI / test (pull_request) Successful in 19m25s
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 08:03:24 -05:00
osobh e80b12521e Merge branch 'feat/p3-m4-wasm-lazy' into feat/p3-wasm-swmr-python
# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
2026-09-27 07:59:14 -05:00
osobh 5d42b4241c Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python
# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
2026-09-27 07:59:14 -05:00
osobhandClaude Opus 5.5 7a4acbce12 docs: openUrl hardening after review (limits, listing passes, CORS tests)
CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB
decode limit, the 4 GiB file limit on wasm32, bodies cut off at their
length, listing passes, the cross-origin tests, and a pre-existing
nondeterministic error choice on cve-2025-2310.h5 that can fail the
native corpus comparison), the viewer README (options, how listing
costs, tests) and the M4 status in docs/design/range-reads.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:54:29 -05:00
osobhandClaude Opus 5.5 b0930708b6 py: boolean-mask keys raise NotImplementedError, not TypeError
h5py supports boolean masks for reads and writes; clawhdf5 supports
neither, so a mask is an unsupported operation (NotImplementedError, as
for every other edit the bindings cannot do), not an invalid key.

Tests: test_unsupported_edits_are_clear_errors (1-D, N-D and per-axis
mask writes, file unchanged) and test_boolean_masks_are_refused (reads).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:48:04 -05:00
osobhandClaude Opus 5.5 39f25e5d4e edit: plan every edit from the file the editor holds, not its path
FileEditor re-opened its path to plan each edit but wrote through the file
it held open, and the Python 'r+' handle re-opened the path after every
edit to read. When the path came to name another file between edits (a
rename or replacement, or a relative path after os.chdir), an edit was laid
out from the other file's metadata and written into the held one,
corrupting it, and later reads came from the other file (the review's
repro: h5py then reports "invalid dataset size, likely file corruption").

The editor now plans from a mapping of its own file (a clone of the held
descriptor, dropped before the edit writes) and canonicalises its path at
open. New FileEditor::reader() opens the held file anew for reading,
without sharing the editor's flock (a mapping of a cloned descriptor holds
the lock until unmapped): through /proc/self/fd on Linux, which follows a
renamed file; elsewhere by path, refused on Unix when the path no longer
names the held file. The Python handle reads through it and keeps no path;
a 'w' file is written at the absolute path it was opened with.

Tests: edit_tests.rs edits_go_to_the_file_held_not_the_path; test_edit.py
test_relative_path_and_chdir and test_path_replaced_between_edits.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:47:03 -05:00
osobhandClaude Opus 5.5 1f7651644b edit: record the maximum before resizing a chunked dataset that has none
FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.

libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).

Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:40:01 -05:00
osobhandClaude Opus 5.5 92fb0830e0 py: in-place editing (clawhdf5.File(path, 'r+')) through FileEditor
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():

- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
  slices and integers with extra leading 1-axes allowed; the exact shape
  for an index list, a scalar only where h5py expands it). Arrays are
  converted as libhdf5 converts them in native byte order (integers
  saturate, floats truncate toward zero and clip, integers go into h5py's
  bool enum by value); other values through
  numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
  dataset is a ValueError instead of libhdf5's arbitrary value. The value
  preparation is a small Python module compiled into the extension
  (src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
  attrs.modify: numeric, bool, complex, bytes and str data of any shape,
  with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
  cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.

Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.

Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.

Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 06:57:36 -05:00
osobhandClaude Opus 5.5 3f45755d58 docs: SWMR reader (range-read M5) in CHANGELOG, README, known issues and designs
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 06:49:04 -05:00
osobhandClaude Opus 5.5 f825a89e23 docs: range-read M4 (openUrl in the browser): changelog, limits, design status
CHANGELOG (Unreleased), known-issues (the browser can open URLs; the
limits of openUrl: round trips per wave of misses, a call holds what it
reads, CORS and validator visibility, the download fallback), the M4
status in docs/design/range-reads.md (why NeedBytes rather than a
Worker, why not clawhdf5-remote's BlockCache, what was tested; the
status paragraph at the top lost a garbled duplicate), the viewer's
README (API, options, how it works, tests; the size table is marked as
predating openUrl), CLAUDE.md and the README crate list.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 06:47:29 -05:00
osobhandClaude Opus 5.5 910d81904c py: remote files (clawhdf5.File(url), File.open_url) through File::storage()
The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.

- Every object of a File now shares one handle (src/handle.rs) that
  runs all file access, metadata included, with the GIL released and
  parses through File::storage() and the clawhdf5_format *_in functions.
  Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
  clawhdf5_remote::storage_for_url (read-only; another mode is a
  ValueError). File.open_url(url, **options) takes the cache and HTTP
  options (block_size, cache_size, headers, retries, timeout,
  allow_full_download, max_full_download, require_validator,
  max_redirects, max_parallel); File.remote_stats gives the block
  cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
  s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
  ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
  OSError, never KeyError/ValueError and never data; `key in group`
  raises it instead of answering False.

Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 06:40:44 -05:00
osobh 7447dce121 Merge branch 'feat/p3-editor-coverage' into feat/p3-remote-editor
# Conflicts:
#	CHANGELOG.md
#	CLAUDE.md
#	docs/design/range-reads.md
2026-09-26 19:17:46 -05:00
osobhandClaude Opus 5.5 8236b0e30a edit: skip heap blocks too small for an attribute, as libhdf5 does
An attribute needing a heap block larger than the next one was refused
("skipping blocks too small for an object", "a first object too large for
the starting block"); once an object's move to dense storage was refused
it refused every new attribute, so 24% of set_attr calls in the review's
random workload failed.

Following H5HF__hdr_update_iter, H5HF__man_iblock_root_create/_double and
H5HF__hdr_skip_blocks, the smaller blocks are now skipped: the iterator
moves past them and they become an indirect free section with a first
row section (serialized, class 1, as H5HF__sect_indirect_serialize writes
it) and ghost normal rows, added as returned space so it merges with a
range skipped just before it (H5HF__sect_indirect_merge_row). Later
objects that best-fit a row section get a block created there
(H5HF__man_iblock_alloc_row / H5HF__sect_indirect_reduce_row: from the
start or end of the range, or from its middle, which splits it, with
libhdf5's span bookkeeping). Heaps with such sections, as libhdf5 writes
them, are now read too (they were refused at open).

dense_skipped_blocks_match_libhdf5 drives every path (merge, split, end,
last entry, row wrap) on earliest/v110/latest files against libhdf5
doing the same edits one session each; heaps, free sections and index
B-trees are equal after every phase. The refusal test now checks the
skip against libhdf5 and keeps a real refusal (last object in a block);
clawhdf5-written heaps get 1-4 KiB attributes too. The three tests fail
on the previous fheap.rs. Random workload refusals: 24% -> 2.2%, all the
documented last-object-in-a-block case.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:50:23 -05:00
osobhandClaude Opus 5.5 c2ae7846c9 docs: label the h5repack sizes in the append-waste figures
306 104 and 49 930 are h5repack of the editor's file; the text read as if
they were h5repack of libhdf5's, which measures 305 954 and 50 188. Both
are now given, from measure_append_waste rerun on 2026-09-26 (file sizes
unchanged).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:50:23 -05:00
osobhandClaude Opus 5.5 159e588550 filters: Fletcher-32 as libhdf5 computes it
Our checksum reduced its sums with `% 65535`; libhdf5's
H5_checksum_fletcher32 folds them with `(s & 0xffff) + (s >> 16)`, which
leaves 0xffff where the modulo leaves 0. On about one chunk in 32768
libhdf5 refused the chunks we wrote and we refused the chunks it wrote.
Every release since v2.1.0 is affected.

clawhdf5_format::checksum::fletcher32 is a port of H5_checksum_fletcher32
and the filter's only implementation. Verification also accepts the
byte-swapped form libhdf5 accepts (1.6.2 and earlier) and the `% 65535`
form earlier releases wrote, so their files stay readable.

The new interop test compares the checksum with libhdf5's own function
(ctypes) on every 1- and 2-byte input and 40 000 random and fold-heavy
inputs, and moves fold-case chunks between h5py and FileBuilder/FileEditor
in both directions; with the old filters.rs the three file tests fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:50:23 -05:00
osobhandClaude Opus 5.5 0e98ffc498 docs: remote files after the adversarial review
CHANGELOG, the clawhdf5-remote and h5rs READMEs and the remote-files
known issues: redirect rules, scaled timeouts (min_speed), URL redaction,
claimed lengths never allocated (download, --max-download), a 200 for a
small file accepted, and ObjectStoreStorage from any thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 18:37:05 -05:00
osobhandClaude Opus 5.5 955dd1c691 docs: range-read milestone M3 — remote files
README: "Reading remote files" (open_url, the range_server and read_url
examples with their real output against the fixtures, h5rs on URLs), the
crate in the crate map and the unreleased highlights. CHANGELOG: the
clawhdf5-remote crate, h5rs URLs, File::storage and
VlResolver::element_in, with the request counts over the conformance
corpus (tank, 2026-09-26, the command given). known-issues: the M2
range-read entry updated (the cache now exists; h5rs reads through
storage) and a new entry for the remote backends' limits (no Python or
browser URLs yet, fixed block size, cloud stores not run against a real
bucket, validators, credentials). Design doc: M3 status with the choices
that differ from the plan (a crate rather than a clawhdf5-io feature,
ureq for HTTP so the default build has no C) and the corpus counts.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 17:29:09 -05:00
osobhandClaude Opus 5.5 0aca0eb724 docs: editor coverage — version-2 B-trees, shrinking, dense attributes, reuse
CHANGELOG (Unreleased): the new FileEditor operations, space reuse, and
the two reader fixes (implicit index grid, object-header continuation
chains). known-issues: the editor's remaining refusals (skipped heap
blocks, heaps with filters or child indirect blocks, freeing a heap
block, implicit-index insertions, ...) and the append-waste sizes before
and after reuse (measure_append_waste, tank 2026-09-26; file sizes are
deterministic). range-reads design: status note on the reader changes.
README and CLAUDE.md: what the editor covers and how to test it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 17:13:31 -05:00
osobhandClaude Opus 5.5 f191dc09d5 docs: range-read milestone M2 — changelog, limits, design status
CHANGELOG (Unreleased): File::open_storage, raw data and v2 B-trees over
Storage, the tests and their corpus results (2026-09-26, tank; conformance
600 of 697, results.json identical to 8f59b2e). known-issues: what
open_storage does not do yet (no remote backend or block cache, read_at
counts of a one-pass read, v1 group lookups, whole-file VDS sources,
zero-copy methods, SWMR growth, hash-order error choice on damaged chunked
datasets). Design: M2 status and the choices made.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 16:49:42 -05:00
osobh 8fadb9f424 Merge branch 'feat/p3-in-place-modify' into feat/p3-range-zfp-edit
# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-py/src/lib.rs
#	crates/clawhdf5/src/error.rs
2026-09-26 14:52:55 -05:00
osobhandClaude Opus 5.5 0e8522cfad clawhdf5-format: the writer skips optional filters that fail, as libhdf5 does
FileBuilder stored every chunk of an LZF or Blosc dataset through the
filter with filter mask 0. libhdf5 counts LZF and Blosc output no smaller
than the chunk as a failure of the optional filter and stores the chunk raw
with the filter's mask bit set. For an LZF chunk whose stream was exactly
the chunk's size, the first libhdf5 rewrite stored raw data at the same
size and kept our stale mask 0 in the index, and h5py could no longer read
the dataset.

precompress_chunks now runs chunks through compress_chunk_masked (as
FileEditor does since f7e2ab1), sequentially and on the parallel path, and
build_chunked_data_from_precompressed records each chunk's real mask in
every index the writer builds: single chunk (layout field), Fixed Array and
Extensible Array filtered elements, and version-2 B-tree type 11 records
(create_datasets_parallel goes through the same path). The writer builds
no version-1 B-tree or implicit index. PrecompressedChunks::chunks gains
the mask. Files whose chunks all compress are byte-identical.

Latent only in the unreleased LZF/Blosc writer (added 2026-09-26); no
tagged release writes either filter.

Tests:
- plugin_filters_interop skipped_optional_filters_are_masked_as_libhdf5_masks_them:
  LZF, shuffle+LZF+fletcher32 and Blosc over random, compressible and
  alternating chunks in every index; masks equal an h5py-written twin's;
  h5py r+ rewrites and extends them; h5py, h5dump and our reader read
  every value. Before: 20 of 24 datasets had masks other than h5py's, and
  with that check disabled h5py failed to read the rewritten datasets
  ("filter returned failure during read").
- plugin_filters_interop files_whose_chunks_all_compress_are_unchanged:
  pins the pre-fix bytes of five all-compressing files.
- chunked_write skipped_lzf_chunks_are_masked_in_every_index (fails before:
  mask 0, want 2).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:45:51 -05:00
osobhandClaude Opus 5.5 b668878129 clawhdf5: FileEditor reports filters it cannot run as Error::Unsupported
A dataset whose filter this build cannot encode (scale-offset, N-Bit, SZIP;
a plugin filter the build lacks) failed with Error::Format("unsupported
filter: 6"), although the editor documents every refused edit as
Error::Unsupported, and the Python bindings raised ValueError rather than
NotImplementedError. Every edit now maps FormatError::UnsupportedFilter to
Error::Unsupported; the file is left untouched as before.

Test: edit_interop unencodable_filters_are_unsupported — h5py scale-offset
datasets (integer with chunks, integer never written, float D-scale):
Error::Unsupported naming the filter, and the file byte for byte unchanged.
Fails on the previous editor (Format(UnsupportedFilter(6))).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 14:13:45 -05:00
osobhandClaude Opus 5.5 677dc5ec7c docs: FileEditor — changelog, limits and leaked space, README example
known-issues records what the editor refuses, that freed space is never
reused (append-workload file sizes measured 2026-09-26 on tank with the
ignored measure_append_waste test; sizes are deterministic), and that there
is no journal.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:35:06 -05:00
osobhandClaude Opus 5.5 2b7065998a docs: ZFP reads (changelog, README feature table, known issues)
ZFP (32013) was the one plugin filter still listed as UnsupportedFilter.
Conformance on tank, `conformance/run.sh --no-fetch` (2026-09-26): 600 of
697 files ok (baseline 599); h5ex_d_zfp.h5 is newly ok.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 13:08:47 -05:00
osobhandClaude Opus 5.5 f2ff2c424f bench: chunked full reads now beat an h5py process pool
CI / test-arm64 (pull_request) Successful in 1m25s
CI / test (pull_request) Successful in 13m0s
Idle-start run on tank at c5334b1 (noisier than the last: compare ratios
within the run). Full reads of deflate data at 16 threads: 4944 MB/s vs
3135 for 16 h5py processes (1.58x; 0.69x-0.76x before). One thread with
the default pool: 6143 MB/s, 15x one h5py call. The concurrent-read
known issue is closed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 12:20:47 -05:00
osobh 9a73299594 Merge branch 'fix/p2b-remaining-conformance' into feat/p2b-scale
# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-format/src/chunked_read.rs
2026-09-26 11:57:35 -05:00
osobh 4b02e7d068 Merge branch 'feat/p2b-blosc2' into feat/p2b-scale
# Conflicts:
#	CHANGELOG.md
2026-09-26 11:57:11 -05:00
osobh d7f07fa5c1 Merge branch 'feat/p2b-writer-btree-internal-nodes' into feat/p2b-scale
# Conflicts:
#	CHANGELOG.md
2026-09-26 11:57:11 -05:00
osobhandClaude Opus 5.5 4c01267b76 test: every scale-offset dataset h5py writes reads as h5py reads it
The scale-offset fix (d110b1d) was covered only by unit vectors from
CVE chunks, and documented as three corner cases. The review found it
is much bigger: of 1480 scale-offset datasets h5py writes (every
integer type i1..u8, f4 and f8, both byte orders, with and without a
fill value, scaleoffset 0..full width), v2.7.0's decoder read 332
differently from h5py: 151 returned wrong values with no error (82
integer datasets with scaleoffset=0 and a wide range, 51 full-width
i4/u4/i8/u8, 18 f4 D-scale datasets with a large range) and 181 failed
to read. The cause in every case is a chunk libhdf5 stores at full
width, whose elements were decoded as offsets from minval.

tests/scaleoffset_interop.rs generates that matrix with h5py at test
time, stores h5py's decoded values uncompressed next to it, and
compares every dataset's bytes. It passes on this branch; with the
filters.rs before d110b1d it reports "332 of 1480 scale-offset datasets
differ from h5py".

CHANGELOG: a Correctness entry stating this was silent wrong data in
every release that decoded scale-offset (v2.2.0 to v2.7.0), replacing
the corner-case wording. docs/known-issues.md: a fixed entry with the
affected cases.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 11:51:30 -05:00
osobhandClaude Opus 5.5 6559a91495 fix: open a file whose cache image cannot load, and fail its objects
For a metadata cache image libhdf5 cannot load, libhdf5 opens the file
and fails the first metadata read (the image loads on the first
H5C_protect after open); h5py reports the error on the root group. The
conformance probe reported it that way, but File::open refused the file,
so the gate counted cve-2025-6269-1..4 and cve-2025-6516 as agreeing
with h5py for behaviour the library did not have.

The library now behaves as the probe reports: File (mmap, buffered and
from_bytes) and MmapFile open the file and every object lookup (dataset,
dataset_at, group, group listings and attributes, VL decoding) fails with
the image's error; LazyFile reads the root group's header at open, so
its open is that first read and fails. Probe and library take the
three-way decision (refuse at open / image loads / image cannot load)
from the same clawhdf5_format::superblock_ext::cache_image_state.

One deliberate difference from libhdf5 remains, documented: after the
failed first read libhdf5 reads the file's own metadata, which the image
was meant to replace and may be stale; here every lookup keeps failing.
File::cache_image_error / MmapFile::cache_image_error expose the error
to code that parses as_bytes() itself; h5rs checks it before reading any
object header (h5rs ls on cve-2025-6269-1 said "invalid object header
version: 0" from the stale bytes).

Test: metadata_cache_image.rs an_image_libhdf5_cannot_load_fails_every_object
(the fixture with its image signature broken; h5py opens that file and
fails the first read with "Bad metadata cache image header signature").
It fails on the previous commit, where File::open refuses the file.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 11:46:28 -05:00
osobhandClaude Opus 5.5 378afa1584 docs: changelog and known issues for the remaining conformance errors
Conformance on tank, conformance/run.sh --no-fetch (2026-09-26): 597 of
697 ok, 6 our-errors (4 corrupt objects HDF5 2.0 reads through a bug, the
Blosc2 and ZFP filters), 2 mismatches (the known h5py big-endian VL bug).
Closes the known-issues entries for metadata cache images,
cve-2024-32624, cve-2020-10810/10812, and unfiltered chunks of the wrong
size; the N-Bit / 64-bit scale-offset entry is recorded as not our bug.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 10:40:18 -05:00
osobhandClaude Opus 5.5 94b6df986c docs: chunked full reads decode in place; small pools no longer block
CHANGELOG entry for the chunked read changes, and the known-issues entry
on concurrent chunked reads updated: both causes it names (per-read page
faults, readers waiting on a small pool) are fixed; the 16-thread
comparison with h5py stays open until re-measured on an idle machine.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 10:36:44 -05:00
osobhandClaude Opus 5.5 56abaec75e docs: Blosc2 reads (read only); ZFP is the one plugin filter left
Record the `blosc2` feature in the changelog, the README's feature table
and the crate table, and mark the Blosc2 half of the known "Filters"
issue fixed (dated, with the conformance run that shows h5ex_d_blosc2
reading). What stays open: ZFP, writing Blosc2, and the Blosc2 features
hdf5plugin never writes (dictionaries, lazy chunks, variable-length
blocks, user-defined codecs and registered filters), which are errors.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 10:22:52 -05:00
osobhandClaude Opus 5.5 193a5f8a82 writer: track attribute creation order with track_order
h5py's track_order=True orders attributes as well as links; the writer
tracked links only. A tracking object's header now sets the attribute
creation order tracked/indexed flags and carries per-message creation
orders, an Attribute Info message holds the next order (inline too),
and dense storage gets a type-9 creation-order index. The file default
applies to datasets, with DatasetBuilder::track_order per dataset; more
than 65 535 attributes on a tracking object is an error (libhdf5's
counter is 2 bytes). The reader lists such attributes in creation
order.

h5py lists them in order (inline, dense, 20 000 on one dataset) and
keeps numbering in r+ mode, including its inline-to-dense move.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 10:17:37 -05:00
osobhandClaude Opus 5.5 d63c76e7ab writer: v2 B-trees with internal nodes (no 65 535-record limit)
Dense link and attribute indexes and the chunk index for several
unlimited dimensions were single leaves, capping them at 65 535
records. btree_v2_write builds trees of any depth, with node capacities
and pointer widths from libhdf5's H5B2__hdr_init arithmetic (now shared
with the reader as btree_v2::node_info) and libhdf5's node sizes (512
dense, 2048 chunks). Indexes that fit the old one-leaf layout are
written byte for byte as before (compared for 10..65 535 links, attrs
and chunks, tracked and filtered).

Tests: 100 000 links (short names; long names with creation order),
70 000 attributes, 200 000 chunks (and 80 000 deflated), read by h5py,
h5dump and clawhdf5 and edited by h5py r+; h5rs check on the same
shapes, asserting depths 2-3.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 10:12:07 -05:00
osobhandClaude Opus 5.5 dda28d6c72 bench: concurrent reads re-measured after the read fixes
CI / test-arm64 (pull_request) Successful in 1m34s
CI / test (pull_request) Successful in 7m30s
Idle tank at 408f69e, h5py re-run in the same session. Contiguous reads
went from 0.25x to 1.44x h5py (full) and 0.12x to 6.3x (256x256
hyperslabs) on one thread; deflate full reads at 8 threads 887 -> 2943
MB/s (h5py processes 3042). Full chunked reads at 16 threads are still
0.69x-0.76x h5py processes; the issue stays open.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 09:28:51 -05:00
osobh 73a01f1256 Merge branch 'feat/p2-python-bindings' into feat/p2-perf-coverage
# Conflicts:
#	CHANGELOG.md
#	README.md
2026-09-26 09:10:57 -05:00
osobh 956e55c76a Merge branch 'feat/p2-writer-groups-links' into feat/p2-perf-coverage
# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-tools/tests/h5rs_interop.rs
2026-09-26 09:10:50 -05:00
osobh 846c35455d Merge branch 'feat/p2-vl-strings' into feat/p2-perf-coverage
# Conflicts:
#	CHANGELOG.md
2026-09-26 09:10:35 -05:00