955fdb660d5121211d06192101d443c12d06fa11
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
304aed5813 |
tests: random editor operations include shrinking, 2-D growth, dense attributes
random_operations_match_a_model (edit_interop) now drives, on every libver (earliest, v114, latest) and filter set (none, gzip + shuffle + fletcher32, LZF): - a dataset with one unlimited dimension and one with two (a version-2 B-tree chunk index under v114/latest), resized to random shapes that shrink and grow any resizable dimension, with block and point writes; - attributes on the first dataset under 20 names with values of random types and sizes (scalars, int64 arrays, short strings, strings above the heap's managed limit), so they move to dense storage on version-2 headers and are replaced by values of other sizes; against a model where shrunk-away elements that come back read as the fill value, compared with our reader and with h5py/numpy every 40 steps, with h5dump and h5rs check; h5py then grows both datasets and adds an attribute. The attribute check also compares h5py's attribute count with libhdf5's object info. CLAWHDF5_EDIT_SEED reruns the workloads with other random choices (seeds 1000-4000 pass). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
dc9cfba6bb |
edit: reuse space freed earlier in the editing session
A FileEditor now keeps the space its edits free — a filtered chunk that moved, chunks a shrink removed, B-tree nodes merged away, a heap's replaced root indirect block or free-space section info, huge objects replaced — and later edits allocate from it (best fit, lowest address among equals, zeroed) before growing the file. An edit never reuses what it frees itself: until it is committed the file still refers to that space. Reused blocks are written in the commit's first phase with the space past the old end of file (nothing on disk refers to them yet), before any existing byte changes, so the crash-safety ordering holds. Space still free when the editor is dropped is leaked, as libhdf5 leaks it without a persistent free-space manager (files that have one, or use paged aggregation, are still refused at open). FileEditor::reusable_bytes reports what is left to reuse. Tests: FreeList merging and best fit, the edit-local rule and the commit split (image unit tests); a shrink followed by regrowth writing the same data reuses every removed chunk and leaves the file size unchanged, while one editor per edit grows the file, h5py/h5dump/h5rs check read both and h5py continues (freed_space_is_reused_within_a_session). measure_append_waste (edit_interop, ignored), same workload, one editor, before -> after (bytes; libhdf5 in brackets), on tank 2026-09-26: gzip chunks 1024, 1000 appends of 100: 307210 -> 306780 (306058); gzip chunks 4096, 2000 appends of 10: 119684 -> 79829 (50292); unfiltered unchanged (no space is freed). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
773f427f16 |
edit: dense attributes, compact-to-dense transition, creation order
FileEditor::set_attr now handles every attribute storage libhdf5 uses for version-2 object headers: - objects that track (and index) attribute creation order: compact attributes carry their creation index in the message header, the Attribute Info message its maximum; - the move to dense storage when an object reaches its compact limit (or an attribute is too large for a header message), as H5O__attr_create does it: a new fractal heap, name index (v2 B-tree type 8) and, when creation order is indexed, creation-order index (type 9); the compact attributes moved over in header message order, their messages freed; - objects already in dense storage (h5py- or clawhdf5-written): new attributes inserted (H5A__dense_insert), an attribute replaced by one of the same encoded size rewritten in its heap object (H5A__dense_write), otherwise removed from both indexes and the heap and inserted anew. edit/fheap.rs follows H5HF: managed objects go to the best-fitting free section of the heap's free-space manager (FSHD/FSSE, kept as libhdf5 keeps it — sorted sections, counts, section info reallocated when its size changes, the manager deleted when empty); otherwise to a new direct block: the root direct block of an empty heap, else the block at the allocation iterator in the root indirect block (created from the root direct block, doubled as needed), with libhdf5's managed-space, allocated space, free space and iterator bookkeeping. Objects above the managed limit are huge objects in their own space, indexed by the huge-object B-tree (type 1). Removed objects return their space merged with adjacent free space. Refused before anything is written: heaps with I/O filters, child indirect blocks, an object larger than the next heap block (libhdf5 skips blocks and records them as free space), free sections other than those inside direct blocks, removing a direct block's last object (libhdf5 frees the block), directly addressed huge objects. Attributes are encoded as libhdf5 does when h5py opens a file r+ (low bound "earliest"): message version 1 (3 for non-ASCII names), simple dataspaces with their maximum dimensions. Header chunks are now visited in libhdf5's order (FIFO), which is also the order attributes move to dense storage in. Tests (edit_coverage_interop): 40 attributes on each of a plain group, a group tracking and indexing creation order, and a dataset (earliest, v110, latest), some above the 4 KiB managed limit, then same-size rewrites: the heap statistics, free-space sections and both index B-trees node for node equal libhdf5's doing the same through h5py; then replacements of other sizes, h5py adds/deletes/rewrites; h5py, h5dump, h5rs check and our reader agree throughout, h5py's attribute count included. clawhdf5-written dense storage (tracked and untracked) is extended the same way; refusals leave the file byte for byte as it was. edit_interop's attribute test now expects dense storage and tracked creation order to work. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
e9c71e5d2e |
edit: version-2 B-tree chunk indexes, shrinking, early allocation
FileEditor can now: - add, move and resize chunks of datasets with two or more unlimited dimensions (version-2 B-tree chunk index, record types 10/11). The new edit/btree2.rs follows libhdf5's H5B2 code: H5B2_update (modify, or insert into a leaf with room, or fall back to H5B2__insert), the preemptive split/redistribute loop with its two retries, split1, split_root (depth growth, node geometry per depth), redistribute2/3, cumulative record counts and the pointer widths H5B2__hdr_init derives, a checksum per node. Removal (H5B2_remove: merge2/3, redistribution, root collapse, the internal-record swap with a leaf) is there too. A dataset without an index yet gets one from the layout message's node size and split/merge percentages. - shrink a chunked dataset along any dimension (resize to a smaller shape), as H5D__set_extent / H5D__chunk_prune_by_extent do: the same chunks visited in the same order; chunks wholly outside the new extent are removed from the index (version-1 B-tree: H5B_remove with its sibling key and link fix-ups and the empty-root case; version-2 B-tree; Fixed/Extensible Array elements reset to the fill element; an implicit index keeps its chunks, as libhdf5 does) and their space noted as free; the part of each partial edge chunk outside the extent is overwritten with the fill value, so elements that come back after a later growth read as fill. - under early allocation, allocate and fill the chunks a growth brings in (H5D__chunk_allocate), which an implicit index needs: libhdf5 refills them, and they may hold the data of chunks pruned earlier. Tests (crates/clawhdf5-tools/tests/edit_coverage_interop.rs): growth in both dimensions of v110/latest files, unfiltered and deflated, gives node-for-node the version-2 B-tree libhdf5 builds (h5py with its chunk cache off, so chunks enter the index in the editor's order), through a depth increase; random chunk order; 60 random shrink/grow/write steps on Extensible Array, version-2 B-tree, Fixed Array, 1-D and implicit datasets (earliest/v110/latest, with and without gzip+shuffle) give the values h5py gets doing the same and the same index shape (version-1 and version-2 B-tree node shapes, Extensible Array statistics); h5py r+ continues on every result; h5dump and h5rs check accept them. edit_interop's version-2 B-tree case now appends instead of expecting a refusal; shrinking is no longer an error in edit_tests. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
b668878129 |
clawhdf5: FileEditor reports filters it cannot run as Error::Unsupported
A dataset whose filter this build cannot encode (scale-offset, N-Bit, SZIP;
a plugin filter the build lacks) failed with Error::Format("unsupported
filter: 6"), although the editor documents every refused edit as
Error::Unsupported, and the Python bindings raised ValueError rather than
NotImplementedError. Every edit now maps FormatError::UnsupportedFilter to
Error::Unsupported; the file is left untouched as before.
Test: edit_interop unencodable_filters_are_unsupported — h5py scale-offset
datasets (integer with chunks, integer never written, float D-scale):
Error::Unsupported naming the filter, and the file byte for byte unchanged.
Fails on the previous editor (Format(UnsupportedFilter(6))).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
||
|
|
485bea0f4f |
clawhdf5: set_attr adds the Attribute Info message a version-2 header needs
libhdf5 counts a version-2 object header's attributes through its Attribute Info message (0x15) and reports none when the header has none. set_attr gave v110/latest groups, the root group and datasets without attributes an attribute message only, so h5py listed the attribute but len(obj.attrs) and H5Oget_info's num_attrs said 0, and stayed wrong after h5py r+ added more. Like H5O__attr_create, the edit now adds the message when a version-2 header lacks it, in the same planned edit: version 0, the header's creation-order track/index flags, maximum creation index 0, undefined fractal heap and B-tree addresses, message flag DONTSHARE — byte for byte what libhdf5 writes. It goes before the attribute (libhdf5's order) when free space holds both, else after it, so a continuation chunk made for the attribute also takes it. Test: edit_interop attribute_count_in_version_2_headers — v110 and latest files, attributes set on the root group, groups and datasets with and without existing attributes: h5py's len/num_attrs/list/values, h5dump -A and our reader agree, also after h5py r+ adds attributes up to and past the compact limit. Fails on the previous editor (h5py len 0). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
f7e2ab12f2 |
clawhdf5: FileEditor skips optional filters that fail, as libhdf5 does
The editor stored every chunk through the whole pipeline with filter mask 0. For LZF that did not shrink a chunk, h5py instead stores it raw with the filter's mask bit set. A chunk the editor stored LZF-encoded at exactly the raw size was then rewritten raw by libhdf5 at the same size; libhdf5 does not touch the index entry when the size is unchanged, so the stale mask 0 stayed and h5py (and h5dump) could no longer read the dataset. clawhdf5_format::filters::compress_chunk_masked runs the pipeline as H5Z_pipeline does: an optional filter (H5Z_FLAG_OPTIONAL) that fails is skipped and its bit set, a mandatory one fails the write, and LZF/Blosc output no smaller than the input counts as failure, as in the reference filters (their output buffer is the input's size). Deflate, LZ4, Zstd, bitshuffle and bzip2 never fail on size in libhdf5 and are kept as before. Test: edit_interop optional_filters_that_fail_are_skipped — the reviewer's repro at every libver: the editor stores the chunk exactly as h5py does (mask 1, size 5; shuffle+LZF+fletcher32 mask 2), h5py r+ rewrites and extends the datasets, and h5py, h5dump and our reader read every value. Fails on the previous editor (mask 0; h5dump cannot read /u8). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
3c89a31df0 |
clawhdf5: FileEditor modifies existing files in place
New clawhdf5::FileEditor opens an HDF5 file (h5py-written at any libver, HDF5 2.0 format included, or clawhdf5-written) under an exclusive flock and changes only what an edit touches: - write_selection/write_all/write_values: compact, contiguous (also late-allocated) and chunked datasets, any selection. Chunks are decoded, updated and re-encoded; a filtered chunk that no longer fits moves to the end of the file unless it is the file's last structure, which grows in place. New chunks go into v1 B-tree, Extensible Array (paged data blocks included), Fixed Array and single-chunk indexes, created on first use. - resize: grow chunked datasets up to maxshape. - set_attr: add/replace compact attributes, in a NIL slot or a new continuation chunk. Each edit is planned in an in-memory image and refused whole (Error::Unsupported) when any part is unsupported (v2 B-tree / implicit new chunks, shrinking, vlen/reference data, dense or order-tracked attributes, cache images, paged/persistent free space). Commit writes and syncs new space before patching existing bytes. Layout v5 (HDF5 2.0) array indexes use 8-byte filtered chunk sizes, as libhdf5 does. Error gains Unsupported/InvalidArgument/Locked and is #[non_exhaustive]; the Python bindings map them. build_attr_message is public. Tests (h5py, h5dump, h5rs check --data after every round; h5py r+ afterwards): appends crossing EA super/data blocks and B-tree splits, the same B-tree node counts and EA statistics as libhdf5 for the same writes (in order, reversed and shuffled; paged blocks), every layout and chunk index overwritten under random selections, attributes to continuation chunks, random operations against a model, refused edits leave the file byte-identical, locking. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |