Range reads M0/M1 (indexed lookups, Storage trait), ZFP, in-place editing #17
@@ -1062,6 +1062,37 @@ and fails their objects (see below).
|
||||
- CI keeps zlib-ng building and tested; the arm64 job no longer needs cmake.
|
||||
|
||||
### Correctness
|
||||
- **`FileBuilder` stored LZF and Blosc chunks with filter mask 0 even when
|
||||
the filter had not shrunk them** (fixed 2026-09-26). Latent in the
|
||||
unreleased LZF/Blosc writer only (added 2026-09-26, "Plugin filters"):
|
||||
no tagged release writes LZF or Blosc, so v2.7.0 and earlier are
|
||||
unaffected. libhdf5 treats LZF and Blosc output no smaller than the chunk
|
||||
as a filter failure and, both being optional filters, stores such a chunk
|
||||
unfiltered with the filter's mask bit set. clawhdf5 stored the filter's
|
||||
output with a clear mask. For an LZF chunk whose stream was exactly the
|
||||
chunk's size (h5py stores `[182, 0, 0, 0, 0]` in a 5-byte chunk raw),
|
||||
the first libhdf5 rewrite of that chunk stored the new data raw at the
|
||||
same size and, the size being unchanged, kept the stale mask 0: h5py
|
||||
then failed to read the dataset ("filter returned failure during read").
|
||||
The whole-file writer now runs chunks through the pipeline as libhdf5
|
||||
does (`clawhdf5_format::filters::compress_chunk_masked`, as `FileEditor`
|
||||
already did) and records each chunk's real mask in every chunk index it
|
||||
builds (single chunk, Fixed Array, Extensible Array, version-2 B-tree;
|
||||
it builds no version-1 B-tree or implicit index), in the sequential and
|
||||
`parallel` paths and `create_datasets_parallel`. Files whose chunks all
|
||||
compress are byte-identical to before. `PrecompressedChunks::chunks` is
|
||||
now `(raw size, stored bytes, filter mask)` (**breaking** for code that
|
||||
reads it). Files written before the fix read correctly; rewrite them
|
||||
(with this build or `h5repack`) before modifying them with libhdf5.
|
||||
Tests: `plugin_filters_interop`
|
||||
`skipped_optional_filters_are_masked_as_libhdf5_masks_them` (LZF,
|
||||
shuffle+LZF+fletcher32 and Blosc, random, compressible and alternating
|
||||
chunks, every index: masks equal an h5py-written twin's; after h5py r+
|
||||
rewrites and extends the datasets, h5py, h5dump and our reader read every
|
||||
value — before the fix 20 of 24 datasets had other masks than h5py's,
|
||||
and h5py could not read the rewritten `[x, 0, 0, 0, 0]` datasets) and
|
||||
`files_whose_chunks_all_compress_are_unchanged`; `chunked_write`
|
||||
`skipped_lzf_chunks_are_masked_in_every_index`.
|
||||
- **Scale-offset data read wrong values in every release that decoded it
|
||||
(v2.2.0 to v2.7.0), silently, on ordinary h5py files** (fixed
|
||||
2026-09-26). Of 1480 scale-offset datasets h5py writes across every
|
||||
|
||||
@@ -17,7 +17,7 @@ use crate::filter_pipeline::{
|
||||
FILTER_LZF, FILTER_PCODEC, FILTER_PCODEC_NAME, FILTER_SHUFFLE, FILTER_ZSTD, FilterDescription,
|
||||
FilterPipeline,
|
||||
};
|
||||
use crate::filters::compress_chunk;
|
||||
use crate::filters::compress_chunk_masked;
|
||||
/// Round a file offset up to the next cache-line boundary.
|
||||
///
|
||||
/// This ensures chunk data starts at an address that is a multiple of the
|
||||
@@ -489,7 +489,12 @@ pub fn split_into_chunks(
|
||||
#[cfg(feature = "parallel")]
|
||||
const PARALLEL_COMPRESS_THRESHOLD: usize = 2;
|
||||
|
||||
/// Compress all chunks, using parallel compression when beneficial.
|
||||
/// Compress all chunks, using parallel compression when beneficial, and
|
||||
/// return each chunk's stored bytes with its filter mask.
|
||||
///
|
||||
/// Chunks run through the pipeline as libhdf5 runs them
|
||||
/// ([`compress_chunk_masked`]): an optional filter that fails — LZF or Blosc
|
||||
/// output no smaller than its input — is skipped and its mask bit set.
|
||||
///
|
||||
/// With the `parallel` feature and more than [`PARALLEL_COMPRESS_THRESHOLD`]
|
||||
/// filtered chunks, compression runs across rayon threads; otherwise it is
|
||||
@@ -499,7 +504,7 @@ fn compress_all_chunks(
|
||||
chunks: &[(Vec<u64>, Vec<u8>)],
|
||||
pipeline: &Option<FilterPipeline>,
|
||||
element_size: u32,
|
||||
) -> Result<Vec<Vec<u8>>, FormatError> {
|
||||
) -> Result<Vec<(Vec<u8>, u32)>, FormatError> {
|
||||
#[cfg(feature = "parallel")]
|
||||
{
|
||||
if let Some(pl) = pipeline
|
||||
@@ -508,7 +513,7 @@ fn compress_all_chunks(
|
||||
use rayon::prelude::*;
|
||||
return chunks
|
||||
.par_iter()
|
||||
.map(|(_offsets, chunk_bytes)| compress_chunk(chunk_bytes, pl, element_size))
|
||||
.map(|(_offsets, chunk_bytes)| compress_chunk_masked(chunk_bytes, pl, element_size))
|
||||
.collect();
|
||||
}
|
||||
}
|
||||
@@ -518,9 +523,9 @@ fn compress_all_chunks(
|
||||
.iter()
|
||||
.map(|(_offsets, chunk_bytes)| {
|
||||
if let Some(pl) = pipeline {
|
||||
compress_chunk(chunk_bytes, pl, element_size)
|
||||
compress_chunk_masked(chunk_bytes, pl, element_size)
|
||||
} else {
|
||||
Ok(chunk_bytes.clone())
|
||||
Ok((chunk_bytes.clone(), 0))
|
||||
}
|
||||
})
|
||||
.collect()
|
||||
@@ -798,8 +803,10 @@ pub fn build_fixed_array_at(
|
||||
/// writer passes eliminates the double-compression that the two-pass layout
|
||||
/// algorithm previously performed.
|
||||
pub struct PrecompressedChunks {
|
||||
/// Per-chunk: (raw_size_bytes, compressed_bytes).
|
||||
pub chunks: Vec<(u64, Vec<u8>)>,
|
||||
/// Per-chunk: (raw_size_bytes, stored_bytes, filter_mask). Bit `i` of
|
||||
/// the mask is set when filter `i` was skipped (an optional filter that
|
||||
/// failed); 0 for every chunk of an unfiltered dataset.
|
||||
pub chunks: Vec<(u64, Vec<u8>, u32)>,
|
||||
pub has_filters: bool,
|
||||
pub element_size: usize,
|
||||
pub shape: Vec<u64>,
|
||||
@@ -834,7 +841,7 @@ pub fn precompress_chunks(
|
||||
let chunks = raw_chunks
|
||||
.into_iter()
|
||||
.zip(compressed)
|
||||
.map(|((_offsets, raw_bytes), c)| (raw_bytes.len() as u64, c))
|
||||
.map(|((_offsets, raw_bytes), (c, mask))| (raw_bytes.len() as u64, c, mask))
|
||||
.collect();
|
||||
|
||||
Ok(PrecompressedChunks {
|
||||
@@ -867,7 +874,7 @@ pub fn build_chunked_data_from_precompressed(
|
||||
let mut data_buf = Vec::new();
|
||||
let mut written_chunks = Vec::with_capacity(num_chunks);
|
||||
|
||||
for (raw_size, compressed) in &pre.chunks {
|
||||
for (raw_size, compressed, filter_mask) in &pre.chunks {
|
||||
let aligned_offset = align_to_cache_line(data_buf.len());
|
||||
if aligned_offset > data_buf.len() {
|
||||
data_buf.resize(aligned_offset, 0u8);
|
||||
@@ -879,7 +886,7 @@ pub fn build_chunked_data_from_precompressed(
|
||||
address,
|
||||
compressed_size,
|
||||
raw_size: *raw_size,
|
||||
filter_mask: 0,
|
||||
filter_mask: *filter_mask,
|
||||
});
|
||||
}
|
||||
|
||||
@@ -916,7 +923,7 @@ pub fn build_chunked_data_from_precompressed(
|
||||
} else {
|
||||
None
|
||||
};
|
||||
let filter_mask = if pre.has_filters { Some(0u32) } else { None };
|
||||
let filter_mask = pre.has_filters.then_some(written_chunks[0].filter_mask);
|
||||
serialize_v4_single_chunk(
|
||||
&chunk_dims_u32,
|
||||
chunk_addr,
|
||||
@@ -1943,6 +1950,98 @@ mod tests {
|
||||
bytes_to_f64(&output)
|
||||
}
|
||||
|
||||
/// Every chunk index the writer builds records each chunk's real filter
|
||||
/// mask: LZF output no smaller than the chunk is skipped (bit 1, behind
|
||||
/// shuffle) and the chunk stored shuffled only; compressible chunks keep
|
||||
/// mask 0. The data reads back through both kinds of chunk.
|
||||
#[cfg(feature = "lzf")]
|
||||
#[test]
|
||||
fn skipped_lzf_chunks_are_masked_in_every_index() {
|
||||
let c = 64usize;
|
||||
// Chunks alternate: random bytes (LZF cannot shrink them), then 7s.
|
||||
let mut state = 0x1234_5678_u64;
|
||||
let data: Vec<f64> = (0..4 * c)
|
||||
.map(|i| {
|
||||
if (i / c).is_multiple_of(2) {
|
||||
state ^= state << 13;
|
||||
state ^= state >> 7;
|
||||
state ^= state << 17;
|
||||
f64::from_bits(state)
|
||||
} else {
|
||||
7.0
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
let raw = f64_to_bytes(&data);
|
||||
let options = ChunkOptions {
|
||||
plugin: Some(PluginFilter::Lzf),
|
||||
..Default::default()
|
||||
};
|
||||
let c64 = c as u64;
|
||||
#[allow(clippy::type_complexity)]
|
||||
let cases: [(&[u64], &[u64], Option<&[u64]>, u8, &[u32]); 4] = [
|
||||
(&[c64], &[c64], None, 1, &[2]),
|
||||
(&[4 * c64], &[c64], None, 3, &[2, 0, 2, 0]),
|
||||
(&[4 * c64], &[c64], Some(&[u64::MAX]), 4, &[2, 0, 2, 0]),
|
||||
(
|
||||
&[2, 2 * c64],
|
||||
&[1, c64],
|
||||
Some(&[u64::MAX, u64::MAX]),
|
||||
5,
|
||||
&[2, 0, 2, 0],
|
||||
),
|
||||
];
|
||||
let base = 0x1000u64;
|
||||
for (shape, chunks, maxshape, index_type, want_masks) in cases {
|
||||
let n: u64 = shape.iter().product();
|
||||
let raw = &raw[..n as usize * 8];
|
||||
let result =
|
||||
build_chunked_data_at_ext(raw, shape, chunks, 8, &options, base, maxshape).unwrap();
|
||||
let mut file = vec![0u8; base as usize];
|
||||
file.extend_from_slice(&result.data_bytes);
|
||||
let layout = DataLayout::parse(&result.layout_message, 8, 8).unwrap();
|
||||
assert!(
|
||||
matches!(&layout, DataLayout::Chunked { chunk_index_type, .. }
|
||||
if *chunk_index_type == Some(index_type)),
|
||||
"{layout:?}"
|
||||
);
|
||||
let dataspace = Dataspace {
|
||||
space_type: DataspaceType::Simple,
|
||||
rank: shape.len() as u8,
|
||||
dimensions: shape.to_vec(),
|
||||
max_dimensions: maxshape.map(<[u64]>::to_vec),
|
||||
};
|
||||
let (mut infos, _) =
|
||||
crate::chunked_read::list_chunks(&file, &layout, &dataspace, 8, 8, 8).unwrap();
|
||||
infos.sort_by(|a, b| a.offsets.cmp(&b.offsets));
|
||||
let masks: Vec<u32> = infos.iter().map(|i| i.filter_mask).collect();
|
||||
assert_eq!(masks, want_masks, "index type {index_type}");
|
||||
for info in &infos {
|
||||
// Skipped chunks are stored at the chunk's size (shuffled).
|
||||
assert_eq!(
|
||||
info.chunk_size == (c * 8) as u32,
|
||||
info.filter_mask != 0,
|
||||
"{info:?}"
|
||||
);
|
||||
}
|
||||
let pipeline = crate::filter_pipeline::FilterPipeline::parse(
|
||||
result.pipeline_message.as_ref().unwrap(),
|
||||
)
|
||||
.unwrap();
|
||||
let out = read_chunked_data(
|
||||
&file,
|
||||
&layout,
|
||||
&dataspace,
|
||||
&make_f64_type(),
|
||||
Some(&pipeline),
|
||||
8,
|
||||
8,
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(out, raw, "index type {index_type}");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ea_roundtrip_1d_inline_only() {
|
||||
let values: Vec<f64> = (0..10).map(|i| i as f64).collect();
|
||||
|
||||
@@ -586,3 +586,457 @@ with h5py.File(sys.argv[1], 'w') as f:
|
||||
let want: Vec<i32> = (0..16).collect();
|
||||
assert_eq!(file.dataset("lzf_ok").unwrap().read_i32().unwrap(), want);
|
||||
}
|
||||
|
||||
/// A family of datasets for the filter-mask tests: element type, chunk
|
||||
/// length along the last dimension, the h5py `create_dataset` keywords of
|
||||
/// the same filters, and how chunk `k` is filled.
|
||||
#[cfg(feature = "lzf")]
|
||||
struct MaskFamily {
|
||||
name: &'static str,
|
||||
/// 1 (`u1`) or 4 (`<i4`).
|
||||
elem: usize,
|
||||
chunk: u64,
|
||||
h5py_kw: &'static str,
|
||||
build: fn(&mut clawhdf5_format::type_builders::DatasetBuilder),
|
||||
fill: MaskFill,
|
||||
}
|
||||
|
||||
#[cfg(feature = "lzf")]
|
||||
#[derive(Clone, Copy)]
|
||||
enum MaskFill {
|
||||
/// `[x, 0, 0, 0, 0]`: LZF's output is exactly the chunk's size, so
|
||||
/// libhdf5's LZF filter fails and h5py stores the chunk raw.
|
||||
FiveBytes,
|
||||
/// Even chunks random bytes, odd chunks one repeated value.
|
||||
Alternating,
|
||||
/// Every chunk one repeated value.
|
||||
Compressible,
|
||||
}
|
||||
|
||||
/// The index shapes of the mask tests: (label, shape, chunk dims, maxshape
|
||||
/// with `u64::MAX` unlimited) for chunk length `c`. Single chunk, Fixed
|
||||
/// Array, Extensible Array and version-2 B-tree, in that order — every
|
||||
/// index the whole-file writer builds.
|
||||
#[cfg(feature = "lzf")]
|
||||
#[allow(clippy::type_complexity)]
|
||||
fn mask_layouts(c: u64) -> Vec<(&'static str, Vec<u64>, Vec<u64>, Option<Vec<u64>>)> {
|
||||
vec![
|
||||
("single", vec![c], vec![c], None),
|
||||
("fixed", vec![4 * c], vec![c], None),
|
||||
("ea", vec![4 * c], vec![c], Some(vec![u64::MAX])),
|
||||
(
|
||||
"bt2",
|
||||
vec![2, 2 * c],
|
||||
vec![1, c],
|
||||
Some(vec![u64::MAX, u64::MAX]),
|
||||
),
|
||||
]
|
||||
}
|
||||
|
||||
/// Raw little-endian bytes of the dataset `fam` fills over `shape`.
|
||||
#[cfg(feature = "lzf")]
|
||||
fn mask_data(fam: &MaskFamily, shape: &[u64], seed: u64) -> Vec<u8> {
|
||||
let c = fam.chunk as usize;
|
||||
let cols = *shape.last().unwrap() as usize;
|
||||
let n: usize = shape.iter().product::<u64>() as usize;
|
||||
let chunks_per_row = cols.div_ceil(c);
|
||||
let mut state = seed;
|
||||
let mut noise = move || {
|
||||
state = state.wrapping_add(0x9E37_79B9_7F4A_7C15);
|
||||
let mut z = state;
|
||||
z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
|
||||
z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
|
||||
z ^ (z >> 31)
|
||||
};
|
||||
let mut out = Vec::with_capacity(n * fam.elem);
|
||||
for i in 0..n {
|
||||
let (row, col) = (i / cols, i % cols);
|
||||
let k = row * chunks_per_row + col / c;
|
||||
let v: u64 = match fam.fill {
|
||||
MaskFill::FiveBytes if col % c == 0 => 182 - k as u64,
|
||||
MaskFill::FiveBytes => 0,
|
||||
MaskFill::Alternating if k.is_multiple_of(2) => noise(),
|
||||
MaskFill::Alternating | MaskFill::Compressible => 7,
|
||||
};
|
||||
out.extend_from_slice(&v.to_le_bytes()[..fam.elem]);
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
#[cfg(feature = "lzf")]
|
||||
fn mask_families() -> Vec<MaskFamily> {
|
||||
#[cfg_attr(not(feature = "blosc"), allow(unused_mut))]
|
||||
let mut v = vec![
|
||||
MaskFamily {
|
||||
name: "lzf5",
|
||||
elem: 1,
|
||||
chunk: 5,
|
||||
h5py_kw: "dict(compression='lzf')",
|
||||
build: |d| {
|
||||
d.with_lzf().without_shuffle();
|
||||
},
|
||||
fill: MaskFill::FiveBytes,
|
||||
},
|
||||
MaskFamily {
|
||||
name: "lzf",
|
||||
elem: 4,
|
||||
chunk: 64,
|
||||
h5py_kw: "dict(compression='lzf')",
|
||||
build: |d| {
|
||||
d.with_lzf().without_shuffle();
|
||||
},
|
||||
fill: MaskFill::Alternating,
|
||||
},
|
||||
MaskFamily {
|
||||
name: "mix",
|
||||
elem: 4,
|
||||
chunk: 8,
|
||||
h5py_kw: "dict(compression='lzf', shuffle=True, fletcher32=True)",
|
||||
build: |d| {
|
||||
d.with_lzf().with_shuffle().with_fletcher32();
|
||||
},
|
||||
fill: MaskFill::Alternating,
|
||||
},
|
||||
MaskFamily {
|
||||
name: "lzfc",
|
||||
elem: 4,
|
||||
chunk: 64,
|
||||
h5py_kw: "dict(compression='lzf')",
|
||||
build: |d| {
|
||||
d.with_lzf().without_shuffle();
|
||||
},
|
||||
fill: MaskFill::Compressible,
|
||||
},
|
||||
];
|
||||
#[cfg(feature = "blosc")]
|
||||
{
|
||||
use clawhdf5_format::chunked_write::{BloscCodec, BloscShuffle};
|
||||
v.push(MaskFamily {
|
||||
name: "blosc",
|
||||
elem: 4,
|
||||
chunk: 64,
|
||||
h5py_kw: "hdf5plugin.Blosc(cname='lz4', clevel=5, shuffle=hdf5plugin.Blosc.SHUFFLE)",
|
||||
build: |d| {
|
||||
d.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
|
||||
},
|
||||
fill: MaskFill::Alternating,
|
||||
});
|
||||
v.push(MaskFamily {
|
||||
name: "blosc0",
|
||||
elem: 4,
|
||||
chunk: 64,
|
||||
h5py_kw: "hdf5plugin.Blosc(cname='lz4', clevel=0, shuffle=hdf5plugin.Blosc.SHUFFLE)",
|
||||
build: |d| {
|
||||
d.with_blosc(BloscCodec::Lz4, 0, BloscShuffle::Byte);
|
||||
},
|
||||
fill: MaskFill::Compressible,
|
||||
});
|
||||
}
|
||||
v
|
||||
}
|
||||
|
||||
/// Builds h5py twins of our datasets and prints, per dataset, the filter
|
||||
/// masks by chunk offset in our file and in the twin.
|
||||
const MASK_TWIN: &str = r#"
|
||||
import sys, numpy as np, h5py
|
||||
try:
|
||||
import hdf5plugin
|
||||
except ImportError:
|
||||
hdf5plugin = None
|
||||
ours, twin, spec = sys.argv[1], sys.argv[2], eval(sys.argv[3])
|
||||
def masks(ds):
|
||||
return sorted((tuple(ds.id.get_chunk_info(k).chunk_offset), ds.id.get_chunk_info(k).filter_mask)
|
||||
for k in range(ds.id.get_num_chunks()))
|
||||
with h5py.File(ours, 'r') as o, h5py.File(twin, 'w', libver='v114') as t:
|
||||
for name, dt, shape, chunks, maxshape, kw, raw in spec:
|
||||
want = np.fromfile(raw, dtype=dt).reshape(shape)
|
||||
assert np.array_equal(o[name][()], want), name
|
||||
t.create_dataset(name, data=want, chunks=chunks, maxshape=maxshape, **eval(kw))
|
||||
print(name, masks(o[name]), '|', masks(t[name]))
|
||||
"#;
|
||||
|
||||
/// h5py (libhdf5) rewrites every chunk of our datasets — random chunks
|
||||
/// become compressible and the other way round; the `[x,0,0,0,0]` chunks
|
||||
/// change in place at the same size — then extends the resizable ones with
|
||||
/// random data, and saves what each dataset must now hold. Prints the
|
||||
/// datasets h5dump can decode (no chunk left LZF-encoded: h5dump has no
|
||||
/// LZF filter).
|
||||
const MASK_REWRITE: &str = r#"
|
||||
import sys, numpy as np, h5py
|
||||
try:
|
||||
import hdf5plugin
|
||||
except ImportError:
|
||||
hdf5plugin = None
|
||||
ours, spec = sys.argv[1], eval(sys.argv[2])
|
||||
rng = np.random.default_rng(3)
|
||||
def noise(shape, dt):
|
||||
return rng.integers(0, 256, int(np.prod(shape)) * np.dtype(dt).itemsize,
|
||||
dtype=np.uint8).view(dt).reshape(shape)
|
||||
dumpable = []
|
||||
with h5py.File(ours, 'r+') as f:
|
||||
for name, dt, shape, chunks, maxshape, kw, raw in spec:
|
||||
d = f[name]
|
||||
want = d[()]
|
||||
for s in d.iter_chunks():
|
||||
blk = want[s]
|
||||
if dt == 'u1':
|
||||
blk.flat[-1] = 1
|
||||
elif (blk == blk.flat[0]).all():
|
||||
blk[...] = noise(blk.shape, dt)
|
||||
else:
|
||||
blk[...] = 5
|
||||
d[...] = want
|
||||
if maxshape is not None:
|
||||
new = tuple(n + c for n, c in zip(shape, chunks))
|
||||
grown = noise(new, dt)
|
||||
grown[tuple(slice(0, n) for n in shape)] = want
|
||||
d.resize(new)
|
||||
d[...] = grown
|
||||
want = grown
|
||||
want.tofile(raw + '.want')
|
||||
pl = d.id.get_create_plist()
|
||||
ids = [pl.get_filter(i)[0] for i in range(pl.get_nfilters())]
|
||||
if 32000 in ids:
|
||||
bit = 1 << ids.index(32000)
|
||||
if not all(d.id.get_chunk_info(k).filter_mask & bit for k in range(d.id.get_num_chunks())):
|
||||
continue
|
||||
dumpable.append(name)
|
||||
with h5py.File(ours, 'r') as f:
|
||||
for name, dt, shape, chunks, maxshape, kw, raw in spec:
|
||||
want = np.fromfile(raw + '.want', dtype=dt).reshape(f[name].shape)
|
||||
assert np.array_equal(f[name][()], want), name
|
||||
print(' '.join(dumpable))
|
||||
"#;
|
||||
|
||||
/// Optional filters that fail are skipped in files `FileBuilder` writes,
|
||||
/// exactly as libhdf5 skips them: an LZF or Blosc output no smaller than the
|
||||
/// chunk leaves the chunk stored unfiltered with the filter's mask bit set.
|
||||
/// The writer used to store every chunk filtered with mask 0. For LZF, a
|
||||
/// chunk whose LZF stream is exactly the chunk's size (`[x,0,0,0,0]`) was
|
||||
/// then corrupted by the first libhdf5 rewrite of it: libhdf5 stores the
|
||||
/// new data raw at the same size and, the size being unchanged, leaves the
|
||||
/// stale mask in the index, so h5py could no longer read the dataset.
|
||||
///
|
||||
/// For every family × every chunk index the writer builds (single chunk,
|
||||
/// Fixed Array, Extensible Array, version-2 B-tree): our masks equal those
|
||||
/// of an h5py-written twin of the same data; then h5py r+ rewrites and
|
||||
/// extends the datasets, and h5py, h5dump (where it has the filter) and our
|
||||
/// reader read every value.
|
||||
#[cfg(feature = "lzf")]
|
||||
#[test]
|
||||
fn skipped_optional_filters_are_masked_as_libhdf5_masks_them() {
|
||||
let modules = if cfg!(feature = "blosc") {
|
||||
"h5py, numpy, hdf5plugin"
|
||||
} else {
|
||||
"h5py, numpy"
|
||||
};
|
||||
if !have_python(modules) {
|
||||
return;
|
||||
}
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let ours = dir.path().join("ours.h5");
|
||||
let twin = dir.path().join("twin.h5");
|
||||
let mut fb = clawhdf5::FileBuilder::new();
|
||||
let mut spec = Vec::new();
|
||||
let mut names = Vec::new();
|
||||
for (fi, fam) in mask_families().iter().enumerate() {
|
||||
for (label, shape, chunks, maxshape) in mask_layouts(fam.chunk) {
|
||||
let name = format!("{}_{label}", fam.name);
|
||||
let data = mask_data(fam, &shape, fi as u64 * 31 + shape.len() as u64);
|
||||
let raw = dir.path().join(format!("{name}.raw"));
|
||||
std::fs::write(&raw, &data).unwrap();
|
||||
let ds = fb.create_dataset(&name);
|
||||
if fam.elem == 1 {
|
||||
ds.with_u8_data(&data);
|
||||
} else {
|
||||
let v: Vec<i32> = data
|
||||
.as_chunks::<4>()
|
||||
.0
|
||||
.iter()
|
||||
.map(|&b| i32::from_le_bytes(b))
|
||||
.collect();
|
||||
ds.with_i32_data(&v);
|
||||
}
|
||||
ds.with_shape(&shape).with_chunks(&chunks);
|
||||
if let Some(ms) = &maxshape {
|
||||
ds.with_maxshape(ms);
|
||||
}
|
||||
(fam.build)(ds);
|
||||
let py_tuple = |v: &[u64]| {
|
||||
let items: Vec<String> = v
|
||||
.iter()
|
||||
.map(|&d| {
|
||||
if d == u64::MAX {
|
||||
"None".into()
|
||||
} else {
|
||||
d.to_string()
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
format!("({},)", items.join(","))
|
||||
};
|
||||
spec.push(format!(
|
||||
"({name:?}, {:?}, {}, {}, {}, {:?}, {:?})",
|
||||
if fam.elem == 1 { "u1" } else { "<i4" },
|
||||
py_tuple(&shape),
|
||||
py_tuple(&chunks),
|
||||
maxshape.as_deref().map_or("None".into(), py_tuple),
|
||||
fam.h5py_kw,
|
||||
raw.to_str().unwrap(),
|
||||
));
|
||||
names.push((name, fam.elem));
|
||||
}
|
||||
}
|
||||
fb.write(&ours).unwrap();
|
||||
let spec = format!("[{}]", spec.join(", "));
|
||||
|
||||
// Our masks are h5py's, chunk by chunk.
|
||||
let out = run_python(
|
||||
MASK_TWIN,
|
||||
&[ours.to_str().unwrap(), twin.to_str().unwrap(), &spec],
|
||||
);
|
||||
let mut skipped = 0;
|
||||
for line in out.lines() {
|
||||
let (name, rest) = line.split_once(' ').unwrap();
|
||||
let (got, want) = rest.split_once(" | ").unwrap();
|
||||
assert_eq!(got, want, "{name}: our filter masks (left) vs h5py's");
|
||||
skipped += usize::from(got.contains("), 1)") || got.contains("), 2)"));
|
||||
}
|
||||
assert_eq!(out.lines().count(), names.len());
|
||||
assert!(
|
||||
skipped >= 12,
|
||||
"too few datasets with skipped filters:\n{out}"
|
||||
);
|
||||
|
||||
// libhdf5 rewrites and extends them; everyone reads the new values.
|
||||
let dumpable = run_python(MASK_REWRITE, &[ours.to_str().unwrap(), &spec]);
|
||||
let plugin_path = run_python("import hdf5plugin; print(hdf5plugin.PLUGIN_PATH)", &[]);
|
||||
let file = File::open(&ours).unwrap();
|
||||
for (name, elem) in &names {
|
||||
let want = std::fs::read(dir.path().join(format!("{name}.raw.want"))).unwrap();
|
||||
let got = file
|
||||
.dataset(name)
|
||||
.unwrap()
|
||||
.read_selection(&Selection::All)
|
||||
.unwrap();
|
||||
assert!(got == want, "{name}: our reader after h5py r+");
|
||||
if !dumpable.split(' ').any(|d| d == name) || Command::new("h5dump").output().is_err() {
|
||||
continue;
|
||||
}
|
||||
let o = Command::new("h5dump")
|
||||
.env("HDF5_PLUGIN_PATH", &plugin_path)
|
||||
.args(["-d", name, "-y", "-w", "0", ours.to_str().unwrap()])
|
||||
.output()
|
||||
.unwrap();
|
||||
assert!(
|
||||
o.status.success(),
|
||||
"h5dump -d {name}: {}",
|
||||
String::from_utf8_lossy(&o.stderr)
|
||||
);
|
||||
let s = String::from_utf8_lossy(&o.stdout).into_owned();
|
||||
let vals: Vec<i64> = s
|
||||
.split_once("DATA {")
|
||||
.and_then(|(_, r)| r.split_once('}'))
|
||||
.map(|(d, _)| {
|
||||
d.split(|c: char| c == ',' || c.is_whitespace())
|
||||
.filter(|t| !t.is_empty())
|
||||
.map(|t| t.parse::<i64>().unwrap())
|
||||
.collect()
|
||||
})
|
||||
.unwrap_or_default();
|
||||
let want_vals: Vec<i64> = want
|
||||
.chunks_exact(*elem)
|
||||
.map(|b| {
|
||||
if *elem == 1 {
|
||||
i64::from(b[0])
|
||||
} else {
|
||||
i64::from(i32::from_le_bytes(b.try_into().unwrap()))
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(vals, want_vals, "h5dump -d {name}");
|
||||
}
|
||||
assert!(
|
||||
dumpable.split(' ').any(|d| d.starts_with("lzf5_")),
|
||||
"{dumpable}"
|
||||
);
|
||||
}
|
||||
|
||||
/// Files whose chunks all compress are written exactly as before optional
|
||||
/// filters could be skipped: every mask is 0 and nothing else changed. The
|
||||
/// hashes are of the files the writer produced before that change.
|
||||
#[cfg(feature = "lzf")]
|
||||
#[test]
|
||||
fn files_whose_chunks_all_compress_are_unchanged() {
|
||||
use clawhdf5_format::checksum::jenkins_lookup3;
|
||||
#[allow(clippy::type_complexity)]
|
||||
#[cfg_attr(not(feature = "blosc"), allow(unused_mut))]
|
||||
let mut cases: Vec<(
|
||||
&str,
|
||||
fn(&mut clawhdf5_format::type_builders::DatasetBuilder),
|
||||
(usize, u32),
|
||||
)> = vec![
|
||||
(
|
||||
"lzf_fixed",
|
||||
|d| {
|
||||
d.with_i32_data(&ramp_i32(4000))
|
||||
.with_chunks(&[500])
|
||||
.with_lzf();
|
||||
},
|
||||
(3965, 449169442),
|
||||
),
|
||||
(
|
||||
"lzf_ea_noshuffle",
|
||||
|d| {
|
||||
d.with_i32_data(&ramp_i32(4000))
|
||||
.with_chunks(&[700])
|
||||
.with_maxshape(&[u64::MAX])
|
||||
.with_lzf()
|
||||
.without_shuffle();
|
||||
},
|
||||
(7213, 4277403206),
|
||||
),
|
||||
(
|
||||
"mix_bt2",
|
||||
|d| {
|
||||
d.with_f64_data(&ramp_f64(40 * 60))
|
||||
.with_shape(&[40, 60])
|
||||
.with_chunks(&[16, 16])
|
||||
.with_maxshape(&[u64::MAX, u64::MAX])
|
||||
.with_lzf()
|
||||
.with_fletcher32();
|
||||
},
|
||||
(11495, 3340532700),
|
||||
),
|
||||
(
|
||||
"lzf_single",
|
||||
|d| {
|
||||
d.with_u8_data(&ramp_u8(3000))
|
||||
.with_chunks(&[3000])
|
||||
.with_lzf();
|
||||
},
|
||||
(546, 690805477),
|
||||
),
|
||||
];
|
||||
#[cfg(feature = "blosc")]
|
||||
cases.push((
|
||||
"blosc_fixed",
|
||||
|d| {
|
||||
use clawhdf5_format::chunked_write::{BloscCodec, BloscShuffle};
|
||||
d.with_i32_data(&ramp_i32(5000))
|
||||
.with_chunks(&[1024])
|
||||
.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
|
||||
},
|
||||
(2776, 4278611376),
|
||||
));
|
||||
for (name, build, want) in &cases {
|
||||
let mut fb = clawhdf5::FileBuilder::new();
|
||||
build(fb.create_dataset("d"));
|
||||
let bytes = fb.finish().unwrap();
|
||||
assert_eq!(
|
||||
(bytes.len(), jenkins_lookup3(&bytes)),
|
||||
*want,
|
||||
"{name}: (length, lookup3 hash) of the file"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -7,6 +7,23 @@ deleting it.
|
||||
|
||||
---
|
||||
|
||||
## LZF/Blosc chunks written with a stale filter mask
|
||||
|
||||
**Status:** fixed 2026-09-26, before any release (the LZF and Blosc writers
|
||||
were added the same day; v2.7.0 and earlier write neither).
|
||||
|
||||
`FileBuilder` stored every chunk of an LZF or Blosc dataset through the
|
||||
filter with filter mask 0. libhdf5 counts LZF and Blosc output no smaller
|
||||
than the chunk as a failure of the (optional) filter and stores the chunk
|
||||
raw with the filter's mask bit set. When a chunk's LZF stream was exactly
|
||||
the chunk's size, the first libhdf5 rewrite of it stored raw data at the
|
||||
same size and left our mask 0 in the index, so h5py could no longer read
|
||||
the dataset. `FileEditor` had the same bug, fixed earlier the same day.
|
||||
Both now use `clawhdf5_format::filters::compress_chunk_masked`, and every
|
||||
chunk index the writer builds records the real mask (see `CHANGELOG.md`).
|
||||
Files written before the fix read correctly; rewrite them before letting
|
||||
libhdf5 modify them.
|
||||
|
||||
## In-place modification (`FileEditor`) limits
|
||||
|
||||
**Status:** open (documented 2026-09-26). `clawhdf5::FileEditor` refuses,
|
||||
|
||||
Reference in New Issue
Block a user