Range reads M0/M1 (indexed lookups, Storage trait), ZFP, in-place editing #17
@@ -1062,6 +1062,37 @@ and fails their objects (see below).
|
|||||||
- CI keeps zlib-ng building and tested; the arm64 job no longer needs cmake.
|
- CI keeps zlib-ng building and tested; the arm64 job no longer needs cmake.
|
||||||
|
|
||||||
### Correctness
|
### Correctness
|
||||||
|
- **`FileBuilder` stored LZF and Blosc chunks with filter mask 0 even when
|
||||||
|
the filter had not shrunk them** (fixed 2026-09-26). Latent in the
|
||||||
|
unreleased LZF/Blosc writer only (added 2026-09-26, "Plugin filters"):
|
||||||
|
no tagged release writes LZF or Blosc, so v2.7.0 and earlier are
|
||||||
|
unaffected. libhdf5 treats LZF and Blosc output no smaller than the chunk
|
||||||
|
as a filter failure and, both being optional filters, stores such a chunk
|
||||||
|
unfiltered with the filter's mask bit set. clawhdf5 stored the filter's
|
||||||
|
output with a clear mask. For an LZF chunk whose stream was exactly the
|
||||||
|
chunk's size (h5py stores `[182, 0, 0, 0, 0]` in a 5-byte chunk raw),
|
||||||
|
the first libhdf5 rewrite of that chunk stored the new data raw at the
|
||||||
|
same size and, the size being unchanged, kept the stale mask 0: h5py
|
||||||
|
then failed to read the dataset ("filter returned failure during read").
|
||||||
|
The whole-file writer now runs chunks through the pipeline as libhdf5
|
||||||
|
does (`clawhdf5_format::filters::compress_chunk_masked`, as `FileEditor`
|
||||||
|
already did) and records each chunk's real mask in every chunk index it
|
||||||
|
builds (single chunk, Fixed Array, Extensible Array, version-2 B-tree;
|
||||||
|
it builds no version-1 B-tree or implicit index), in the sequential and
|
||||||
|
`parallel` paths and `create_datasets_parallel`. Files whose chunks all
|
||||||
|
compress are byte-identical to before. `PrecompressedChunks::chunks` is
|
||||||
|
now `(raw size, stored bytes, filter mask)` (**breaking** for code that
|
||||||
|
reads it). Files written before the fix read correctly; rewrite them
|
||||||
|
(with this build or `h5repack`) before modifying them with libhdf5.
|
||||||
|
Tests: `plugin_filters_interop`
|
||||||
|
`skipped_optional_filters_are_masked_as_libhdf5_masks_them` (LZF,
|
||||||
|
shuffle+LZF+fletcher32 and Blosc, random, compressible and alternating
|
||||||
|
chunks, every index: masks equal an h5py-written twin's; after h5py r+
|
||||||
|
rewrites and extends the datasets, h5py, h5dump and our reader read every
|
||||||
|
value — before the fix 20 of 24 datasets had other masks than h5py's,
|
||||||
|
and h5py could not read the rewritten `[x, 0, 0, 0, 0]` datasets) and
|
||||||
|
`files_whose_chunks_all_compress_are_unchanged`; `chunked_write`
|
||||||
|
`skipped_lzf_chunks_are_masked_in_every_index`.
|
||||||
- **Scale-offset data read wrong values in every release that decoded it
|
- **Scale-offset data read wrong values in every release that decoded it
|
||||||
(v2.2.0 to v2.7.0), silently, on ordinary h5py files** (fixed
|
(v2.2.0 to v2.7.0), silently, on ordinary h5py files** (fixed
|
||||||
2026-09-26). Of 1480 scale-offset datasets h5py writes across every
|
2026-09-26). Of 1480 scale-offset datasets h5py writes across every
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ use crate::filter_pipeline::{
|
|||||||
FILTER_LZF, FILTER_PCODEC, FILTER_PCODEC_NAME, FILTER_SHUFFLE, FILTER_ZSTD, FilterDescription,
|
FILTER_LZF, FILTER_PCODEC, FILTER_PCODEC_NAME, FILTER_SHUFFLE, FILTER_ZSTD, FilterDescription,
|
||||||
FilterPipeline,
|
FilterPipeline,
|
||||||
};
|
};
|
||||||
use crate::filters::compress_chunk;
|
use crate::filters::compress_chunk_masked;
|
||||||
/// Round a file offset up to the next cache-line boundary.
|
/// Round a file offset up to the next cache-line boundary.
|
||||||
///
|
///
|
||||||
/// This ensures chunk data starts at an address that is a multiple of the
|
/// This ensures chunk data starts at an address that is a multiple of the
|
||||||
@@ -489,7 +489,12 @@ pub fn split_into_chunks(
|
|||||||
#[cfg(feature = "parallel")]
|
#[cfg(feature = "parallel")]
|
||||||
const PARALLEL_COMPRESS_THRESHOLD: usize = 2;
|
const PARALLEL_COMPRESS_THRESHOLD: usize = 2;
|
||||||
|
|
||||||
/// Compress all chunks, using parallel compression when beneficial.
|
/// Compress all chunks, using parallel compression when beneficial, and
|
||||||
|
/// return each chunk's stored bytes with its filter mask.
|
||||||
|
///
|
||||||
|
/// Chunks run through the pipeline as libhdf5 runs them
|
||||||
|
/// ([`compress_chunk_masked`]): an optional filter that fails — LZF or Blosc
|
||||||
|
/// output no smaller than its input — is skipped and its mask bit set.
|
||||||
///
|
///
|
||||||
/// With the `parallel` feature and more than [`PARALLEL_COMPRESS_THRESHOLD`]
|
/// With the `parallel` feature and more than [`PARALLEL_COMPRESS_THRESHOLD`]
|
||||||
/// filtered chunks, compression runs across rayon threads; otherwise it is
|
/// filtered chunks, compression runs across rayon threads; otherwise it is
|
||||||
@@ -499,7 +504,7 @@ fn compress_all_chunks(
|
|||||||
chunks: &[(Vec<u64>, Vec<u8>)],
|
chunks: &[(Vec<u64>, Vec<u8>)],
|
||||||
pipeline: &Option<FilterPipeline>,
|
pipeline: &Option<FilterPipeline>,
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
) -> Result<Vec<Vec<u8>>, FormatError> {
|
) -> Result<Vec<(Vec<u8>, u32)>, FormatError> {
|
||||||
#[cfg(feature = "parallel")]
|
#[cfg(feature = "parallel")]
|
||||||
{
|
{
|
||||||
if let Some(pl) = pipeline
|
if let Some(pl) = pipeline
|
||||||
@@ -508,7 +513,7 @@ fn compress_all_chunks(
|
|||||||
use rayon::prelude::*;
|
use rayon::prelude::*;
|
||||||
return chunks
|
return chunks
|
||||||
.par_iter()
|
.par_iter()
|
||||||
.map(|(_offsets, chunk_bytes)| compress_chunk(chunk_bytes, pl, element_size))
|
.map(|(_offsets, chunk_bytes)| compress_chunk_masked(chunk_bytes, pl, element_size))
|
||||||
.collect();
|
.collect();
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -518,9 +523,9 @@ fn compress_all_chunks(
|
|||||||
.iter()
|
.iter()
|
||||||
.map(|(_offsets, chunk_bytes)| {
|
.map(|(_offsets, chunk_bytes)| {
|
||||||
if let Some(pl) = pipeline {
|
if let Some(pl) = pipeline {
|
||||||
compress_chunk(chunk_bytes, pl, element_size)
|
compress_chunk_masked(chunk_bytes, pl, element_size)
|
||||||
} else {
|
} else {
|
||||||
Ok(chunk_bytes.clone())
|
Ok((chunk_bytes.clone(), 0))
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
.collect()
|
.collect()
|
||||||
@@ -798,8 +803,10 @@ pub fn build_fixed_array_at(
|
|||||||
/// writer passes eliminates the double-compression that the two-pass layout
|
/// writer passes eliminates the double-compression that the two-pass layout
|
||||||
/// algorithm previously performed.
|
/// algorithm previously performed.
|
||||||
pub struct PrecompressedChunks {
|
pub struct PrecompressedChunks {
|
||||||
/// Per-chunk: (raw_size_bytes, compressed_bytes).
|
/// Per-chunk: (raw_size_bytes, stored_bytes, filter_mask). Bit `i` of
|
||||||
pub chunks: Vec<(u64, Vec<u8>)>,
|
/// the mask is set when filter `i` was skipped (an optional filter that
|
||||||
|
/// failed); 0 for every chunk of an unfiltered dataset.
|
||||||
|
pub chunks: Vec<(u64, Vec<u8>, u32)>,
|
||||||
pub has_filters: bool,
|
pub has_filters: bool,
|
||||||
pub element_size: usize,
|
pub element_size: usize,
|
||||||
pub shape: Vec<u64>,
|
pub shape: Vec<u64>,
|
||||||
@@ -834,7 +841,7 @@ pub fn precompress_chunks(
|
|||||||
let chunks = raw_chunks
|
let chunks = raw_chunks
|
||||||
.into_iter()
|
.into_iter()
|
||||||
.zip(compressed)
|
.zip(compressed)
|
||||||
.map(|((_offsets, raw_bytes), c)| (raw_bytes.len() as u64, c))
|
.map(|((_offsets, raw_bytes), (c, mask))| (raw_bytes.len() as u64, c, mask))
|
||||||
.collect();
|
.collect();
|
||||||
|
|
||||||
Ok(PrecompressedChunks {
|
Ok(PrecompressedChunks {
|
||||||
@@ -867,7 +874,7 @@ pub fn build_chunked_data_from_precompressed(
|
|||||||
let mut data_buf = Vec::new();
|
let mut data_buf = Vec::new();
|
||||||
let mut written_chunks = Vec::with_capacity(num_chunks);
|
let mut written_chunks = Vec::with_capacity(num_chunks);
|
||||||
|
|
||||||
for (raw_size, compressed) in &pre.chunks {
|
for (raw_size, compressed, filter_mask) in &pre.chunks {
|
||||||
let aligned_offset = align_to_cache_line(data_buf.len());
|
let aligned_offset = align_to_cache_line(data_buf.len());
|
||||||
if aligned_offset > data_buf.len() {
|
if aligned_offset > data_buf.len() {
|
||||||
data_buf.resize(aligned_offset, 0u8);
|
data_buf.resize(aligned_offset, 0u8);
|
||||||
@@ -879,7 +886,7 @@ pub fn build_chunked_data_from_precompressed(
|
|||||||
address,
|
address,
|
||||||
compressed_size,
|
compressed_size,
|
||||||
raw_size: *raw_size,
|
raw_size: *raw_size,
|
||||||
filter_mask: 0,
|
filter_mask: *filter_mask,
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -916,7 +923,7 @@ pub fn build_chunked_data_from_precompressed(
|
|||||||
} else {
|
} else {
|
||||||
None
|
None
|
||||||
};
|
};
|
||||||
let filter_mask = if pre.has_filters { Some(0u32) } else { None };
|
let filter_mask = pre.has_filters.then_some(written_chunks[0].filter_mask);
|
||||||
serialize_v4_single_chunk(
|
serialize_v4_single_chunk(
|
||||||
&chunk_dims_u32,
|
&chunk_dims_u32,
|
||||||
chunk_addr,
|
chunk_addr,
|
||||||
@@ -1943,6 +1950,98 @@ mod tests {
|
|||||||
bytes_to_f64(&output)
|
bytes_to_f64(&output)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Every chunk index the writer builds records each chunk's real filter
|
||||||
|
/// mask: LZF output no smaller than the chunk is skipped (bit 1, behind
|
||||||
|
/// shuffle) and the chunk stored shuffled only; compressible chunks keep
|
||||||
|
/// mask 0. The data reads back through both kinds of chunk.
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
#[test]
|
||||||
|
fn skipped_lzf_chunks_are_masked_in_every_index() {
|
||||||
|
let c = 64usize;
|
||||||
|
// Chunks alternate: random bytes (LZF cannot shrink them), then 7s.
|
||||||
|
let mut state = 0x1234_5678_u64;
|
||||||
|
let data: Vec<f64> = (0..4 * c)
|
||||||
|
.map(|i| {
|
||||||
|
if (i / c).is_multiple_of(2) {
|
||||||
|
state ^= state << 13;
|
||||||
|
state ^= state >> 7;
|
||||||
|
state ^= state << 17;
|
||||||
|
f64::from_bits(state)
|
||||||
|
} else {
|
||||||
|
7.0
|
||||||
|
}
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
let raw = f64_to_bytes(&data);
|
||||||
|
let options = ChunkOptions {
|
||||||
|
plugin: Some(PluginFilter::Lzf),
|
||||||
|
..Default::default()
|
||||||
|
};
|
||||||
|
let c64 = c as u64;
|
||||||
|
#[allow(clippy::type_complexity)]
|
||||||
|
let cases: [(&[u64], &[u64], Option<&[u64]>, u8, &[u32]); 4] = [
|
||||||
|
(&[c64], &[c64], None, 1, &[2]),
|
||||||
|
(&[4 * c64], &[c64], None, 3, &[2, 0, 2, 0]),
|
||||||
|
(&[4 * c64], &[c64], Some(&[u64::MAX]), 4, &[2, 0, 2, 0]),
|
||||||
|
(
|
||||||
|
&[2, 2 * c64],
|
||||||
|
&[1, c64],
|
||||||
|
Some(&[u64::MAX, u64::MAX]),
|
||||||
|
5,
|
||||||
|
&[2, 0, 2, 0],
|
||||||
|
),
|
||||||
|
];
|
||||||
|
let base = 0x1000u64;
|
||||||
|
for (shape, chunks, maxshape, index_type, want_masks) in cases {
|
||||||
|
let n: u64 = shape.iter().product();
|
||||||
|
let raw = &raw[..n as usize * 8];
|
||||||
|
let result =
|
||||||
|
build_chunked_data_at_ext(raw, shape, chunks, 8, &options, base, maxshape).unwrap();
|
||||||
|
let mut file = vec![0u8; base as usize];
|
||||||
|
file.extend_from_slice(&result.data_bytes);
|
||||||
|
let layout = DataLayout::parse(&result.layout_message, 8, 8).unwrap();
|
||||||
|
assert!(
|
||||||
|
matches!(&layout, DataLayout::Chunked { chunk_index_type, .. }
|
||||||
|
if *chunk_index_type == Some(index_type)),
|
||||||
|
"{layout:?}"
|
||||||
|
);
|
||||||
|
let dataspace = Dataspace {
|
||||||
|
space_type: DataspaceType::Simple,
|
||||||
|
rank: shape.len() as u8,
|
||||||
|
dimensions: shape.to_vec(),
|
||||||
|
max_dimensions: maxshape.map(<[u64]>::to_vec),
|
||||||
|
};
|
||||||
|
let (mut infos, _) =
|
||||||
|
crate::chunked_read::list_chunks(&file, &layout, &dataspace, 8, 8, 8).unwrap();
|
||||||
|
infos.sort_by(|a, b| a.offsets.cmp(&b.offsets));
|
||||||
|
let masks: Vec<u32> = infos.iter().map(|i| i.filter_mask).collect();
|
||||||
|
assert_eq!(masks, want_masks, "index type {index_type}");
|
||||||
|
for info in &infos {
|
||||||
|
// Skipped chunks are stored at the chunk's size (shuffled).
|
||||||
|
assert_eq!(
|
||||||
|
info.chunk_size == (c * 8) as u32,
|
||||||
|
info.filter_mask != 0,
|
||||||
|
"{info:?}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
let pipeline = crate::filter_pipeline::FilterPipeline::parse(
|
||||||
|
result.pipeline_message.as_ref().unwrap(),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
let out = read_chunked_data(
|
||||||
|
&file,
|
||||||
|
&layout,
|
||||||
|
&dataspace,
|
||||||
|
&make_f64_type(),
|
||||||
|
Some(&pipeline),
|
||||||
|
8,
|
||||||
|
8,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(out, raw, "index type {index_type}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn ea_roundtrip_1d_inline_only() {
|
fn ea_roundtrip_1d_inline_only() {
|
||||||
let values: Vec<f64> = (0..10).map(|i| i as f64).collect();
|
let values: Vec<f64> = (0..10).map(|i| i as f64).collect();
|
||||||
|
|||||||
@@ -586,3 +586,457 @@ with h5py.File(sys.argv[1], 'w') as f:
|
|||||||
let want: Vec<i32> = (0..16).collect();
|
let want: Vec<i32> = (0..16).collect();
|
||||||
assert_eq!(file.dataset("lzf_ok").unwrap().read_i32().unwrap(), want);
|
assert_eq!(file.dataset("lzf_ok").unwrap().read_i32().unwrap(), want);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// A family of datasets for the filter-mask tests: element type, chunk
|
||||||
|
/// length along the last dimension, the h5py `create_dataset` keywords of
|
||||||
|
/// the same filters, and how chunk `k` is filled.
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
struct MaskFamily {
|
||||||
|
name: &'static str,
|
||||||
|
/// 1 (`u1`) or 4 (`<i4`).
|
||||||
|
elem: usize,
|
||||||
|
chunk: u64,
|
||||||
|
h5py_kw: &'static str,
|
||||||
|
build: fn(&mut clawhdf5_format::type_builders::DatasetBuilder),
|
||||||
|
fill: MaskFill,
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
#[derive(Clone, Copy)]
|
||||||
|
enum MaskFill {
|
||||||
|
/// `[x, 0, 0, 0, 0]`: LZF's output is exactly the chunk's size, so
|
||||||
|
/// libhdf5's LZF filter fails and h5py stores the chunk raw.
|
||||||
|
FiveBytes,
|
||||||
|
/// Even chunks random bytes, odd chunks one repeated value.
|
||||||
|
Alternating,
|
||||||
|
/// Every chunk one repeated value.
|
||||||
|
Compressible,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The index shapes of the mask tests: (label, shape, chunk dims, maxshape
|
||||||
|
/// with `u64::MAX` unlimited) for chunk length `c`. Single chunk, Fixed
|
||||||
|
/// Array, Extensible Array and version-2 B-tree, in that order — every
|
||||||
|
/// index the whole-file writer builds.
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
#[allow(clippy::type_complexity)]
|
||||||
|
fn mask_layouts(c: u64) -> Vec<(&'static str, Vec<u64>, Vec<u64>, Option<Vec<u64>>)> {
|
||||||
|
vec![
|
||||||
|
("single", vec![c], vec![c], None),
|
||||||
|
("fixed", vec![4 * c], vec![c], None),
|
||||||
|
("ea", vec![4 * c], vec![c], Some(vec![u64::MAX])),
|
||||||
|
(
|
||||||
|
"bt2",
|
||||||
|
vec![2, 2 * c],
|
||||||
|
vec![1, c],
|
||||||
|
Some(vec![u64::MAX, u64::MAX]),
|
||||||
|
),
|
||||||
|
]
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Raw little-endian bytes of the dataset `fam` fills over `shape`.
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
fn mask_data(fam: &MaskFamily, shape: &[u64], seed: u64) -> Vec<u8> {
|
||||||
|
let c = fam.chunk as usize;
|
||||||
|
let cols = *shape.last().unwrap() as usize;
|
||||||
|
let n: usize = shape.iter().product::<u64>() as usize;
|
||||||
|
let chunks_per_row = cols.div_ceil(c);
|
||||||
|
let mut state = seed;
|
||||||
|
let mut noise = move || {
|
||||||
|
state = state.wrapping_add(0x9E37_79B9_7F4A_7C15);
|
||||||
|
let mut z = state;
|
||||||
|
z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
|
||||||
|
z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
|
||||||
|
z ^ (z >> 31)
|
||||||
|
};
|
||||||
|
let mut out = Vec::with_capacity(n * fam.elem);
|
||||||
|
for i in 0..n {
|
||||||
|
let (row, col) = (i / cols, i % cols);
|
||||||
|
let k = row * chunks_per_row + col / c;
|
||||||
|
let v: u64 = match fam.fill {
|
||||||
|
MaskFill::FiveBytes if col % c == 0 => 182 - k as u64,
|
||||||
|
MaskFill::FiveBytes => 0,
|
||||||
|
MaskFill::Alternating if k.is_multiple_of(2) => noise(),
|
||||||
|
MaskFill::Alternating | MaskFill::Compressible => 7,
|
||||||
|
};
|
||||||
|
out.extend_from_slice(&v.to_le_bytes()[..fam.elem]);
|
||||||
|
}
|
||||||
|
out
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
fn mask_families() -> Vec<MaskFamily> {
|
||||||
|
#[cfg_attr(not(feature = "blosc"), allow(unused_mut))]
|
||||||
|
let mut v = vec![
|
||||||
|
MaskFamily {
|
||||||
|
name: "lzf5",
|
||||||
|
elem: 1,
|
||||||
|
chunk: 5,
|
||||||
|
h5py_kw: "dict(compression='lzf')",
|
||||||
|
build: |d| {
|
||||||
|
d.with_lzf().without_shuffle();
|
||||||
|
},
|
||||||
|
fill: MaskFill::FiveBytes,
|
||||||
|
},
|
||||||
|
MaskFamily {
|
||||||
|
name: "lzf",
|
||||||
|
elem: 4,
|
||||||
|
chunk: 64,
|
||||||
|
h5py_kw: "dict(compression='lzf')",
|
||||||
|
build: |d| {
|
||||||
|
d.with_lzf().without_shuffle();
|
||||||
|
},
|
||||||
|
fill: MaskFill::Alternating,
|
||||||
|
},
|
||||||
|
MaskFamily {
|
||||||
|
name: "mix",
|
||||||
|
elem: 4,
|
||||||
|
chunk: 8,
|
||||||
|
h5py_kw: "dict(compression='lzf', shuffle=True, fletcher32=True)",
|
||||||
|
build: |d| {
|
||||||
|
d.with_lzf().with_shuffle().with_fletcher32();
|
||||||
|
},
|
||||||
|
fill: MaskFill::Alternating,
|
||||||
|
},
|
||||||
|
MaskFamily {
|
||||||
|
name: "lzfc",
|
||||||
|
elem: 4,
|
||||||
|
chunk: 64,
|
||||||
|
h5py_kw: "dict(compression='lzf')",
|
||||||
|
build: |d| {
|
||||||
|
d.with_lzf().without_shuffle();
|
||||||
|
},
|
||||||
|
fill: MaskFill::Compressible,
|
||||||
|
},
|
||||||
|
];
|
||||||
|
#[cfg(feature = "blosc")]
|
||||||
|
{
|
||||||
|
use clawhdf5_format::chunked_write::{BloscCodec, BloscShuffle};
|
||||||
|
v.push(MaskFamily {
|
||||||
|
name: "blosc",
|
||||||
|
elem: 4,
|
||||||
|
chunk: 64,
|
||||||
|
h5py_kw: "hdf5plugin.Blosc(cname='lz4', clevel=5, shuffle=hdf5plugin.Blosc.SHUFFLE)",
|
||||||
|
build: |d| {
|
||||||
|
d.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
|
||||||
|
},
|
||||||
|
fill: MaskFill::Alternating,
|
||||||
|
});
|
||||||
|
v.push(MaskFamily {
|
||||||
|
name: "blosc0",
|
||||||
|
elem: 4,
|
||||||
|
chunk: 64,
|
||||||
|
h5py_kw: "hdf5plugin.Blosc(cname='lz4', clevel=0, shuffle=hdf5plugin.Blosc.SHUFFLE)",
|
||||||
|
build: |d| {
|
||||||
|
d.with_blosc(BloscCodec::Lz4, 0, BloscShuffle::Byte);
|
||||||
|
},
|
||||||
|
fill: MaskFill::Compressible,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
v
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Builds h5py twins of our datasets and prints, per dataset, the filter
|
||||||
|
/// masks by chunk offset in our file and in the twin.
|
||||||
|
const MASK_TWIN: &str = r#"
|
||||||
|
import sys, numpy as np, h5py
|
||||||
|
try:
|
||||||
|
import hdf5plugin
|
||||||
|
except ImportError:
|
||||||
|
hdf5plugin = None
|
||||||
|
ours, twin, spec = sys.argv[1], sys.argv[2], eval(sys.argv[3])
|
||||||
|
def masks(ds):
|
||||||
|
return sorted((tuple(ds.id.get_chunk_info(k).chunk_offset), ds.id.get_chunk_info(k).filter_mask)
|
||||||
|
for k in range(ds.id.get_num_chunks()))
|
||||||
|
with h5py.File(ours, 'r') as o, h5py.File(twin, 'w', libver='v114') as t:
|
||||||
|
for name, dt, shape, chunks, maxshape, kw, raw in spec:
|
||||||
|
want = np.fromfile(raw, dtype=dt).reshape(shape)
|
||||||
|
assert np.array_equal(o[name][()], want), name
|
||||||
|
t.create_dataset(name, data=want, chunks=chunks, maxshape=maxshape, **eval(kw))
|
||||||
|
print(name, masks(o[name]), '|', masks(t[name]))
|
||||||
|
"#;
|
||||||
|
|
||||||
|
/// h5py (libhdf5) rewrites every chunk of our datasets — random chunks
|
||||||
|
/// become compressible and the other way round; the `[x,0,0,0,0]` chunks
|
||||||
|
/// change in place at the same size — then extends the resizable ones with
|
||||||
|
/// random data, and saves what each dataset must now hold. Prints the
|
||||||
|
/// datasets h5dump can decode (no chunk left LZF-encoded: h5dump has no
|
||||||
|
/// LZF filter).
|
||||||
|
const MASK_REWRITE: &str = r#"
|
||||||
|
import sys, numpy as np, h5py
|
||||||
|
try:
|
||||||
|
import hdf5plugin
|
||||||
|
except ImportError:
|
||||||
|
hdf5plugin = None
|
||||||
|
ours, spec = sys.argv[1], eval(sys.argv[2])
|
||||||
|
rng = np.random.default_rng(3)
|
||||||
|
def noise(shape, dt):
|
||||||
|
return rng.integers(0, 256, int(np.prod(shape)) * np.dtype(dt).itemsize,
|
||||||
|
dtype=np.uint8).view(dt).reshape(shape)
|
||||||
|
dumpable = []
|
||||||
|
with h5py.File(ours, 'r+') as f:
|
||||||
|
for name, dt, shape, chunks, maxshape, kw, raw in spec:
|
||||||
|
d = f[name]
|
||||||
|
want = d[()]
|
||||||
|
for s in d.iter_chunks():
|
||||||
|
blk = want[s]
|
||||||
|
if dt == 'u1':
|
||||||
|
blk.flat[-1] = 1
|
||||||
|
elif (blk == blk.flat[0]).all():
|
||||||
|
blk[...] = noise(blk.shape, dt)
|
||||||
|
else:
|
||||||
|
blk[...] = 5
|
||||||
|
d[...] = want
|
||||||
|
if maxshape is not None:
|
||||||
|
new = tuple(n + c for n, c in zip(shape, chunks))
|
||||||
|
grown = noise(new, dt)
|
||||||
|
grown[tuple(slice(0, n) for n in shape)] = want
|
||||||
|
d.resize(new)
|
||||||
|
d[...] = grown
|
||||||
|
want = grown
|
||||||
|
want.tofile(raw + '.want')
|
||||||
|
pl = d.id.get_create_plist()
|
||||||
|
ids = [pl.get_filter(i)[0] for i in range(pl.get_nfilters())]
|
||||||
|
if 32000 in ids:
|
||||||
|
bit = 1 << ids.index(32000)
|
||||||
|
if not all(d.id.get_chunk_info(k).filter_mask & bit for k in range(d.id.get_num_chunks())):
|
||||||
|
continue
|
||||||
|
dumpable.append(name)
|
||||||
|
with h5py.File(ours, 'r') as f:
|
||||||
|
for name, dt, shape, chunks, maxshape, kw, raw in spec:
|
||||||
|
want = np.fromfile(raw + '.want', dtype=dt).reshape(f[name].shape)
|
||||||
|
assert np.array_equal(f[name][()], want), name
|
||||||
|
print(' '.join(dumpable))
|
||||||
|
"#;
|
||||||
|
|
||||||
|
/// Optional filters that fail are skipped in files `FileBuilder` writes,
|
||||||
|
/// exactly as libhdf5 skips them: an LZF or Blosc output no smaller than the
|
||||||
|
/// chunk leaves the chunk stored unfiltered with the filter's mask bit set.
|
||||||
|
/// The writer used to store every chunk filtered with mask 0. For LZF, a
|
||||||
|
/// chunk whose LZF stream is exactly the chunk's size (`[x,0,0,0,0]`) was
|
||||||
|
/// then corrupted by the first libhdf5 rewrite of it: libhdf5 stores the
|
||||||
|
/// new data raw at the same size and, the size being unchanged, leaves the
|
||||||
|
/// stale mask in the index, so h5py could no longer read the dataset.
|
||||||
|
///
|
||||||
|
/// For every family × every chunk index the writer builds (single chunk,
|
||||||
|
/// Fixed Array, Extensible Array, version-2 B-tree): our masks equal those
|
||||||
|
/// of an h5py-written twin of the same data; then h5py r+ rewrites and
|
||||||
|
/// extends the datasets, and h5py, h5dump (where it has the filter) and our
|
||||||
|
/// reader read every value.
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
#[test]
|
||||||
|
fn skipped_optional_filters_are_masked_as_libhdf5_masks_them() {
|
||||||
|
let modules = if cfg!(feature = "blosc") {
|
||||||
|
"h5py, numpy, hdf5plugin"
|
||||||
|
} else {
|
||||||
|
"h5py, numpy"
|
||||||
|
};
|
||||||
|
if !have_python(modules) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let ours = dir.path().join("ours.h5");
|
||||||
|
let twin = dir.path().join("twin.h5");
|
||||||
|
let mut fb = clawhdf5::FileBuilder::new();
|
||||||
|
let mut spec = Vec::new();
|
||||||
|
let mut names = Vec::new();
|
||||||
|
for (fi, fam) in mask_families().iter().enumerate() {
|
||||||
|
for (label, shape, chunks, maxshape) in mask_layouts(fam.chunk) {
|
||||||
|
let name = format!("{}_{label}", fam.name);
|
||||||
|
let data = mask_data(fam, &shape, fi as u64 * 31 + shape.len() as u64);
|
||||||
|
let raw = dir.path().join(format!("{name}.raw"));
|
||||||
|
std::fs::write(&raw, &data).unwrap();
|
||||||
|
let ds = fb.create_dataset(&name);
|
||||||
|
if fam.elem == 1 {
|
||||||
|
ds.with_u8_data(&data);
|
||||||
|
} else {
|
||||||
|
let v: Vec<i32> = data
|
||||||
|
.as_chunks::<4>()
|
||||||
|
.0
|
||||||
|
.iter()
|
||||||
|
.map(|&b| i32::from_le_bytes(b))
|
||||||
|
.collect();
|
||||||
|
ds.with_i32_data(&v);
|
||||||
|
}
|
||||||
|
ds.with_shape(&shape).with_chunks(&chunks);
|
||||||
|
if let Some(ms) = &maxshape {
|
||||||
|
ds.with_maxshape(ms);
|
||||||
|
}
|
||||||
|
(fam.build)(ds);
|
||||||
|
let py_tuple = |v: &[u64]| {
|
||||||
|
let items: Vec<String> = v
|
||||||
|
.iter()
|
||||||
|
.map(|&d| {
|
||||||
|
if d == u64::MAX {
|
||||||
|
"None".into()
|
||||||
|
} else {
|
||||||
|
d.to_string()
|
||||||
|
}
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
format!("({},)", items.join(","))
|
||||||
|
};
|
||||||
|
spec.push(format!(
|
||||||
|
"({name:?}, {:?}, {}, {}, {}, {:?}, {:?})",
|
||||||
|
if fam.elem == 1 { "u1" } else { "<i4" },
|
||||||
|
py_tuple(&shape),
|
||||||
|
py_tuple(&chunks),
|
||||||
|
maxshape.as_deref().map_or("None".into(), py_tuple),
|
||||||
|
fam.h5py_kw,
|
||||||
|
raw.to_str().unwrap(),
|
||||||
|
));
|
||||||
|
names.push((name, fam.elem));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
fb.write(&ours).unwrap();
|
||||||
|
let spec = format!("[{}]", spec.join(", "));
|
||||||
|
|
||||||
|
// Our masks are h5py's, chunk by chunk.
|
||||||
|
let out = run_python(
|
||||||
|
MASK_TWIN,
|
||||||
|
&[ours.to_str().unwrap(), twin.to_str().unwrap(), &spec],
|
||||||
|
);
|
||||||
|
let mut skipped = 0;
|
||||||
|
for line in out.lines() {
|
||||||
|
let (name, rest) = line.split_once(' ').unwrap();
|
||||||
|
let (got, want) = rest.split_once(" | ").unwrap();
|
||||||
|
assert_eq!(got, want, "{name}: our filter masks (left) vs h5py's");
|
||||||
|
skipped += usize::from(got.contains("), 1)") || got.contains("), 2)"));
|
||||||
|
}
|
||||||
|
assert_eq!(out.lines().count(), names.len());
|
||||||
|
assert!(
|
||||||
|
skipped >= 12,
|
||||||
|
"too few datasets with skipped filters:\n{out}"
|
||||||
|
);
|
||||||
|
|
||||||
|
// libhdf5 rewrites and extends them; everyone reads the new values.
|
||||||
|
let dumpable = run_python(MASK_REWRITE, &[ours.to_str().unwrap(), &spec]);
|
||||||
|
let plugin_path = run_python("import hdf5plugin; print(hdf5plugin.PLUGIN_PATH)", &[]);
|
||||||
|
let file = File::open(&ours).unwrap();
|
||||||
|
for (name, elem) in &names {
|
||||||
|
let want = std::fs::read(dir.path().join(format!("{name}.raw.want"))).unwrap();
|
||||||
|
let got = file
|
||||||
|
.dataset(name)
|
||||||
|
.unwrap()
|
||||||
|
.read_selection(&Selection::All)
|
||||||
|
.unwrap();
|
||||||
|
assert!(got == want, "{name}: our reader after h5py r+");
|
||||||
|
if !dumpable.split(' ').any(|d| d == name) || Command::new("h5dump").output().is_err() {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let o = Command::new("h5dump")
|
||||||
|
.env("HDF5_PLUGIN_PATH", &plugin_path)
|
||||||
|
.args(["-d", name, "-y", "-w", "0", ours.to_str().unwrap()])
|
||||||
|
.output()
|
||||||
|
.unwrap();
|
||||||
|
assert!(
|
||||||
|
o.status.success(),
|
||||||
|
"h5dump -d {name}: {}",
|
||||||
|
String::from_utf8_lossy(&o.stderr)
|
||||||
|
);
|
||||||
|
let s = String::from_utf8_lossy(&o.stdout).into_owned();
|
||||||
|
let vals: Vec<i64> = s
|
||||||
|
.split_once("DATA {")
|
||||||
|
.and_then(|(_, r)| r.split_once('}'))
|
||||||
|
.map(|(d, _)| {
|
||||||
|
d.split(|c: char| c == ',' || c.is_whitespace())
|
||||||
|
.filter(|t| !t.is_empty())
|
||||||
|
.map(|t| t.parse::<i64>().unwrap())
|
||||||
|
.collect()
|
||||||
|
})
|
||||||
|
.unwrap_or_default();
|
||||||
|
let want_vals: Vec<i64> = want
|
||||||
|
.chunks_exact(*elem)
|
||||||
|
.map(|b| {
|
||||||
|
if *elem == 1 {
|
||||||
|
i64::from(b[0])
|
||||||
|
} else {
|
||||||
|
i64::from(i32::from_le_bytes(b.try_into().unwrap()))
|
||||||
|
}
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
assert_eq!(vals, want_vals, "h5dump -d {name}");
|
||||||
|
}
|
||||||
|
assert!(
|
||||||
|
dumpable.split(' ').any(|d| d.starts_with("lzf5_")),
|
||||||
|
"{dumpable}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Files whose chunks all compress are written exactly as before optional
|
||||||
|
/// filters could be skipped: every mask is 0 and nothing else changed. The
|
||||||
|
/// hashes are of the files the writer produced before that change.
|
||||||
|
#[cfg(feature = "lzf")]
|
||||||
|
#[test]
|
||||||
|
fn files_whose_chunks_all_compress_are_unchanged() {
|
||||||
|
use clawhdf5_format::checksum::jenkins_lookup3;
|
||||||
|
#[allow(clippy::type_complexity)]
|
||||||
|
#[cfg_attr(not(feature = "blosc"), allow(unused_mut))]
|
||||||
|
let mut cases: Vec<(
|
||||||
|
&str,
|
||||||
|
fn(&mut clawhdf5_format::type_builders::DatasetBuilder),
|
||||||
|
(usize, u32),
|
||||||
|
)> = vec![
|
||||||
|
(
|
||||||
|
"lzf_fixed",
|
||||||
|
|d| {
|
||||||
|
d.with_i32_data(&ramp_i32(4000))
|
||||||
|
.with_chunks(&[500])
|
||||||
|
.with_lzf();
|
||||||
|
},
|
||||||
|
(3965, 449169442),
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"lzf_ea_noshuffle",
|
||||||
|
|d| {
|
||||||
|
d.with_i32_data(&ramp_i32(4000))
|
||||||
|
.with_chunks(&[700])
|
||||||
|
.with_maxshape(&[u64::MAX])
|
||||||
|
.with_lzf()
|
||||||
|
.without_shuffle();
|
||||||
|
},
|
||||||
|
(7213, 4277403206),
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"mix_bt2",
|
||||||
|
|d| {
|
||||||
|
d.with_f64_data(&ramp_f64(40 * 60))
|
||||||
|
.with_shape(&[40, 60])
|
||||||
|
.with_chunks(&[16, 16])
|
||||||
|
.with_maxshape(&[u64::MAX, u64::MAX])
|
||||||
|
.with_lzf()
|
||||||
|
.with_fletcher32();
|
||||||
|
},
|
||||||
|
(11495, 3340532700),
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"lzf_single",
|
||||||
|
|d| {
|
||||||
|
d.with_u8_data(&ramp_u8(3000))
|
||||||
|
.with_chunks(&[3000])
|
||||||
|
.with_lzf();
|
||||||
|
},
|
||||||
|
(546, 690805477),
|
||||||
|
),
|
||||||
|
];
|
||||||
|
#[cfg(feature = "blosc")]
|
||||||
|
cases.push((
|
||||||
|
"blosc_fixed",
|
||||||
|
|d| {
|
||||||
|
use clawhdf5_format::chunked_write::{BloscCodec, BloscShuffle};
|
||||||
|
d.with_i32_data(&ramp_i32(5000))
|
||||||
|
.with_chunks(&[1024])
|
||||||
|
.with_blosc(BloscCodec::Lz4, 5, BloscShuffle::Byte);
|
||||||
|
},
|
||||||
|
(2776, 4278611376),
|
||||||
|
));
|
||||||
|
for (name, build, want) in &cases {
|
||||||
|
let mut fb = clawhdf5::FileBuilder::new();
|
||||||
|
build(fb.create_dataset("d"));
|
||||||
|
let bytes = fb.finish().unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
(bytes.len(), jenkins_lookup3(&bytes)),
|
||||||
|
*want,
|
||||||
|
"{name}: (length, lookup3 hash) of the file"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@@ -7,6 +7,23 @@ deleting it.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## LZF/Blosc chunks written with a stale filter mask
|
||||||
|
|
||||||
|
**Status:** fixed 2026-09-26, before any release (the LZF and Blosc writers
|
||||||
|
were added the same day; v2.7.0 and earlier write neither).
|
||||||
|
|
||||||
|
`FileBuilder` stored every chunk of an LZF or Blosc dataset through the
|
||||||
|
filter with filter mask 0. libhdf5 counts LZF and Blosc output no smaller
|
||||||
|
than the chunk as a failure of the (optional) filter and stores the chunk
|
||||||
|
raw with the filter's mask bit set. When a chunk's LZF stream was exactly
|
||||||
|
the chunk's size, the first libhdf5 rewrite of it stored raw data at the
|
||||||
|
same size and left our mask 0 in the index, so h5py could no longer read
|
||||||
|
the dataset. `FileEditor` had the same bug, fixed earlier the same day.
|
||||||
|
Both now use `clawhdf5_format::filters::compress_chunk_masked`, and every
|
||||||
|
chunk index the writer builds records the real mask (see `CHANGELOG.md`).
|
||||||
|
Files written before the fix read correctly; rewrite them before letting
|
||||||
|
libhdf5 modify them.
|
||||||
|
|
||||||
## In-place modification (`FileEditor`) limits
|
## In-place modification (`FileEditor`) limits
|
||||||
|
|
||||||
**Status:** open (documented 2026-09-26). `clawhdf5::FileEditor` refuses,
|
**Status:** open (documented 2026-09-26). `clawhdf5::FileEditor` refuses,
|
||||||
|
|||||||
Reference in New Issue
Block a user