format: inline the version-1 message loop again (ObjectHeader::parse back to 8f59b2e's speed)
4313917kept parse_v1_messages out of line (#[inline(never)]) so the generic parser would not carry the loop; sinceef428d7the slice path is compiled once in this crate, and the call itself was the remaining cost of the chunk queue: A/B builds of object_header_parse_x401 with only this attribute changed put #[inline(never)] and no attribute at 24.5-24.9 us and #[inline] at 23.6-24.0 us, with8f59b2eat 23.6-24.1 us. Lazily creating the chunk list only when a continuation is found (tried too) measured no faster and was not kept. Same code otherwise: every chunk-queue check (65,536 chunks, cycles, file size budget, one chunk buffer at a time, libhdf5 order, overlap allowed) is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -295,7 +295,13 @@ impl ObjectHeader {
|
||||
/// The messages of one version-1 chunk: each checked and appended to
|
||||
/// `messages` (NIL ones dropped), each continuation added to `spans`.
|
||||
/// Returns how many messages (NIL ones included) the chunk holds.
|
||||
#[inline(never)]
|
||||
///
|
||||
/// Inlined into the chunk loop: kept out of line (`#[inline(never)]`,
|
||||
/// 4313917), the call cost `ObjectHeader::parse` about 2.5 ns per
|
||||
/// header, 4% on small version-1 headers (A/B builds, 2026-09-27; see
|
||||
/// `BENCHMARKS.md`). Without an attribute the compiler keeps it out of
|
||||
/// line too.
|
||||
#[inline]
|
||||
fn parse_v1_messages(
|
||||
data: &[u8],
|
||||
offset_size: u8,
|
||||
|
||||
Reference in New Issue
Block a user