fix(read): decode on the calling thread when rayon's pool has one thread
Full reads of chunked datasets handed their chunks to rayon. With a one-thread pool (concurrent_read --decode-threads 1, RAYON_NUM_THREADS=1) every thread reading through a File queued behind that single worker, so 16 readers decoded on one core: per-thread CPU time showed one thread doing all the decoding and the readers almost none, and full reads stopped at about 2x one thread. The cached full-read path and the uncached reader behind verify_provenance now decode inline when the pool cannot parallelise (parallel_read::pool_can_parallelise). The File's chunk cache was the suspect but not the cause: datasets over its budget were already read without inserting, and skipping its lookups gained only a few percent at 16 threads. The regression test keeps a one-thread global pool's worker busy and requires a full read and verify_provenance to finish anyway; before the fix both waited for the worker (timed out). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -27,6 +27,20 @@ pub fn should_use_parallel(chunk_count: usize) -> bool {
|
||||
chunk_count > PARALLEL_THRESHOLD
|
||||
}
|
||||
|
||||
/// Whether handing a read's chunks to rayon can decode them faster than the
|
||||
/// calling thread would alone.
|
||||
///
|
||||
/// `false` when the pool the work would go to (the current pool inside a
|
||||
/// rayon worker, else the global one) has a single thread. Handing work to
|
||||
/// that pool is then worse than useless: the caller blocks while the one
|
||||
/// worker decodes, and every other thread reading at the same time queues
|
||||
/// behind the same worker, so N reader threads decode on one core. (That is
|
||||
/// how full reads with `--decode-threads 1` stopped scaling at about 2x in
|
||||
/// the `concurrent_read` benchmark.)
|
||||
pub fn pool_can_parallelise() -> bool {
|
||||
rayon::current_num_threads() > 1
|
||||
}
|
||||
|
||||
/// Decompress chunks in parallel using lane-partitioned assignment.
|
||||
///
|
||||
/// Instead of naive `par_iter`, chunks are deterministically assigned to lanes
|
||||
|
||||
Reference in New Issue
Block a user