The worker's input-materialization paths (download_to_directory,
fs_util::hardlink_directory_tree_recursive, fs_util::set_perms_recursive, the directory
cache's clone/construct paths) issue a tokio::spawn_blocking and touch per-process shared
state per file. An action whose input tree carries a hermetic toolchain is thousands of
files; at ~200 concurrent actions this produces millions of acquisitions of two per-process
global locks:
- tokio's blocking-pool spawner mutex (every
spawn_blocking takes it, plus a condvar wake), and
- the store's
EvictingMap state mutex (nativelink-util/src/evicting_map.rs —
parking_lot::Mutex around the whole LRU; taken on every blob get/insert/touch).
The critical sections are short, but on a large multi-socket guest the cost is the handoff:
each contended lock transfer is a cross-socket cache-line migration, and with hundreds of
waiters the queue drains at cache-coherency speed. Userspace backtraces of the worker under
load (eu-stack, ~900 threads, v1.5.2 + the fadvise shim from Finding 1 already active):
450 parking_lot::raw_mutex::RawMutex::lock_slow
389 parking_lot::condvar::Condvar::wait_until_internal (blocking-pool waiters)
153 tokio::runtime::blocking::pool::Spawner::spawn_task (holds/queues on the pool lock)
113 nativelink_util::fs_util::hardlink_directory_tree_recursive::{closure}
54 nativelink_util::fs_util::set_perms_recursive_impl::{closure}
26 nativelink_worker::directory_cache::DirectoryCache::copy_file_to::{closure}
Steady state at that moment: 1.0–1.5 actions/s with 200 actions in "executing", ~7% CPU on
255 vCPUs, ~1% PSI-IO — spawned commands themselves finished in 0.3–2 s.
The worker's input-materialization paths (
download_to_directory,fs_util::hardlink_directory_tree_recursive,fs_util::set_perms_recursive, the directorycache's clone/construct paths) issue a
tokio::spawn_blockingand touch per-process sharedstate per file. An action whose input tree carries a hermetic toolchain is thousands of
files; at ~200 concurrent actions this produces millions of acquisitions of two per-process
global locks:
spawn_blockingtakes it, plus a condvar wake), andEvictingMapstate mutex (nativelink-util/src/evicting_map.rs—parking_lot::Mutexaround the whole LRU; taken on every blob get/insert/touch).The critical sections are short, but on a large multi-socket guest the cost is the handoff:
each contended lock transfer is a cross-socket cache-line migration, and with hundreds of
waiters the queue drains at cache-coherency speed. Userspace backtraces of the worker under
load (
eu-stack, ~900 threads, v1.5.2 + the fadvise shim from Finding 1 already active):Steady state at that moment: 1.0–1.5 actions/s with 200 actions in "executing", ~7% CPU on
255 vCPUs, ~1% PSI-IO — spawned commands themselves finished in 0.3–2 s.