Bug #22224
Updated by yaroslavmarkin (Yaroslav Markin) about 2 months ago
Hi! Fair warning: a significant portion of this bugreport and experimentation/benchmarking was done by an agent.
Sorry for the slop, but it was so much faster to benchmark and find a root issue.
**Proposed PR: https://github.com/ruby/ruby/pull/18176**
## Summary
With YJIT enabled, `vm_make_env_each` calls `rb_yjit_invalidate_ep_is_bp` on every Proc/lambda environment materialization. That function enters `with_vm_lock` unconditionally (the only early return covers boot, before `INVARIANTS` is initialized). In multi-Ractor mode this lock acquisition is `rb_jit_vm_lock_then_barrier`: the VM lock plus a stop-all-Ractors barrier.
The result: any workload that creates Procs loses parallel scalability under Ractors, and beyond ~2 Ractors adding workers makes the whole process slower in absolute terms. In the repro below, 8 Ractors with YJIT run ~150x slower than 1 Ractor with YJIT, and ~195x slower than 8 Ractors without YJIT, on the same code.
Note the invalidation itself is not the cost. After the first escape of a given iseq, its `no_ep_escape_iseqs` entry is an empty set forever (`ep_is_bp()` returns false for it, so no new blocks are ever registered), yet every subsequent materialization still pays lock + global barrier to look up the entry and iterate nothing.
## Standalone repro (no gems)
```ruby
# Usage: ruby [--yjit] yjit_ractor_repro.rb [proc|calc] [n_ractors] [seconds]
MODE = (ARGV[0] || "proc").to_sym
N = (ARGV[1] || 1).to_i
DUR = (ARGV[2] || 3).to_f
def make_proc(i)
x = i
-> { x + 1 } # captures x: the frame env is materialized on every call
end
def work_proc(n)
s = 0
n.times { |i| s += make_proc(i).call }
s
end
def work_calc(n)
s = 0
n.times { |i| s += (i * i) % 7 } # same shape, no Proc escapes
s
end
def bench_loop(mode, dur)
count = 0
deadline = Process.clock_gettime(Process::CLOCK_MONOTONIC) + dur
while Process.clock_gettime(Process::CLOCK_MONOTONIC) < deadline
mode == :proc ? work_proc(1000) : work_calc(1000)
count += 1000
end
count
end
bench_loop(MODE, 0.5) # warm up: get the methods YJIT-compiled before measuring
total =
if N == 1
bench_loop(MODE, DUR)
else
N.times.map { Ractor.new(MODE, DUR) { |m, d| bench_loop(m, d) } }.sum(&:value)
end
puts format("mode=%s yjit=%s ractors=%d throughput=%.2fM iters/s",
MODE, RubyVM::YJIT.enabled?, N, total / DUR / 1_000_000.0)
```
Results on Apple M1 Pro (10 cores), macOS, `ruby 4.0.6 (2026-07-14 revision 03b6d3f889) +PRISM [arm64-darwin25]`, 3 s per cell:
| workload | ractors | interpreter (M iters/s) | --yjit (M iters/s) |
|----------|---------|-------------------------|--------------------|
| calc (control, no Proc) | 1 | 21.3 | 129.4 |
| calc | 4 | 83.0 | 495.3 |
| calc | 8 | 162.2 | **969.1** |
| proc | 1 | 4.5 | 6.0 |
| proc | 4 | 7.7 | **0.14** |
| proc | 8 | 7.8 | **0.04** |
The control shows this is not a general YJIT-vs-Ractor problem: on the Proc-free workload YJIT scales superbly (969M iters/s at 8 Ractors). Only the environment-materializing workload collapses, and only with YJIT on.
## Where the time goes
Native sampling (macOS `sample`) of a loaded multi-Ractor process shows almost all threads parked in `__psynch_cvwait`, with the barrier initiated from:
```
vm_make_env_each
-> rb_yjit_invalidate_ep_is_bp (yjit/src/invariants.rs)
-> with_vm_lock -> rb_jit_vm_lock_then_barrier
-> rb_ractor_sched_barrier_start
(victims: rb_ractor_sched_barrier_join / ractor_sched_barrier_join_wait_locked)
```
In a 6 s sample of a 5-Ractor process, barrier-related frames appear ~9,000 times vs ~50 in the single-Ractor run of the same workload. Ruby-side profilers cannot see this: the wait time is attributed as diffuse "self time" across whatever frames are on top, which is presumably why it has gone unreported.
Current code (`yjit/src/invariants.rs`, same on master as of 2026-08-03):
```rust
pub extern "C" fn rb_yjit_invalidate_ep_is_bp(iseq: IseqPtr) {
// Skip tracking EP escapes on boot. We don't need to invalidate anything during boot.
if unsafe { INVARIANTS.is_none() } {
return;
}
with_vm_lock(src_loc!(), || {
let no_ep_escape_iseqs = &mut Invariants::get_instance().no_ep_escape_iseqs;
match no_ep_escape_iseqs.get_mut(&iseq) {
Some(blocks) => {
for block in mem::take(blocks) {
invalidate_block_version(&block);
incr_counter!(invalidate_ep_escape);
}
}
None => {
no_ep_escape_iseqs.insert(iseq, HashSet::new());
}
}
});
}
```
## Real-world impact
Rails 8.1 enables YJIT by default, and a Rails request materializes many Proc environments (middleware blocks, route handling, view rendering), so any Rails app served by a multi-Ractor server hits this out of the box. Found while investigating https://github.com/yaroslav/kino/issues/6, where a stock Rails 8.1 app showed worker scaling inverting: more Ractors, less total throughput.
Rails 8.1.3 health-check endpoint (`/up`, no database), `ab -c 64 -k`, kino `:ractor` mode, N workers x 1 thread, requests/sec:
| workers | YJIT on (Rails default) | YJIT off |
|---------|-------------------------|----------|
| 1 | 6,416 | 3,964 |
| 2 | 5,036 | 6,999 |
| 5 | 2,800 | 13,269 |
| 8 | 2,003 | 12,596 |
Fully disabling GC changed the YJIT-on numbers by only ~7%, ruling out GC barriers as the driver; the native profile above identifies the initiator.