Project

General

Profile

Bug #22224

Updated by yaroslavmarkin (Yaroslav Markin) about 2 months ago

Hi! Fair warning: a significant portion of this bugreport and experimentation/benchmarking was done by an agent. 

 Sorry for the slop, but it was so much faster to benchmark and find a root issue. 

 **Proposed PR: https://github.com/ruby/ruby/pull/18176** 

 ## Summary 

 With YJIT enabled, `vm_make_env_each` calls `rb_yjit_invalidate_ep_is_bp` on every Proc/lambda environment materialization. That function enters `with_vm_lock` unconditionally (the only early return covers boot, before `INVARIANTS` is initialized). In multi-Ractor mode this lock acquisition is `rb_jit_vm_lock_then_barrier`: the VM lock plus a stop-all-Ractors barrier. 

 The result: any workload that creates Procs loses parallel scalability under Ractors, and beyond ~2 Ractors adding workers makes the whole process slower in absolute terms. In the repro below, 8 Ractors with YJIT run ~150x slower than 1 Ractor with YJIT, and ~195x slower than 8 Ractors without YJIT, on the same code. 

 Note the invalidation itself is not the cost. After the first escape of a given iseq, its `no_ep_escape_iseqs` entry is an empty set forever (`ep_is_bp()` returns false for it, so no new blocks are ever registered), yet every subsequent materialization still pays lock + global barrier to look up the entry and iterate nothing. 

 ## Standalone repro (no gems) 

 ```ruby 
 # Usage: ruby [--yjit] yjit_ractor_repro.rb [proc|calc] [n_ractors] [seconds] 

 MODE = (ARGV[0] || "proc").to_sym 
 N = (ARGV[1] || 1).to_i 
 DUR = (ARGV[2] || 3).to_f 

 def make_proc(i) 
   x = i 
   -> { x + 1 } # captures x: the frame env is materialized on every call 
 end 

 def work_proc(n) 
   s = 0 
   n.times { |i| s += make_proc(i).call } 
   s 
 end 

 def work_calc(n) 
   s = 0 
   n.times { |i| s += (i * i) % 7 } # same shape, no Proc escapes 
   s 
 end 

 def bench_loop(mode, dur) 
   count = 0 
   deadline = Process.clock_gettime(Process::CLOCK_MONOTONIC) + dur 
   while Process.clock_gettime(Process::CLOCK_MONOTONIC) < deadline 
     mode == :proc ? work_proc(1000) : work_calc(1000) 
     count += 1000 
   end 
   count 
 end 

 bench_loop(MODE, 0.5) # warm up: get the methods YJIT-compiled before measuring 

 total = 
   if N == 1 
     bench_loop(MODE, DUR) 
   else 
     N.times.map { Ractor.new(MODE, DUR) { |m, d| bench_loop(m, d) } }.sum(&:value) 
   end 

 puts format("mode=%s yjit=%s ractors=%d throughput=%.2fM iters/s", 
   MODE, RubyVM::YJIT.enabled?, N, total / DUR / 1_000_000.0) 
 ``` 

 Results on Apple M1 Pro (10 cores), macOS, `ruby 4.0.6 (2026-07-14 revision 03b6d3f889) +PRISM [arm64-darwin25]`, 3 s per cell: 

 | workload | ractors | interpreter (M iters/s) | --yjit (M iters/s) | 
 |----------|---------|-------------------------|--------------------| 
 | calc (control, no Proc) | 1 | 21.3 | 129.4 | 
 | calc | 4 | 83.0 | 495.3 | 
 | calc | 8 | 162.2 | **969.1** | 
 | proc | 1 | 4.5 | 6.0 | 
 | proc | 4 | 7.7 | **0.14** | 
 | proc | 8 | 7.8 | **0.04** | 

 The control shows this is not a general YJIT-vs-Ractor problem: on the Proc-free workload YJIT scales superbly (969M iters/s at 8 Ractors). Only the environment-materializing workload collapses, and only with YJIT on. 

 ## Where the time goes 

 Native sampling (macOS `sample`) of a loaded multi-Ractor process shows almost all threads parked in `__psynch_cvwait`, with the barrier initiated from: 

 ``` 
 vm_make_env_each 
 -> rb_yjit_invalidate_ep_is_bp          (yjit/src/invariants.rs) 
 -> with_vm_lock -> rb_jit_vm_lock_then_barrier 
 -> rb_ractor_sched_barrier_start 
    (victims: rb_ractor_sched_barrier_join / ractor_sched_barrier_join_wait_locked) 
 ``` 

 In a 6 s sample of a 5-Ractor process, barrier-related frames appear ~9,000 times vs ~50 in the single-Ractor run of the same workload. Ruby-side profilers cannot see this: the wait time is attributed as diffuse "self time" across whatever frames are on top, which is presumably why it has gone unreported. 

 Current code (`yjit/src/invariants.rs`, same on master as of 2026-08-03): 

 ```rust 
 pub extern "C" fn rb_yjit_invalidate_ep_is_bp(iseq: IseqPtr) { 
     // Skip tracking EP escapes on boot. We don't need to invalidate anything during boot. 
     if unsafe { INVARIANTS.is_none() } { 
         return; 
     } 

     with_vm_lock(src_loc!(), || { 
         let no_ep_escape_iseqs = &mut Invariants::get_instance().no_ep_escape_iseqs; 
         match no_ep_escape_iseqs.get_mut(&iseq) { 
             Some(blocks) => { 
                 for block in mem::take(blocks) { 
                     invalidate_block_version(&block); 
                     incr_counter!(invalidate_ep_escape); 
                 } 
             } 
             None => { 
                 no_ep_escape_iseqs.insert(iseq, HashSet::new()); 
             } 
         } 
     }); 
 } 
 ``` 

 ## Real-world impact 

 Rails 8.1 enables YJIT by default, and a Rails request materializes many Proc environments (middleware blocks, route handling, view rendering), so any Rails app served by a multi-Ractor server hits this out of the box. Found while investigating https://github.com/yaroslav/kino/issues/6, where a stock Rails 8.1 app showed worker scaling inverting: more Ractors, less total throughput. 

 Rails 8.1.3 health-check endpoint (`/up`, no database), `ab -c 64 -k`, kino `:ractor` mode, N workers x 1 thread, requests/sec: 

 | workers | YJIT on (Rails default) | YJIT off | 
 |---------|-------------------------|----------| 
 | 1 | 6,416 | 3,964 | 
 | 2 | 5,036 | 6,999 | 
 | 5 | 2,800 | 13,269 | 
 | 8 | 2,003 | 12,596 | 

 Fully disabling GC changed the YJIT-on numbers by only ~7%, ruling out GC barriers as the driver; the native profile above identifies the initiator. 

Back