Project

General

Profile

Feature #22255

Updated by ko1 (Koichi Sasada) about 2 months ago

## Abstract 

 Add a `timeout:` keyword to the three ways a Ractor waits for a message: 

 ```ruby 
 Ractor::Port#receive(timeout: nil)        # -> msg or nil 
 Ractor.receive(timeout: nil)              # -> msg or nil 
 Ractor.select(*ports, timeout: nil)       # -> [port, msg] or nil 
 ``` 

 They return `nil` once the timeout passes. `timeout: 0` never blocks: it takes a 
 message if one is already there and returns `nil` otherwise. 

 ## Background 

 Today a Ractor wait cannot be bounded. Once a Ractor calls `receive`, it waits 
 until a message arrives, so there is no way to write 

 * a worker that gives up and reports progress if nothing arrives for a while, 
 * a shutdown that stops waiting after a grace period, 
 * a non-blocking "is there anything for me?" check. 

 The usual workarounds are worse than a timeout: a watchdog Ractor that sends a 
 dummy message, or a helper thread that closes the port. Both add a Ractor or a 
 thread per wait and change the shape of the program. 

 Every other blocking wait in Ruby already has this: `Thread::Queue#pop(timeout:)`, 
 `ConditionVariable#wait(mutex, timeout)`, `Thread#join(limit)`, `IO#wait(timeout)`, 
 `Mutex#sleep(timeout)`. 

 ## Specification 

 | call | result | 
 |---|---| 
 | `port.receive` | blocks until a message arrives (unchanged) | 
 | `port.receive(timeout: 1.5)` | the message, or `nil` after 1.5 seconds | 
 | `port.receive(timeout: 0)` | the message if one is already queued, else `nil`, without blocking | 
 | `port.receive(timeout: -1)` | `ArgumentError` | 
 | `Ractor.select(a, b, timeout: 1.5)` | `[port, msg]`, or `nil` after 1.5 seconds | 

 The timeout behaves as it does elsewhere in Ruby: `nil` when it passes, and 
 `timeout: 0` polls without blocking, as in `Thread::Queue#pop(timeout:)`, 
 `IO.select` and `IO#wait`. 

 * `timeout` accepts what `sleep` accepts (Integer, Float, Rational); a non-numeric 
   value raises `TypeError`. 
 * A negative timeout raises `ArgumentError`, as `IO.select` does. Note that 
   `Thread::Queue#pop(timeout: -1)` instead returns `nil` right away, so the two 
   existing APIs already disagree here; this follows `IO.select`. 
 * `timeout: nil` (the default) means no timeout, i.e. today's behaviour. 
 * A closed port still raises `Ractor::ClosedError`, timeout or not. 
 * The timeout bounds how long the call blocks; it does not cut delivery off. A 
   message that arrives while the timeout is being reported is still returned. 
   Nothing is lost either way: a message only leaves the queue when it is returned. 

 ### Why `nil` 

 `nil` on timeout is what `Thread::Queue#pop(timeout:)`, `IO.select`, `IO#wait` and 
 `Thread#join` all return. It is ambiguous when `nil` itself is a valid message, 
 which is the same ambiguity `Queue#pop` has. The alternatives are a new exception 
 class (`Ractor::TimeoutError`) or a sentinel object; both were judged heavier than 
 the ambiguity is worth, but this is worth confirming. 

 ### Why `timeout: 0` rather than a separate `try_receive` 

 `timeout: 0` covers the non-blocking case with no new method name, and it is a 
 genuine fast path in the implementation: it converts nothing, reads no clock and 
 never parks. It costs about what a `receive` of an already-waiting message costs 
 (~70ns vs ~150ns in a micro benchmark), against ~1.7us for a `receive(timeout: 10)` 
 that has to compute a deadline. 

 ## Implementation 

 The wait stays where it is. `rb_ractor_sched_wait()` keeps parking the thread in 
 the thread scheduler, so **an M:N thread still hands its native thread back** 
 instead of becoming a dedicated one. Measured with 50 Ractors waiting at once, a 
 timed `receive` uses the same 18 native threads an untimed one does. 

 How the deadline is taken depends on the thread: 

 * A **dedicated native thread** parks on its own condvar, so it takes the deadline 
   there, the way `native_cond_sleep()` does. The condvar it waits on is the one a 
   send already signals, so nothing else is involved. 
 * An **M:N thread** has no condvar of its own, so its deadline is armed on the 
   timer thread as a timeout-only wheel entry. The timer thread then wakes it 
   through `thread_sched_to_ready_common()`, which is exactly how a send wakes it. 

 That gives an M:N waiter two wakers, so the send path takes an armed timeout back 
 before waking, skips a thread a fired timeout already made runnable, and bumps the 
 scheduler event serial so a timeout that has not fired cannot wake it twice. 

 Both kinds of wait exist on every pthread platform, including builds without the 
 timer wheel (`USE_MN_THREADS == 0`, e.g. s390x-linux), where every thread is 
 dedicated. On win32 the wait is already a condvar wait, which takes the timeout 
 directly. 

 Timer resolution follows the thread kind, exactly as `sleep` already does on the 
 same thread: a dedicated thread gets the condvar's resolution, an M:N thread gets 
 the timer wheel's 1ms tick. Measured, `receive(timeout: 0.0002)` and 
 `sleep(0.0002)` return in 0.275ms on the main thread and 1.15ms inside a Ractor. 

 ## Verification 

 Besides `test/ruby/test_ractor.rb` additions: an attack suite of 13 scenarios 
 (timeout racing with send, `Port#close`, `Thread#kill`, `Thread#raise`, Ractor 
 termination, GC and compaction, `fork`, 64 waiters expiring in the same instant, 
 all-Ractors-waiting deadlock detection) and a sweep that injects a trap, `SIGINT`, 
 `Thread#kill` and `Thread#raise` at four points of the wait, including the instant 
 the timer fires. Run against a release build, a `VM_CHECK_MODE` build and a 
 `USE_MN_THREADS=0` build, with `RUBY_MN_THREADS=0` and `=1`. 

 Of ~120,000 interrupt injections, ~78,000 landed while the thread was parked: no 
 missed wakeup, no wait cut short by a trap, no assertion. 2,000,000 arm/disarm 
 cycles showed no timer-wheel growth. 

 ## Open questions 

 1. `nil` on timeout, or a `Ractor::TimeoutError`? 
 1a. A negative timeout: `ArgumentError` as in `IO.select`, or `nil` as in 
     `Thread::Queue#pop`? 
 2. Should `Ractor.select` return something other than `nil`, given it already 
    returns an Array? 

 ## Patch 

 https://github.com/ruby/ruby/pull/18418 

 ## Notes (by ko1) 

 * Above description and a patch is written by Claude Code. 
 * I believe this proposal is enough simple, acceptable. 

Back