Project

General

Profile

Bug #22196

Updated by ioquatix (Samuel Williams) 3 months ago

# Heap-use-after-free in `fiber_switch` reading terminated fiber's status (regression via Bug #21955 backport) 

 Ruby 3.4.10 intermittently segfaults while an async-http workload terminates task fibers after truncated chunked HTTP responses. The equivalent workload completed 20 high-churn attempts on Ruby 3.4.9 without a crash. ASan shows a heap-use-after-free: the outgoing fiber's `rb_fiber_t` is freed by GC while `fiber_switch` still reads its status on the `Fiber#transfer` path. 

 ## Environment 

 ```text 
 ruby 3.4.10 (2026-06-30 revision 2b0b7728dc) +PRISM [aarch64-linux] 
 ``` 

 - Linux `aarch64` Docker image. 
 - `async-1.32.1`, `async-http-0.64.2`, and `async-pool-0.10.3`. 
 - The ASan build uses `-fsanitize=address -fno-omit-frame-pointer -g -O0` and has YJIT disabled. 

 ## Expected behavior 

 An incomplete chunked HTTP response should fail its async task with `EOFError`; the VM should not crash while cleaning up the task fiber. 

 ## Actual behavior 

 Ruby exits with a segmentation fault in the fiber stack pool: 

 ```text 
 -- C level backtrace -- 
 fiber_pool_stack_release 
 fiber_stack_release 
 fiber_switch 
 rb_fiber_start 
 ``` 

 One GDB run on a source-built 3.4.10 reached: 

 ```text 
 #0 fiber_pool_stack_release (stack=0xffffc4012ea0) at cont.c:770 
     pool = 0x0 
 #1 fiber_stack_release (fiber=0xffffc4012b00) at cont.c:903 
 #2 fiber_switch (...) at cont.c:2728 
 ``` 

 We first hit this in a production service; the harness below reproduces the same C-level frames using only async-http against a local server, with no application code. 

 ## AddressSanitizer result 

 An ASan build makes the bug deterministic. With 50 in-process request bursts, all five stock 3.4.10 runs report: 

 ```text 
 ERROR: AddressSanitizer: heap-use-after-free 
 READ of size 1 
     #0 fiber_switch /src/ruby-3.4.10/cont.c:2727 
 ... 
 freed by thread T2 here: 
     cont_free /src/ruby-3.4.10/cont.c:1079 
     fiber_free /src/ruby-3.4.10/cont.c:1164 
 ... 
 previously allocated by thread T3 here: 
     fiber_t_alloc /src/ruby-3.4.10/cont.c:1986 
 SUMMARY: AddressSanitizer: heap-use-after-free 
     /src/ruby-3.4.10/cont.c:2727 in fiber_switch 
 ``` 

 Line 2727 of `cont.c` is the new `FIBER_TERMINATED_P(fiber)` condition (see Regression candidate below). Thus the first invalid operation is reading the freed `rb_fiber_t` object's status. It occurs before `fiber_pool_stack_release`; the latter is a downstream manifestation when the stale status reads as terminated and the release branch continues. The full report is attached as `stock-1.log`. 

 ## Minimal reproduction harness 

 The attached harness consists of: 

 - `repro.rb`: local raw HTTP server plus direct async-http client; 
 - `Gemfile` and `Gemfile.lock`: pinned async dependencies; 
 - `Dockerfile` and `run.sh`: repeat the workload on a stock Ruby image. 

 Run: 

 ```sh 
 ./run.sh 3.4.10 20 
 ``` 

 This is probabilistic; successful runs are expected. The harness retains a log for every attempt and reports the first interpreter crash. 

 The workload synchronizes groups of ten failures by default: 

 ```sh 
 FAULT_BARRIER=10 ./run.sh 3.4.10 50 
 ``` 

 Run the same source comparison under ASan (deterministic): 

 ```sh 
 BATCHES=50 ./run-asan.sh 5 
 ``` 

 `./run.sh 3.4.9 20` completes all 20 attempts without a crash; 3.4.9 predates the `148f263a` backport. 

 ## Regression candidate 

 Ruby 3.4.10 includes ruby_3_4 commit [`148f263a`](https://github.com/ruby/ruby/commit/148f263a0d299e4532c753470a282144c9104820), the backport of [`dc1777d`](https://github.com/ruby/ruby/commit/dc1777d01770ab62ec99ff6fa4cf622098f44968) for [Bug #21955](https://bugs.ruby-lang.org/issues/21955). 

 That changes `fiber_switch` as follows: 

 ```diff 
 -      if (resuming_fiber && FIBER_TERMINATED_P(fiber)) { 
 +      if (FIBER_TERMINATED_P(fiber)) { 
          fiber_stack_release(fiber); 
      } 
 ``` 

 `Fiber#transfer` passes a null `resuming_fiber`, so the old condition short-circuited without reading `fiber->status`. The new condition dereferences the stale `fiber` pointer before deciding whether to release its stack. 

 We have not tested ruby-head/3.5, but the same condition is present there since `dc1777d`, so we expect it to be affected. 

 ## Controlled source comparison 

 The harness also includes: 

 ```sh 
 BATCHES=20 ./run-ab.sh 20 
 ``` 

 It builds the published Ruby 3.4.10 source twice with the same compiler and workload: 

 1. stock 3.4.10; 
 2. 3.4.10 with only the above condition restored to `resuming_fiber && FIBER_TERMINATED_P(fiber)`. 

 | Build | Attempts | `fiber_pool_stack_release` crashes | 
 | --- | ---: | ---: | 
 | stock 3.4.10 | 20 | 4 | 
 | 3.4.10 with Bug #21955 condition reverted | 20 | 0 | 

 The stock run also had one additional process abort with: 

 ```text 
 malloc(): unaligned tcache chunk detected 
 ``` 

 All 20 reverted-build processes completed successfully. These runs used `BATCHES=20`, 100 requests per batch, concurrency 20, and a barrier of 10 simultaneous truncated responses. The image is built from the Ruby 3.4.10 release tarball with `-O0 -g3 -fno-omit-frame-pointer`; the only source difference between the two images is the condition shown above. 

 The ASan A/B uses `BATCHES=50`: all 5 stock runs report the above heap-use-after-free, while all 5 reverted runs complete cleanly. 

 ## Request 

 Could the fiber maintainers review why a `rb_fiber_t` can be finalized by GC while `fiber_switch` still holds it as the outgoing fiber? Reinstating the `resuming_fiber` guard avoids the dereference on the transfer path, but a proper fix likely needs to ensure that the outgoing fiber remains live through the switch or otherwise avoid reading it after its lifetime ends. 

Back