Bug #22406
openProcess.fork can leave the child unable to resolve hostnames: getaddrinfo helper threads take glibc's resolver lock after releasing the fork lock
Description
On Linux with glibc, a process forked while another thread's getaddrinfo helper pthread is exiting can inherit glibc's resolver-configuration lock in the locked state. Every hostname lookup in the child then blocks forever. Feature #20590 keeps fork out of the getaddrinfo(3) call itself, but the helper releases that protection before its thread exit takes the same lock.
This is not the macOS problem in #21790, #21876 and #21969. Those are a crash in Apple's resolver that every child hits after any lookup in the parent. This one is on Linux with glibc. It is a race with another thread's lookup, and the child waits on a lock forever rather than crashing.
Reproduce process¶
Reproduced with each of these, in the official ruby:<version>-slim images (Debian, glibc 2.41, aarch64):
ruby 4.0.7 (2026-09-15 revision 229531a6cf) +PRISM [aarch64-linux]
ruby 3.4.11 (2026-09-23 revision 592f1ffdb3) +PRISM [aarch64-linux]
ruby 4.0.6 (2026-07-14 revision 03b6d3f889) +PRISM [aarch64-linux]
ruby 3.4.10 (2026-06-30 revision 2b0b7728dc) +PRISM [aarch64-linux]
master (dbabf7d789) has the same code: rb_thread_prevent_fork still releases the fork lock before the helper returns, and the helpers are still wrapped at raddrinfo.c L552 and L3193.
No network is needed. The background lookups and the probe both resolve localhost from /etc/hosts.
require "socket"
# Threads that keep resolving a hostname, as an HTTP poller in the parent would.
4.times do
Thread.new do
loop do
Addrinfo.getaddrinfo("localhost", 80, nil, :STREAM)
Thread.pass
end
end
end
sleep 0.5
hung = 0
20_000.times do
reader, writer = IO.pipe
pid = fork do
reader.close
lookup = Thread.new { Addrinfo.getaddrinfo("localhost", 80, :INET, :STREAM) }
writer.write(lookup.join(2) ? "ok" : "hung")
exit!(0)
end
writer.close
hung += 1 if reader.read == "hung"
reader.close
Process.wait(pid)
end
puts "#{RUBY_DESCRIPTION}: #{hung} of 20000 children could not resolve localhost"
As a control, change the background threads to resolve "127.0.0.1". That address never reaches getaddrinfo(3), so no helper thread is started.
Result¶
Runs on aarch64 Linux in Docker:
| Ruby | Background threads resolve | Children that could not resolve, per run of 20,000 forks |
|---|---|---|
| 4.0.7 | localhost |
6, 3 |
| 4.0.7 | 127.0.0.1 (control) |
0 |
| 3.4.11 | localhost |
3, 2 |
| 3.4.11 | 127.0.0.1 (control) |
0 |
| 4.0.6 | localhost |
10, 3, 8, 2 |
| 4.0.6 | 127.0.0.1 (control) |
0 |
| 3.4.10 | localhost |
1, 4, 4, 1 |
| 3.4.10 | 127.0.0.1 (control) |
0 |
This is the stack of the stuck helper thread in a hung child, from eu-stack with libc6-dbg installed (3.4.10, glibc 2.41):
#0 __lll_lock_wait_private
#1 get_locked_global
#2 __resolv_conf_get_current
#3 __res_vinit
#4 maybe_init
#5 __resolv_context_get
#6 gethostbyname2_r@@GLIBC_2.17
#7 getaddrinfo
#8 do_getaddrinfo
#9 rb_thread_prevent_fork
#10 start_thread
#11 thread_start
The futex it waits on resolves to libc's static lock (libc+0x1b7dd8 in that build), the lock get_locked_global takes.
Expected result¶
Every child resolves localhost, as every child does in the control runs.
Feature #20590 made fork wait for in-flight getaddrinfo calls so that a child cannot inherit a resolver lock held by another thread. A child should not inherit one from the part of a helper thread's life that comes after the call either.
Analysis¶
- Each helper thread's start routine runs
getaddrinfo(3)insiderb_thread_prevent_fork, which holdsrb_thread_fork_rw_lockfor reading.rb_fork_rubytakes that lock for writing. - The read lock is released as soon as
getaddrinforeturns. Then the helper thread exits, and glibc's thread exit takes the resolver lock:start_threadcalls__libc_thread_freeres,- which calls
__res_thread_freeresand then__res_iclose(&_res, true), - which calls
__resolv_conf_detach, - which calls
get_locked_global, locking the staticlockinresolv/resolv_conf.c.
- glibc's
forkresets the malloc, stdio, NSS and dynamic loader locks in the child, but not this one.
So a fork that runs after rb_thread_release_fork_lock() in the helper, and before the helper's thread exit releases the resolver lock, copies that lock held by a thread that does not exist in the child.
The only resolver callers in the reproduction are Ruby's own helper threads, and the fork lock keeps fork out of their getaddrinfo calls. So the lock the child inherited was held by a helper thread that was exiting.
The window is probably hit more often than its size suggests (inference). When Process.fork arrives while a helper is inside getaddrinfo, rb_thread_acquire_fork_lock blocks. The helper's pthread_rwlock_unlock then wakes it just as that same helper moves on to exit and take the resolver lock.
Ruby keeps the waiting Ruby thread interruptible, so Timeout and resolv_timeout still fire. Lookups of IP literals still work, because they never reach getaddrinfo(3). So in practice every new connection by hostname times out for the life of the process.
Source:
- Ruby at v3_4_11; v4.0.7 and master have the same structure, see the master links above:
- glibc 2.41, the version in the reproduction. 2.35 has the same code on these paths.
- nptl/pthread_create.c L459:
start_threadcalls__libc_thread_freeres. - malloc/thread-freeres.c L35: that calls
__res_thread_freeres. - resolv/res-close.c L129-L142 and L117:
__res_iclose(&_res, true)calls__resolv_conf_detach. - resolv/resolv_conf.c L81 defines
lock. L648 takes it at thread exit, and L130 is where the child waits. - posix/fork.c L88-L108: the locks fork resets in the child.
- nptl/pthread_create.c L459:
Impact¶
This affects preforking servers whose parent process runs background threads that connect by hostname.
We found it in a Pitchfork deployment. A feature-flag SDK in the parent polls an HTTP endpoint every 30 seconds, and roughly one in tens of thousands of forked workers comes up unable to open any connection by hostname. Pitchfork's documentation for Pitchfork.prevent_fork warns about getaddrinfo(3) in background threads.
Possible directions¶
These are not patches, just the options I considered:
- Wait for helpers to finish exiting. Keep fork out until helper threads have exited, not only until they have finished
getaddrinfo. For example, fork could join helpers that have returned but not yet exited, though today the helpers are detached. - Reuse resolver threads. Keep a small pool instead of exiting after each lookup, so the detach that runs at thread exit happens only at process exit.
glibc could also reset this lock in the child, as it does the NSS locks. But glibc's position is that only async-signal-safe functions are supported in the child of a multithreaded process.
Workaround¶
- Avoid hostname lookups in background threads of a process that forks, or wrap them so the fork waits until the thread is fully idle.
- To contain the damage, check resolution in the child after fork and exit if it hangs. For example, run
Addrinfo.getaddrinfo("localhost", 80, :INET, :STREAM)in a thread and join it with a short timeout.
cc @byroot (Jean Boussier) (Feature #20590), @mame (Yusuke Endoh) (Feature #19965), @shioimm (Misaki Shioi) (the Happy Eyeballs getaddrinfo helper)
Updated by dug (Doug E) 25 minutes ago
- Description updated (diff)