Project

General

Profile

Actions

Bug #22406

open

Process.fork can leave the child unable to resolve hostnames: getaddrinfo helper threads take glibc's resolver lock after releasing the fork lock

Bug #22406: Process.fork can leave the child unable to resolve hostnames: getaddrinfo helper threads take glibc's resolver lock after releasing the fork lock

Added by dug (Doug E) about 1 hour ago. Updated 25 minutes ago.

Status:
Open
Assignee:
-
Target version:
-
ruby -v:
ruby 4.0.7 (2026-09-15 revision 229531a6cf) +PRISM [aarch64-linux]
[ruby-core:126921]

Description

On Linux with glibc, a process forked while another thread's getaddrinfo helper pthread is exiting can inherit glibc's resolver-configuration lock in the locked state. Every hostname lookup in the child then blocks forever. Feature #20590 keeps fork out of the getaddrinfo(3) call itself, but the helper releases that protection before its thread exit takes the same lock.

This is not the macOS problem in #21790, #21876 and #21969. Those are a crash in Apple's resolver that every child hits after any lookup in the parent. This one is on Linux with glibc. It is a race with another thread's lookup, and the child waits on a lock forever rather than crashing.

Reproduce process

Reproduced with each of these, in the official ruby:<version>-slim images (Debian, glibc 2.41, aarch64):

ruby 4.0.7 (2026-09-15 revision 229531a6cf) +PRISM [aarch64-linux]
ruby 3.4.11 (2026-09-23 revision 592f1ffdb3) +PRISM [aarch64-linux]
ruby 4.0.6 (2026-07-14 revision 03b6d3f889) +PRISM [aarch64-linux]
ruby 3.4.10 (2026-06-30 revision 2b0b7728dc) +PRISM [aarch64-linux]

master (dbabf7d789) has the same code: rb_thread_prevent_fork still releases the fork lock before the helper returns, and the helpers are still wrapped at raddrinfo.c L552 and L3193.

No network is needed. The background lookups and the probe both resolve localhost from /etc/hosts.

require "socket"

# Threads that keep resolving a hostname, as an HTTP poller in the parent would.
4.times do
  Thread.new do
    loop do
      Addrinfo.getaddrinfo("localhost", 80, nil, :STREAM)
      Thread.pass
    end
  end
end
sleep 0.5

hung = 0
20_000.times do
  reader, writer = IO.pipe
  pid = fork do
    reader.close
    lookup = Thread.new { Addrinfo.getaddrinfo("localhost", 80, :INET, :STREAM) }
    writer.write(lookup.join(2) ? "ok" : "hung")
    exit!(0)
  end
  writer.close
  hung += 1 if reader.read == "hung"
  reader.close
  Process.wait(pid)
end
puts "#{RUBY_DESCRIPTION}: #{hung} of 20000 children could not resolve localhost"

As a control, change the background threads to resolve "127.0.0.1". That address never reaches getaddrinfo(3), so no helper thread is started.

Result

Runs on aarch64 Linux in Docker:

Ruby Background threads resolve Children that could not resolve, per run of 20,000 forks
4.0.7 localhost 6, 3
4.0.7 127.0.0.1 (control) 0
3.4.11 localhost 3, 2
3.4.11 127.0.0.1 (control) 0
4.0.6 localhost 10, 3, 8, 2
4.0.6 127.0.0.1 (control) 0
3.4.10 localhost 1, 4, 4, 1
3.4.10 127.0.0.1 (control) 0

This is the stack of the stuck helper thread in a hung child, from eu-stack with libc6-dbg installed (3.4.10, glibc 2.41):

#0  __lll_lock_wait_private
#1  get_locked_global
#2  __resolv_conf_get_current
#3  __res_vinit
#4  maybe_init
#5  __resolv_context_get
#6  gethostbyname2_r@@GLIBC_2.17
#7  getaddrinfo
#8  do_getaddrinfo
#9  rb_thread_prevent_fork
#10 start_thread
#11 thread_start

The futex it waits on resolves to libc's static lock (libc+0x1b7dd8 in that build), the lock get_locked_global takes.

Expected result

Every child resolves localhost, as every child does in the control runs.

Feature #20590 made fork wait for in-flight getaddrinfo calls so that a child cannot inherit a resolver lock held by another thread. A child should not inherit one from the part of a helper thread's life that comes after the call either.

Analysis

  • Each helper thread's start routine runs getaddrinfo(3) inside rb_thread_prevent_fork, which holds rb_thread_fork_rw_lock for reading. rb_fork_ruby takes that lock for writing.
  • The read lock is released as soon as getaddrinfo returns. Then the helper thread exits, and glibc's thread exit takes the resolver lock:
    • start_thread calls __libc_thread_freeres,
    • which calls __res_thread_freeres and then __res_iclose(&_res, true),
    • which calls __resolv_conf_detach,
    • which calls get_locked_global, locking the static lock in resolv/resolv_conf.c.
  • glibc's fork resets the malloc, stdio, NSS and dynamic loader locks in the child, but not this one.

So a fork that runs after rb_thread_release_fork_lock() in the helper, and before the helper's thread exit releases the resolver lock, copies that lock held by a thread that does not exist in the child.

The only resolver callers in the reproduction are Ruby's own helper threads, and the fork lock keeps fork out of their getaddrinfo calls. So the lock the child inherited was held by a helper thread that was exiting.

The window is probably hit more often than its size suggests (inference). When Process.fork arrives while a helper is inside getaddrinfo, rb_thread_acquire_fork_lock blocks. The helper's pthread_rwlock_unlock then wakes it just as that same helper moves on to exit and take the resolver lock.

Ruby keeps the waiting Ruby thread interruptible, so Timeout and resolv_timeout still fire. Lookups of IP literals still work, because they never reach getaddrinfo(3). So in practice every new connection by hostname times out for the life of the process.

Source:

Impact

This affects preforking servers whose parent process runs background threads that connect by hostname.

We found it in a Pitchfork deployment. A feature-flag SDK in the parent polls an HTTP endpoint every 30 seconds, and roughly one in tens of thousands of forked workers comes up unable to open any connection by hostname. Pitchfork's documentation for Pitchfork.prevent_fork warns about getaddrinfo(3) in background threads.

Possible directions

These are not patches, just the options I considered:

  • Wait for helpers to finish exiting. Keep fork out until helper threads have exited, not only until they have finished getaddrinfo. For example, fork could join helpers that have returned but not yet exited, though today the helpers are detached.
  • Reuse resolver threads. Keep a small pool instead of exiting after each lookup, so the detach that runs at thread exit happens only at process exit.

glibc could also reset this lock in the child, as it does the NSS locks. But glibc's position is that only async-signal-safe functions are supported in the child of a multithreaded process.

Workaround

  • Avoid hostname lookups in background threads of a process that forks, or wrap them so the fork waits until the thread is fully idle.
  • To contain the damage, check resolution in the child after fork and exit if it hangs. For example, run Addrinfo.getaddrinfo("localhost", 80, :INET, :STREAM) in a thread and join it with a short timeout.

cc @byroot (Jean Boussier) (Feature #20590), @mame (Yusuke Endoh) (Feature #19965), @shioimm (Misaki Shioi) (the Happy Eyeballs getaddrinfo helper)

Updated by dug (Doug E) 25 minutes ago Actions #1

  • Description updated (diff)
Actions

Also available in: PDF Atom