Project

General

Profile

Actions

Bug #22423

closed

String#initialize(encoding:) stores an inflated capacity and writes past the malloc'd buffer on append (rb_str_init termlen double-adjust)

Bug #22423: String#initialize(encoding:) stores an inflated capacity and writes past the malloc'd buffer on append (rb_str_init termlen double-adjust)

Added by sushie (Bug Bounty) 1 day ago. Updated about 6 hours ago.

Status:
Closed
Assignee:
-
Target version:
-
ruby -v:
ruby 4.0.7 (2026-09-15 revision 229531a6cf) +PRISM [arm64-darwin25]
[ruby-core:127061]

Description

Description

String#initialize with encoding: and capacity: keywords stores a capacity
that is LARGER than the buffer actually allocated, so later in-place growth
(<< / concat) writes past the end of the malloc'd region.

Root cause: in rb_str_init (string.c) the buffer is sized with the NEW
encoding's terminator length, the source bytes are copied, and then
rb_enc_associate(str, enc) re-adjusts aux.capa against the OLD encoding's
terminator length (via rb_str_change_terminator_length). When the source
string's encoding has a longer terminator (UTF-16/32 -> UTF-8), the stored
capacity ends up larger than the allocation.

This is the same construct that was fixed in str_shared_replace by commit
323c1df196 ("[Bug #22387] Fix termlen in str_shared_replace"); rb_str_init
has the identical pattern and is still unchanged on master (verified
byte-identical between the v4_0_7 tag and current master).

Reproducer

Uses only the fiddle stdlib, which ships with the default distribution
(the library path below is the homebrew macOS layout; adjust to your install):

require "fiddle"

h = Fiddle::Handle.new("/opt/homebrew/opt/ruby/lib/libruby.4.0.dylib")
$rb_str_capacity = Fiddle::Function.new(h["rb_str_capacity"], [Fiddle::TYPE_VOIDP], Fiddle::TYPE_SIZE_T)
def capa(s) = $rb_str_capacity.call(Fiddle.dlwrap(s))

S = Class.new(String)
orig   = ("\u{30AF}" * 5).encode("UTF-16LE")   # 10 bytes, 2-byte terminator
orig32 = ("\u{30AF}" * 5).encode("UTF-32LE")   # 20 bytes, 4-byte terminator

a = S.new(orig,   encoding: "UTF-8", capacity: 200)
b = S.new(orig32, encoding: "UTF-8", capacity: 200)
puts "UTF-16LE -> UTF-8 : capa=#{capa(a)}"
puts "UTF-32LE -> UTF-8 : capa=#{capa(b)}"
a << ("x" * 191)
puts "after append 191B : len=#{a.bytesize} capa=#{capa(a)}"

Actual output (4.0.7; identical behavior on master):

UTF-16LE -> UTF-8 : capa=201
UTF-32LE -> UTF-8 : capa=203
after append 191B : len=201 capa=201

The buffer is allocated as capacity + termlen(UTF-8) = 201 bytes, but the
stored capacity is 201 (UTF-16LE input) / 203 (UTF-32LE input). Appending 191
bytes grows the string to 201 bytes, which does not exceed the stored
capacity, so no reallocation happens and TERM_FILL writes the NUL terminator
at index 201 - one byte past the 201-byte malloc region (up to 3 bytes past
for the UTF-32LE case).

Worst case: capacity: 207 allocates 208 bytes (207 + 1), stores capacity
210, and appending 207 bytes to the 20 existing content bytes writes through
index 227 - 19 bytes past the allocation, into adjacent heap memory.

Expected behavior

rb_str_capacity should never exceed the allocation (capacity + termlen(new encoding)), and in-place growth must reallocate before writing past the
buffer.

Proposed fix

In rb_str_init, associate the new encoding with rb_enc_raw_set instead of
rb_enc_associate, so the capacity is not re-adjusted against the old
terminator length (mirroring 323c1df196); alternatively reorder so the buffer
is sized against the encoding that is actually set. Suggested change in
string.c rb_str_init:

rb_enc_raw_set(str, enc);

instead of rb_enc_associate(str, enc).

Affected versions

All releases up to and including 4.0.7, and current master.

Note

Reported initially through the Ruby HackerOne program; filing here per the
maintainers' suggestion so the fix can land and be backported.

Updated by sushie (Bug Bounty) 1 day ago Actions #1 [ruby-core:127062]

A proposed patch was shared by nobu in the HackerOne thread for this issue: it prevents rb_enc_cr_str_exact_copy from copying the old encoding's metadata when the encoding is being changed, and uses rb_enc_raw_set in rb_str_init to avoid the terminator-length recalculation (patch by nobu, to be attached by the maintainers).

Updated by nobu (Nobuyoshi Nakada) about 7 hours ago Actions #2

  • Backport changed from 3.3: UNKNOWN, 3.4: UNKNOWN, 4.0: UNKNOWN to 3.3: REQUIRED, 3.4: REQUIRED, 4.0: REQUIRED

Updated by nobu (Nobuyoshi Nakada) about 6 hours ago Actions #3

  • Status changed from Open to Closed

Applied in changeset git|2d08b5f42f2cbed4a4e6ecf8d62987a4f21f98c1.


[Bug #22423] Preserve encoding with String capacity

Actions

Also available in: PDF Atom