Bug #22423
closedString#initialize(encoding:) stores an inflated capacity and writes past the malloc'd buffer on append (rb_str_init termlen double-adjust)
Description
Description¶
String#initialize with encoding: and capacity: keywords stores a capacity
that is LARGER than the buffer actually allocated, so later in-place growth
(<< / concat) writes past the end of the malloc'd region.
Root cause: in rb_str_init (string.c) the buffer is sized with the NEW
encoding's terminator length, the source bytes are copied, and then
rb_enc_associate(str, enc) re-adjusts aux.capa against the OLD encoding's
terminator length (via rb_str_change_terminator_length). When the source
string's encoding has a longer terminator (UTF-16/32 -> UTF-8), the stored
capacity ends up larger than the allocation.
This is the same construct that was fixed in str_shared_replace by commit
323c1df196 ("[Bug #22387] Fix termlen in str_shared_replace"); rb_str_init
has the identical pattern and is still unchanged on master (verified
byte-identical between the v4_0_7 tag and current master).
Reproducer¶
Uses only the fiddle stdlib, which ships with the default distribution
(the library path below is the homebrew macOS layout; adjust to your install):
require "fiddle"
h = Fiddle::Handle.new("/opt/homebrew/opt/ruby/lib/libruby.4.0.dylib")
$rb_str_capacity = Fiddle::Function.new(h["rb_str_capacity"], [Fiddle::TYPE_VOIDP], Fiddle::TYPE_SIZE_T)
def capa(s) = $rb_str_capacity.call(Fiddle.dlwrap(s))
S = Class.new(String)
orig = ("\u{30AF}" * 5).encode("UTF-16LE") # 10 bytes, 2-byte terminator
orig32 = ("\u{30AF}" * 5).encode("UTF-32LE") # 20 bytes, 4-byte terminator
a = S.new(orig, encoding: "UTF-8", capacity: 200)
b = S.new(orig32, encoding: "UTF-8", capacity: 200)
puts "UTF-16LE -> UTF-8 : capa=#{capa(a)}"
puts "UTF-32LE -> UTF-8 : capa=#{capa(b)}"
a << ("x" * 191)
puts "after append 191B : len=#{a.bytesize} capa=#{capa(a)}"
Actual output (4.0.7; identical behavior on master):
The buffer is allocated as capacity + termlen(UTF-8) = 201 bytes, but the
stored capacity is 201 (UTF-16LE input) / 203 (UTF-32LE input). Appending 191
bytes grows the string to 201 bytes, which does not exceed the stored
capacity, so no reallocation happens and TERM_FILL writes the NUL terminator
at index 201 - one byte past the 201-byte malloc region (up to 3 bytes past
for the UTF-32LE case).
Worst case: capacity: 207 allocates 208 bytes (207 + 1), stores capacity
210, and appending 207 bytes to the 20 existing content bytes writes through
index 227 - 19 bytes past the allocation, into adjacent heap memory.
Expected behavior¶
rb_str_capacity should never exceed the allocation (capacity + termlen(new encoding)), and in-place growth must reallocate before writing past the
buffer.
Proposed fix¶
In rb_str_init, associate the new encoding with rb_enc_raw_set instead of
rb_enc_associate, so the capacity is not re-adjusted against the old
terminator length (mirroring 323c1df196); alternatively reorder so the buffer
is sized against the encoding that is actually set. Suggested change in
string.c rb_str_init:
instead of rb_enc_associate(str, enc).
Affected versions¶
All releases up to and including 4.0.7, and current master.
Note¶
Reported initially through the Ruby HackerOne program; filing here per the
maintainers' suggestion so the fix can land and be backported.