Bug #22390
openruby 3.4 -> 4.0 upgrade GC changes causing >2x RSS increase
Description
Hello, I've been trying to upgrade a fairly sizeable monolith, which has been running on ruby 3.4.10 in production, to ruby 4 (have been targeting 4.0.6 more specifically, although the below should apply the same to 4.0.7). While we didn't experience anything out of the ordinary in staging, unfortunately, we started seeing our more intensive workloads (job processors) grow in memory usage over the first 40m of running on ruby 4, up to the point where the servers would hit OOM, at which point we decided to rollback and investigate. This workload calls Process.warmup after booting the application and then fork multiple processes (not sure if this information is relevant).
We figured out that the culprit of this change in GC behaviour was this commit from @peter.zhu; this commit actually another bug for which we have a mitigation in ruby 3.4, however it changed quite a few more things in the heuristics used to decide when to give empty pages back to the OS.
Before the commit, heap_pages_free_unused_pages used objspace->heap_pages.allocatable_slots to decide, and this number would be reset to 0 on every sweep, and grow when expanding the heap, which resulted in every empty page being freed on every major GC.
After the commit, this changed to using heap_pages_freeable_pages, which gets computed on every major GC, and matched against GC_HEAP_FREE_SLOTS_MAX_RATIO. In our case at least, this doesn't work out, because in the aforementioned workload, objspace_available_slots grows at the same pace as heap_pages_freeable_pages; this makes it so that the 65% max ratio is never reached for this type of steadily churning live sets; by the time the size pools settle, the remainder of empty pages is never given back. I also tried to pull claude to figure some of this out, and it mentioned that the per pool 20% reservation alongside the single (per process) empty pages management somehow plays a role, but I couldn't make sense of whether that was legit or just an hallucination.
I'm attaching a script which I coded with some AI help to roughly mimic the objects being created per job iteration. Bear in mind that, in our case in production, we're talking about many more (OOM hits at ~5Gb, and we've seen the ruby 4 builds grow up to 4Gb before we rolled back), but this represents the type of objects created in a way that makes each of the pools grow independently. If you run it locally, you should see the ruby 4 example use ~2x the memory, mostly because of the empty pages being kept around; major gcs will sometimes free them back to the OS, but I suspect that the reason for it was some other condition triggering a major GC (oldmalloc perhaps?) and causing heap_pages_freeable_pages to be recalculated as a result.
We continue experimenting with the GC knobs to see if we can find a combination which would allow us to run ruby 4 in production. We're looking at pre-allocating a bunch of slots in pools 1-5 (as per upper bounds from observations using GC.stat_heap), while lowering GC_HEAP_OLDOBJECT_LIMIT_FACTOR to force major GC to run more often (a tad more, not too much, not to hurt latency) and hopefully updating freeable pages stats, while also lowering GC_HEAP_FREE_SLOTS_MAX_RATIO as well to a value that makes more sense. If anyone has suggestions on how to further get to optimal values, I'm all ears. Still, while I accept that special workloads require special tuning, I don't find this one unusual enough to justify this type of memory usage increase just due to a ruby upgrade, hence why I file this as a bug.
Files
No data to display