Project

General

Profile

Actions

Bug #20526

open

File.open(encoding: "bom|utf-8") converts "\r\n" to "\n" on Windows

Bug #20526: File.open(encoding: "bom|utf-8") converts "\r\n" to "\n" on Windows

Added by kou (Kouhei Sutou) over 2 years ago. Updated 13 days ago.

Status:
Open
Assignee:
Target version:
-
ruby -v:
ruby 3.2.2 (2023-03-30 revision e51014f9c0) [x64-mingw-ucrt]
[ruby-core:118182]
Tags:

Description

I'm not sure whether this is an intentional behavior or not but it seems that encoding: "utf-8" doesn't change newline conversion but encoding: "bom|utf-8" changes newline conversion:

File.write("a.txt", "a\r\n")
File.read("a.txt").bytes # => [97, 13, 10]
File.open("a.txt", encoding: "utf-8") {|f| f.read.bytes} # => [97, 10, 10]
File.open("a.txt", encoding: "bom|utf-8") {|f| f.read.bytes} # => [97, 10] XXX: \r\n -> \n
File.open("a.txt", encoding: "bom|utf-8", universal_newline: false) {|f| f.read.bytes} # => [97, 13, 10]

Note that the XXX: line in the above codes. Is this an intentional behavior?

Updated by nobu (Nobuyoshi Nakada) over 2 years ago Actions #1 [ruby-core:118183]

Probably a bug at push back after BOM look ahead.

BTW, on Windows, File.write and File.read are in text mode by default.
That file would be 4 bytes, "a\r\r\n" in binary.

Updated by nobu (Nobuyoshi Nakada) over 2 years ago Actions #2

  • Backport changed from 3.1: UNKNOWN, 3.2: UNKNOWN, 3.3: UNKNOWN to 3.1: REQUIRED, 3.2: REQUIRED, 3.3: REQUIRED

Updated by kou (Kouhei Sutou) over 2 years ago Actions #3

  • Description updated (diff)

Updated by hsbt (Hiroshi SHIBATA) about 2 years ago Actions #4

  • Target version deleted (3.2)

Updated by YO4 (Yoshinao Muramatsu) over 1 year ago Actions #5 [ruby-core:120417]

There are similar strangeness around an encoding specifiers.

preparations

RUBY_VERSION # => "3.3.5"
File.write("a.txt", "a\r\n")
File.binread("a.txt").bytes # => [97, 13, 13, 10]

experimentations

File.open("a.txt")                         {|f| f.read.bytes} # => [97, 13, 10] # expected(msvcrt[_*] newline)
File.open("a.txt", "r:utf-8")              {|f| f.read.bytes} # => [97, 13, 10] # expected
File.open("a.txt", "r", encoding: "utf-8") {|f| f.read.bytes} # => [97, 13, 10] # expected
File.open("a.txt", encoding: "utf-8")      {|f| f.read.bytes} # => [97, 10, 10] # XXX: universal newline enabled?

The omission of the mode parameter seems to enable universal newline.

File.open("a.txt", "rt:utf-8")                  {|f| f.read.bytes} # => [97, 10, 10] # expected(universal newline)
File.open("a.txt", "rt:bom|utf-8")              {|f| f.read.bytes} # => [97, 10] # XXX
File.open("a.txt", "rt", encoding: "utf-8")     {|f| f.read.bytes} # => [97, 10, 10] # expected(universal newline)
File.open("a.txt", "rt", encoding: "bom|utf-8") {|f| f.read.bytes} # => [97, 10] # XXX

XXX: This is odd because universal newline and msvcrt newline appear to be cooperating.

Updated by hsbt (Hiroshi SHIBATA) over 1 year ago Actions #6

  • Tags set to windows

Updated by hsbt (Hiroshi SHIBATA) over 1 year ago Actions #7

  • Tags changed from windows to win

Updated by hsbt (Hiroshi SHIBATA) about 1 month ago Actions #8

  • Assignee set to windows

Updated by YO4 (Yoshinao Muramatsu) 16 days ago Actions #9 [ruby-core:126454]

When a file is opened with encoding: "bom|utf-8", it appears that the fd holds the FTEXT flag as of the point where io_strip_bom() is called.

File.write("a.txt", "a\r\n") writes "a\r\r\n".
Then rb_io_getbyte is called from io_strip_bom, fill rbuf with crlf conversion via _read(). As a result rbuf holds "\a\r\n".
Newline conversion makes this into "a\n" # => [97, 10].
IO#rewind makes this clear.

File.write("a.txt", "a\r\n")
File.open("a.txt", encoding: "bom|utf-8") {|f| f.read.bytes} # => [97, 10]
File.open("a.txt", encoding: "bom|utf-8") {|f| f.rewind; f.read.bytes;} # => [97, 10, 10]
File.open("a.txt", encoding: "bom|utf-8", universal_newline: false) {|f| f.rewind; f.read.bytes} # => [97, 13, 13, 10]

Strangely, encoding: "utf-8" and encoding: "bom:utf-8" seem to enable universal newline conversion.
This affects newline conversion and the handling of Ctrl-Z character.

File.write("a.txt", "a\r\n\x1a\r")
File.open("a.txt", encoding: "utf-8") {|f| f.read.bytes} # => [97, 10, 10, 26, 10]
File.open("a.txt", "r:utf-8") {|f| f.read.bytes} # => [97, 13, 10]

Updated by YO4 (Yoshinao Muramatsu) 13 days ago Actions #10 [ruby-core:126475]

This issue is caused by a combination of two problems.

The first is that universal newline is incorrectly enabled when calling open(name, encoding: "utf-8").
The second is that, in io.c, io_strip_bom() uses rb_io_getbyte(), which is affected by Bug #21691.
As a result, the rbuf contains data that has been converted to CRLF by _read(), and converting this data using universal newline causes the issue.

Fixing the first problem will make open(name, encoding: "utf-8") work correctly, but the same problem will remain for open(name, encoding: "utf-8", universal_newline: true).

To address the root cause of the second, Bug #21691, I have proposed Bug #21765.
The first problem remains the core task of this issue.

Actions

Also available in: PDF Atom