This project is closed and read-only.
Feature #2350
closedUnicode specific functionality on String in 1.9
Added by Manfred (Manfred Stienstra) almost 17 years ago. Updated over 15 years ago.
Description
=begin
I was wondering is there are any plans to include Unicode aware methods for Unicode encodings on String? For example, upcase and downcase only handle ASCII characters at the moment.
cafe = "Café"
cafe.encoding # => #Encoding:UTF-8
"Café".upcase # => CAFé
=end
Files
| signature.asc (207 Bytes) signature.asc | Cezary (Cezary Baginski), 03/23/2011 03:23 AM | ||
| signature.asc (207 Bytes) signature.asc | Cezary (Cezary Baginski), 04/12/2011 08:17 PM | ||
| signature.asc (207 Bytes) signature.asc | Cezary (Cezary Baginski), 04/12/2011 08:18 PM |
Updated by matz (Yukihiro Matsumoto) almost 17 years ago
Actions
#1
=begin
Hi,
In message "Re: [ruby-core:26650] [Feature #2350] Unicode specific functionality on String in 1.9"
on Mon, 9 Nov 2009 23:29:42 +0900, Manfred Stienstra redmine@ruby-lang.org writes:
|I was wondering is there are any plans to include Unicode aware methods for Unicode encodings on String? For example, upcase and downcase only handle ASCII characters at the moment.
|
|cafe = "Café"
|cafe.encoding # => #Encoding:UTF-8
|"Café".upcase # => CAFé
As far as I understand, the Unicode case conversion requires
additional language information for e.g. Turkish i. And some
conversion does not round-trip e.g. German SS. Use unicode gem
instead.
=end
Updated by Manfred (Manfred Stienstra) almost 17 years ago
Actions
#2
=begin
Yes, case conversions require the Unicode database and specific locale implementations. Thank you for your answer!
=end
Updated by mame (Yusuke Endoh) over 16 years ago
Actions
#3
- Status changed from Open to Rejected
=begin
Hi,
Matz seemed to reject this ticket, and OP seemed to be satisfied
with matz's answer. So I close the ticket.
--
Yusuke Endoh mame@tsg.ne.jp
=end
Updated by now (Nikolai Weibull) over 16 years ago
Actions
#4
=begin
On Thu, Mar 25, 2010 at 14:45, Yusuke Endoh redmine@ruby-lang.org wrote:
Issue #2350 has been updated by Yusuke Endoh.
Matz seemed to reject this ticket, and OP seemed to be satisfied
with matz's answer. So I close the ticket.
How would I be able to hook in my character-encodings library into
Ruby 1.9 Strings? I would like to override, for example, #upcase for
all Strings that have a Unicode encoding. Is this possible?
Thanks!
=end
Updated by naruse (Yui NARUSE) over 16 years ago
Actions
#5
=begin
(2010/03/26 0:02), Nikolai Weibull wrote:
On Thu, Mar 25, 2010 at 14:45, Yusuke Endohredmine@ruby-lang.org wrote:
Issue #2350 has been updated by Yusuke Endoh.
Matz seemed to reject this ticket, and OP seemed to be satisfied
with matz's answer. So I close the ticket.How would I be able to hook in my character-encodings library into
Ruby 1.9 Strings? I would like to override, for example, #upcase for
all Strings that have a Unicode encoding. Is this possible?
You can hook String methods, Ruby doesn't forbid it.
But I think, people want both ASCII version and Unicode version of upcase.
So you should name your Unicode methods another names.
--
NARUSE, Yui naruse@airemix.jp
=end
Updated by now (Nikolai Weibull) over 16 years ago
Actions
#6
=begin
On Thu, Mar 25, 2010 at 18:24, NARUSE, Yui naruse@airemix.jp wrote:
(2010/03/26 0:02), Nikolai Weibull wrote:
On Thu, Mar 25, 2010 at 14:45, Yusuke Endohredmine@ruby-lang.org wrote:
Issue #2350 has been updated by Yusuke Endoh.
Matz seemed to reject this ticket, and OP seemed to be satisfied
with matz's answer. So I close the ticket.How would I be able to hook in my character-encodings library into
Ruby 1.9 Strings? I would like to override, for example, #upcase for
all Strings that have a Unicode encoding. Is this possible?You can hook String methods, Ruby doesn't forbid it.
Yes, I can do something like
class String
def unicodify
extend Encoding::Character::Unicode
end
end
but I was wondering if there was a way to do it without having to do
String.new.unicodify.upcase
But I think, people want both ASCII version and Unicode version of upcase.
So you should name your Unicode methods another names.
Why would they want that? Having an ASCII-only version of #upcase
makes no sense for a Unicode String more than supporting #upcase
requires that you load the Unicode character database information,
which takes up quite a lot of memory.
I want to transparently deal with this kind of thing. I know that the
Ruby way is to be explicit about encodings and I actually like that,
but that’s only something I care about at creation, not when invoking
methods on the String.
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#7
[ruby-core:35522]
=begin
On Thu, Mar 25, 2010 at 19:33, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 18:24, NARUSE, Yui naruse@airemix.jp wrote:
(2010/03/26 0:02), Nikolai Weibull wrote:
I was wondering if there was a way to do it without having to do
String.new.unicodify.upcase
But I think, people want both ASCII version and Unicode version of upcase.
So you should name your Unicode methods another names.
Why would they want that? Having an ASCII-only version of #upcase
makes no sense for a Unicode String more than supporting #upcase
requires that you load the Unicode character database information,
which takes up quite a lot of memory.
So, what’s the reasoning here? Having "äbc".upcase return "äBC" makes
absolutely no sense and means that quite a few methods on String are
completely useless in a m18n context.
=end
Updated by judofyr (Magnus Holm) over 15 years ago
Actions
#8
[ruby-core:35524]
=begin
The problem is that the definition of #upcase doesn't only depend on the
encoding used, but also the language of the encoded text. For instance, if
you're writing in Turkish, you would expect "i".upcase to return a dotted
uppcase I: http://www.i18nguy.com/unicode/turkish-i18n.html
http://www.i18nguy.com/unicode/turkish-i18n.htmlDoing this properly is
really hard and needs to have a lot of flexibility, especially when it
comes to non-Western languages. It's far easier for everyone that the
built-in #upcase is simple and fast and you'll have to be explicit about any
other I18n stuff IMO.
// Magnus Holm
On Fri, Mar 18, 2011 at 11:19, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 19:33, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 18:24, NARUSE, Yui naruse@airemix.jp wrote:
(2010/03/26 0:02), Nikolai Weibull wrote:
I was wondering if there was a way to do it without having to do
String.new.unicodify.upcase
But I think, people want both ASCII version and Unicode version of
upcase.
So you should name your Unicode methods another names.Why would they want that? Having an ASCII-only version of #upcase
makes no sense for a Unicode String more than supporting #upcase
requires that you load the Unicode character database information,
which takes up quite a lot of memory.So, what’s the reasoning here? Having "äbc".upcase return "äBC" makes
absolutely no sense and means that quite a few methods on String are
completely useless in a m18n context.
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#9
[ruby-core:35525]
=begin
On Fri, Mar 18, 2011 at 11:53, Magnus Holm judofyr@gmail.com wrote:
The problem is that the definition of #upcase doesn't only depend on the
encoding used, but also the language of the encoded text. For instance, if
you're writing in Turkish, you would expect "i".upcase to return a dotted
uppcase I: http://www.i18nguy.com/unicode/turkish-i18n.html
I know. The same goes for ‘i’ in Lithuanian.
Doing this properly is really hard and needs to have a lot of flexibility,
especially when it comes to non-Western languages.
This is simply not true. Unicode defines how to deal with case
conversions. I’m not saying that the Unicode standard is infallible,
but we can at least adhere to it. I’m not saying that Unicode is the
only encoding that we should care about, but if we support the Unicode
transfer formats, why not support other interesting parts of the
standard?
It's far easier for everyone that the built-in #upcase is
simple and fast and you'll have to be explicit about any
other I18n stuff IMO.
Easy, perhaps, but hardly useful.
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII. I was looking for an
actual solution to this problem. I have a library
(character-encodings) that does support these conversions, based on
locale and the Unicode character database (UCD). How do we make it
easy for the user to deal with m18n? I mean, if I say
-- coding: utf-8 --¶
puts "äbc".upcase
I expect this to do the right thing for Unicode under the current locale.
As Unicode defines how to deal with case conversions, if I tell Ruby
that “this String is encoded as UTF-8” (or, in this case, “strings in
this file are encoded as UTF-8”), I expect Ruby to respond “OK, I’ll
use the Unicode rules that govern methods like #upcase for that
String”.
The UCD requires a lot of memory, so I suggested that a library, such
as character-encodings, should be able to seamlessly add this kind of
behavior without requiring the user to write "äbc".unicodify.upcase,
if the UCD can’t be included in standard Ruby runtime.
But, come to think of it, doesn’t Oniguruma need most of the UCD
information, so isn’t most of it already included in the Ruby runtime?
Adding casing information perhaps wouldn’t require much additional
space.
If this isn’t of interest, then I’m still looking for a way to
override #upcase for Strings that use the UTF-8 encoding without
resorting to alias_method or extend (as shown earlier in this
discussion). This seems impossible to do at the moment, as Encoding
is a completely opaque object.
=end
Updated by Cezary (Cezary Baginski) over 15 years ago
Actions
#10
[ruby-core:35541]
=begin
On Fri, Mar 18, 2011 at 09:52:27PM +0900, Nikolai Weibull wrote:
On Fri, Mar 18, 2011 at 11:53, Magnus Holm judofyr@gmail.com wrote:
It's far easier for everyone that the built-in #upcase is
simple and fast and you'll have to be explicit about any
other I18n stuff IMO.Easy, perhaps, but hardly useful.
A agree - for human interaction it is completely useless. I tend to
think of #upcase as just a convenience method for dealing with ASCII
only system level functionality, e.g. paths on filesystems,
environment variables, html tags, (un)capitalizing to get class names,
database table names, etc.
Anything else is "no-op" or "undefined" for me.
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII.
I would probably go one step further and disallow upcase and friends
for any non-US-ASCII string for this reason. At least issue a warning.
If this isn’t of interest, then I’m still looking for a way to
override #upcase for Strings that use the UTF-8 encoding without
resorting to alias_method or extend (as shown earlier in this
discussion). This seems impossible to do at the moment, as Encoding
is a completely opaque object.
Correct me if I am wrong, but even "upper case" as a concept is not
common among all languages - an implementation detail for specific
cases at best.
For example, in German, you may want a more meaningful 'to_noun'
instead of 'capitalize'. For Japanese some may want upcase as a no-op
and some as a hack to convert to katakana. For case insensitivity,
probably a "normalize" method would be more descriptive.
Out of curiosity: in what specific case is utf upcase necessary?
--
Cezary Baginski
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#11
[ruby-core:35549]
=begin
On Tue, Mar 22, 2011 at 18:30, Cezary cezary.baginski@gmail.com wrote:
On Fri, Mar 18, 2011 at 09:52:27PM +0900, Nikolai Weibull wrote:
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII.
I would probably go one step further and disallow upcase and friends
for any non-US-ASCII string for this reason. At least issue a warning.
For Unicode there actually are well-defined casing rules.
For example, in German, you may want a more meaningful 'to_noun'
instead of 'capitalize'. For Japanese some may want upcase as a no-op
and some as a hack to convert to katakana. For case insensitivity,
probably a "normalize" method would be more descriptive.
This is perhaps true, but beside the point.
Out of curiosity: in what specific case is utf upcase necessary?
That’s a good question. It’s perhaps not a common operation, but text
editors and regular expression engines most likely need it. Even if
their utility is limited, returning incorrect results is worse.
=end
Updated by carl.hoerberg (Carl Hörberg) over 15 years ago
Actions
#12
[ruby-core:35626]
=begin
We need it to allow class names in foreign languages. Today "Åtgärd" ain't recognized as a constant, and there for can't be uses a class name.
=end
Updated by naruse (Yui NARUSE) over 15 years ago
Actions
#13
[ruby-core:35660]
=begin
About String#upcase, our current answer is simple: use ICU.
https://github.com/jarib/ffi-icu
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#14
[ruby-core:35522]
=begin
On Thu, Mar 25, 2010 at 19:33, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 18:24, NARUSE, Yui naruse@airemix.jp wrote:
(2010/03/26 0:02), Nikolai Weibull wrote:
I was wondering if there was a way to do it without having to do
String.new.unicodify.upcase
But I think, people want both ASCII version and Unicode version of upcase.
So you should name your Unicode methods another names.
Why would they want that? Having an ASCII-only version of #upcase
makes no sense for a Unicode String more than supporting #upcase
requires that you load the Unicode character database information,
which takes up quite a lot of memory.
So, what’s the reasoning here? Having "äbc".upcase return "äBC" makes
absolutely no sense and means that quite a few methods on String are
completely useless in a m18n context.
=end
Updated by judofyr (Magnus Holm) over 15 years ago
Actions
#15
[ruby-core:35524]
=begin
The problem is that the definition of #upcase doesn't only depend on the
encoding used, but also the language of the encoded text. For instance, if
you're writing in Turkish, you would expect "i".upcase to return a dotted
uppcase I: http://www.i18nguy.com/unicode/turkish-i18n.html
http://www.i18nguy.com/unicode/turkish-i18n.htmlDoing this properly is
really hard and needs to have a lot of flexibility, especially when it
comes to non-Western languages. It's far easier for everyone that the
built-in #upcase is simple and fast and you'll have to be explicit about any
other I18n stuff IMO.
// Magnus Holm
On Fri, Mar 18, 2011 at 11:19, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 19:33, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 18:24, NARUSE, Yui naruse@airemix.jp wrote:
(2010/03/26 0:02), Nikolai Weibull wrote:
I was wondering if there was a way to do it without having to do
String.new.unicodify.upcase
But I think, people want both ASCII version and Unicode version of
upcase.
So you should name your Unicode methods another names.Why would they want that? Having an ASCII-only version of #upcase
makes no sense for a Unicode String more than supporting #upcase
requires that you load the Unicode character database information,
which takes up quite a lot of memory.So, what’s the reasoning here? Having "äbc".upcase return "äBC" makes
absolutely no sense and means that quite a few methods on String are
completely useless in a m18n context.
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#16
[ruby-core:35525]
=begin
On Fri, Mar 18, 2011 at 11:53, Magnus Holm judofyr@gmail.com wrote:
The problem is that the definition of #upcase doesn't only depend on the
encoding used, but also the language of the encoded text. For instance, if
you're writing in Turkish, you would expect "i".upcase to return a dotted
uppcase I: http://www.i18nguy.com/unicode/turkish-i18n.html
I know. The same goes for ‘i’ in Lithuanian.
Doing this properly is really hard and needs to have a lot of flexibility,
especially when it comes to non-Western languages.
This is simply not true. Unicode defines how to deal with case
conversions. I’m not saying that the Unicode standard is infallible,
but we can at least adhere to it. I’m not saying that Unicode is the
only encoding that we should care about, but if we support the Unicode
transfer formats, why not support other interesting parts of the
standard?
It's far easier for everyone that the built-in #upcase is
simple and fast and you'll have to be explicit about any
other I18n stuff IMO.
Easy, perhaps, but hardly useful.
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII. I was looking for an
actual solution to this problem. I have a library
(character-encodings) that does support these conversions, based on
locale and the Unicode character database (UCD). How do we make it
easy for the user to deal with m18n? I mean, if I say
-- coding: utf-8 --¶
puts "äbc".upcase
I expect this to do the right thing for Unicode under the current locale.
As Unicode defines how to deal with case conversions, if I tell Ruby
that “this String is encoded as UTF-8” (or, in this case, “strings in
this file are encoded as UTF-8”), I expect Ruby to respond “OK, I’ll
use the Unicode rules that govern methods like #upcase for that
String”.
The UCD requires a lot of memory, so I suggested that a library, such
as character-encodings, should be able to seamlessly add this kind of
behavior without requiring the user to write "äbc".unicodify.upcase,
if the UCD can’t be included in standard Ruby runtime.
But, come to think of it, doesn’t Oniguruma need most of the UCD
information, so isn’t most of it already included in the Ruby runtime?
Adding casing information perhaps wouldn’t require much additional
space.
If this isn’t of interest, then I’m still looking for a way to
override #upcase for Strings that use the UTF-8 encoding without
resorting to alias_method or extend (as shown earlier in this
discussion). This seems impossible to do at the moment, as Encoding
is a completely opaque object.
=end
Updated by Cezary (Cezary Baginski) over 15 years ago
Actions
#17
[ruby-core:35541]
=begin
On Fri, Mar 18, 2011 at 09:52:27PM +0900, Nikolai Weibull wrote:
On Fri, Mar 18, 2011 at 11:53, Magnus Holm judofyr@gmail.com wrote:
It's far easier for everyone that the built-in #upcase is
simple and fast and you'll have to be explicit about any
other I18n stuff IMO.Easy, perhaps, but hardly useful.
A agree - for human interaction it is completely useless. I tend to
think of #upcase as just a convenience method for dealing with ASCII
only system level functionality, e.g. paths on filesystems,
environment variables, html tags, (un)capitalizing to get class names,
database table names, etc.
Anything else is "no-op" or "undefined" for me.
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII.
I would probably go one step further and disallow upcase and friends
for any non-US-ASCII string for this reason. At least issue a warning.
If this isn’t of interest, then I’m still looking for a way to
override #upcase for Strings that use the UTF-8 encoding without
resorting to alias_method or extend (as shown earlier in this
discussion). This seems impossible to do at the moment, as Encoding
is a completely opaque object.
Correct me if I am wrong, but even "upper case" as a concept is not
common among all languages - an implementation detail for specific
cases at best.
For example, in German, you may want a more meaningful 'to_noun'
instead of 'capitalize'. For Japanese some may want upcase as a no-op
and some as a hack to convert to katakana. For case insensitivity,
probably a "normalize" method would be more descriptive.
Out of curiosity: in what specific case is utf upcase necessary?
--
Cezary Baginski
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#18
[ruby-core:35549]
=begin
On Tue, Mar 22, 2011 at 18:30, Cezary cezary.baginski@gmail.com wrote:
On Fri, Mar 18, 2011 at 09:52:27PM +0900, Nikolai Weibull wrote:
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII.
I would probably go one step further and disallow upcase and friends
for any non-US-ASCII string for this reason. At least issue a warning.
For Unicode there actually are well-defined casing rules.
For example, in German, you may want a more meaningful 'to_noun'
instead of 'capitalize'. For Japanese some may want upcase as a no-op
and some as a hack to convert to katakana. For case insensitivity,
probably a "normalize" method would be more descriptive.
This is perhaps true, but beside the point.
Out of curiosity: in what specific case is utf upcase necessary?
That’s a good question. It’s perhaps not a common operation, but text
editors and regular expression engines most likely need it. Even if
their utility is limited, returning incorrect results is worse.
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#19
[ruby-core:35522]
=begin
On Thu, Mar 25, 2010 at 19:33, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 18:24, NARUSE, Yui naruse@airemix.jp wrote:
(2010/03/26 0:02), Nikolai Weibull wrote:
I was wondering if there was a way to do it without having to do
String.new.unicodify.upcase
But I think, people want both ASCII version and Unicode version of upcase.
So you should name your Unicode methods another names.
Why would they want that? Having an ASCII-only version of #upcase
makes no sense for a Unicode String more than supporting #upcase
requires that you load the Unicode character database information,
which takes up quite a lot of memory.
So, what’s the reasoning here? Having "äbc".upcase return "äBC" makes
absolutely no sense and means that quite a few methods on String are
completely useless in a m18n context.
=end
Updated by judofyr (Magnus Holm) over 15 years ago
Actions
#20
[ruby-core:35524]
=begin
The problem is that the definition of #upcase doesn't only depend on the
encoding used, but also the language of the encoded text. For instance, if
you're writing in Turkish, you would expect "i".upcase to return a dotted
uppcase I: http://www.i18nguy.com/unicode/turkish-i18n.html
http://www.i18nguy.com/unicode/turkish-i18n.htmlDoing this properly is
really hard and needs to have a lot of flexibility, especially when it
comes to non-Western languages. It's far easier for everyone that the
built-in #upcase is simple and fast and you'll have to be explicit about any
other I18n stuff IMO.
// Magnus Holm
On Fri, Mar 18, 2011 at 11:19, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 19:33, Nikolai Weibull now@bitwi.se wrote:
On Thu, Mar 25, 2010 at 18:24, NARUSE, Yui naruse@airemix.jp wrote:
(2010/03/26 0:02), Nikolai Weibull wrote:
I was wondering if there was a way to do it without having to do
String.new.unicodify.upcase
But I think, people want both ASCII version and Unicode version of
upcase.
So you should name your Unicode methods another names.Why would they want that? Having an ASCII-only version of #upcase
makes no sense for a Unicode String more than supporting #upcase
requires that you load the Unicode character database information,
which takes up quite a lot of memory.So, what’s the reasoning here? Having "äbc".upcase return "äBC" makes
absolutely no sense and means that quite a few methods on String are
completely useless in a m18n context.
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#21
[ruby-core:35525]
=begin
On Fri, Mar 18, 2011 at 11:53, Magnus Holm judofyr@gmail.com wrote:
The problem is that the definition of #upcase doesn't only depend on the
encoding used, but also the language of the encoded text. For instance, if
you're writing in Turkish, you would expect "i".upcase to return a dotted
uppcase I: http://www.i18nguy.com/unicode/turkish-i18n.html
I know. The same goes for ‘i’ in Lithuanian.
Doing this properly is really hard and needs to have a lot of flexibility,
especially when it comes to non-Western languages.
This is simply not true. Unicode defines how to deal with case
conversions. I’m not saying that the Unicode standard is infallible,
but we can at least adhere to it. I’m not saying that Unicode is the
only encoding that we should care about, but if we support the Unicode
transfer formats, why not support other interesting parts of the
standard?
It's far easier for everyone that the built-in #upcase is
simple and fast and you'll have to be explicit about any
other I18n stuff IMO.
Easy, perhaps, but hardly useful.
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII. I was looking for an
actual solution to this problem. I have a library
(character-encodings) that does support these conversions, based on
locale and the Unicode character database (UCD). How do we make it
easy for the user to deal with m18n? I mean, if I say
-- coding: utf-8 --¶
puts "äbc".upcase
I expect this to do the right thing for Unicode under the current locale.
As Unicode defines how to deal with case conversions, if I tell Ruby
that “this String is encoded as UTF-8” (or, in this case, “strings in
this file are encoded as UTF-8”), I expect Ruby to respond “OK, I’ll
use the Unicode rules that govern methods like #upcase for that
String”.
The UCD requires a lot of memory, so I suggested that a library, such
as character-encodings, should be able to seamlessly add this kind of
behavior without requiring the user to write "äbc".unicodify.upcase,
if the UCD can’t be included in standard Ruby runtime.
But, come to think of it, doesn’t Oniguruma need most of the UCD
information, so isn’t most of it already included in the Ruby runtime?
Adding casing information perhaps wouldn’t require much additional
space.
If this isn’t of interest, then I’m still looking for a way to
override #upcase for Strings that use the UTF-8 encoding without
resorting to alias_method or extend (as shown earlier in this
discussion). This seems impossible to do at the moment, as Encoding
is a completely opaque object.
=end
Updated by Cezary (Cezary Baginski) over 15 years ago
Actions
#22
[ruby-core:35541]
=begin
On Fri, Mar 18, 2011 at 09:52:27PM +0900, Nikolai Weibull wrote:
On Fri, Mar 18, 2011 at 11:53, Magnus Holm judofyr@gmail.com wrote:
It's far easier for everyone that the built-in #upcase is
simple and fast and you'll have to be explicit about any
other I18n stuff IMO.Easy, perhaps, but hardly useful.
A agree - for human interaction it is completely useless. I tend to
think of #upcase as just a convenience method for dealing with ASCII
only system level functionality, e.g. paths on filesystems,
environment variables, html tags, (un)capitalizing to get class names,
database table names, etc.
Anything else is "no-op" or "undefined" for me.
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII.
I would probably go one step further and disallow upcase and friends
for any non-US-ASCII string for this reason. At least issue a warning.
If this isn’t of interest, then I’m still looking for a way to
override #upcase for Strings that use the UTF-8 encoding without
resorting to alias_method or extend (as shown earlier in this
discussion). This seems impossible to do at the moment, as Encoding
is a completely opaque object.
Correct me if I am wrong, but even "upper case" as a concept is not
common among all languages - an implementation detail for specific
cases at best.
For example, in German, you may want a more meaningful 'to_noun'
instead of 'capitalize'. For Japanese some may want upcase as a no-op
and some as a hack to convert to katakana. For case insensitivity,
probably a "normalize" method would be more descriptive.
Out of curiosity: in what specific case is utf upcase necessary?
--
Cezary Baginski
=end
Updated by now (Nikolai Weibull) over 15 years ago
Actions
#23
[ruby-core:35549]
=begin
On Tue, Mar 22, 2011 at 18:30, Cezary cezary.baginski@gmail.com wrote:
On Fri, Mar 18, 2011 at 09:52:27PM +0900, Nikolai Weibull wrote:
My point is that the current #upcase (and similar methods) is
basically useless for anything other than ASCII.
I would probably go one step further and disallow upcase and friends
for any non-US-ASCII string for this reason. At least issue a warning.
For Unicode there actually are well-defined casing rules.
For example, in German, you may want a more meaningful 'to_noun'
instead of 'capitalize'. For Japanese some may want upcase as a no-op
and some as a hack to convert to katakana. For case insensitivity,
probably a "normalize" method would be more descriptive.
This is perhaps true, but beside the point.
Out of curiosity: in what specific case is utf upcase necessary?
That’s a good question. It’s perhaps not a common operation, but text
editors and regular expression engines most likely need it. Even if
their utility is limited, returning incorrect results is worse.
=end