Project

General

Profile

Feature #22231

Updated by himura467 (Akito Shitara) 2 months ago

## Use case 

 I am writing nack-ruby (https://github.com/nsgi-org/nack-ruby), a Ruby implementation of NSGI, a host/guest interface `IO::Buffer` has no search primitive: no `#index`, no delimiter scan, no substring search. This matters for web applications. The host hands each HTTP servers that hand request field data to Ruby the application as an zero-copy `IO::Buffer` view views over its their own socket read buffer, so a request reaches buffer. The zero-copy path ends at the application without its bytes being copied. 

 That holds right up until the application parses inside a field, which is where request handling actually begins: first delimiter scan, such as splitting a query string on `&` and `=`, scanning `&`, finding a body for a multipart boundary, cutting or scanning a header value at a delimiter. `IO::Buffer` has no way value, which is where request processing usually begins. 

 The same applies to any protocol parser built on `IO::Buffer`: line framing, chunked transfer encoding, netstrings, length-prefixed records. All of them must find a byte, so at that point delimiter before they can decide what to slice. 

 Today there are two ways to do it, and both give up what the only options are: buffer was for: 

 ```ruby 
 buffer.get_string.index("&")          # copies the field whole region to find one offset 

 i = 0                                 # no copy, but a method call per byte 
 i += 1 while i < buffer.size && buffer.get_value(:U8, i) != 0x26 
 ``` 

 The first throws away the zero-copy path the host went to the trouble of providing, and allocates Searching a `String` proportional to the field; copying a 64 KiB body by copying it out to search it measures is 3.7x slower than searching it in place, and produces allocates 66 KB of garbage per call. The second keeps the buffer byte loop allocates nothing but is roughly 1570x slower. Either way the buffer stops paying for itself at the first delimiter. 

 This is not specific to NSGI. Any protocol parser built on `IO::Buffer`, such as line framing, chunked transfer encoding, or length-prefixed records, has to find a delimiter before it can decide what to slice. 

 ## Specification 

 ```ruby 
 buffer.index(object, offset = 0, length = size - offset) # => Integer or nil 
 ``` 

 `object` may be an `Integer` byte value (`0..255`), a `String`, or another `IO::Buffer`. 

 ```ruby 
 buffer = IO::Buffer.for("Hello World") 

 buffer.index("World")                    # => 6 
 buffer.index("o".ord)                    # => 4 
 buffer.index(IO::Buffer.for("World"))    # => 6 
 buffer.index("!")                        # => nil 
 buffer.index("o", 5)                     # => 7     (absolute, not relative to offset) 

 buffer.slice(6, 5).index("o")            # => 1     (relative to the slice) 
 ``` 

 - Searching is byte-oriented; a `String`'s encoding is ignored. 
 - An empty `object` matches at `offset`, as with `String#index`. 
 - An `object` longer than the range returns `nil`. 
 - An out-of-range `offset` or `length` raises `ArgumentError` (see below). 

 Single-byte values use `memchr`. Longer values reuse `rb_memsearch`, the portable substring search that already backs `String#index`. 

 ## Open questions 

 ### Out-of-range `offset`/`length` raises, where `String#index` returns `nil`. 

 All ten existing `(offset, length)` methods in `io_buffer.c` route through `io_buffer_validate_range` and raise, and the class never clamps. `String#index` also has no `length` parameter, so matching it does not settle what `length` should do; the Ruby-wide convention there is to clamp, as in `"hello".byteslice(0, 1000)`, which is what `IO::Buffer` declines to do elsewhere. I chose consistency within the class, and am happy to switch if the `String` reading is preferred. 

 ### An `Integer` outside `0..255` raises, where `#clear` masks it. 

 `buffer.clear(256)` fills with `0`, and `String#setbyte` masks too. Masking a value being searched for seems worse than masking one being written, since `index(256)` would quietly search for `\x00` and could return a match. 

 ## Follow-ups 

 `#rindex` for a reverse scan, and an `#each_until(delimiter)` framing iterator yielding successive delimited slices, in the same spirit as the vectorizable `#and!` / `#or!` / `#bit_count` family. Both build on `#index`, so this ticket is scoped to it.

Back