Help with performance (file io)

There are many interesting points in that post. However, I disagree that the task is mostly I/O bound. In fact, it seems to me the task is mostly CPU bound, spending most of its time in filter_line (in the “read” version). I fiddled with it a bit, and was able to reduce the read version to about 1sec (from 1.8) with the following approach:

defp filter_line(line) do
  filter_line(line,<<>>)
end

defp filter_line(<<c::utf8, rest::binary>>, integer) when c in ?0..?9 do
  filter_line(rest, <<integer::binary, c>>)
end
defp filter_line(<<?,::utf8, _::binary>>, integer) do
  num = String.to_integer(integer)
  rem(num, 2) == 0 || rem(num, 5) == 0
end
defp filter_line(<<_other::utf8, rest::binary>>, _integer) do
  filter_line(rest, <<>>)
end

Basically, I’m iterating one codepoint at a time, accumulating integer codepoints (0..9) until the first observed comma. Then I convert the accumulated string of digits to int, and do the rem math. By doing this I don’t need to split the entire line by comma (since I don’t care about the rest anyway), and also don’t need to do additional regex scan for the terminating integer.

I didn’t do much to analyze the correctness besides verifying that the output file size is the same. I’m also not saying this is the best (or optimal) approach. The point is, as I hinted earlier in this thread, that most of the action is taking place in a few CPU bound functions (assuming you perform I/O efficiently, which you do with File.read! and File.write!). If Ruby’s versions of those operations are implemented in C, then Ruby will probably be faster.