Provided my code has no bugs, see for yourself:

I've noticed persistent differences in the results of your code and mine. (Well, the versions of your code as included in ikegami's benchmarks.) They take the form of 2 to 3 subsequences that differ by a count of 1 or 2. I suspect that this is an easy bug to fix... but since it's your code, I figure you can debug and fix it. ;-)

That all said, I'd like to point out that 1) including the file reads in your benchmark obscures the issue. As long as the machine has plenty of memory (and maybe yours doesn't) you are contaminating your results with two different methods of reading the data. And 2) you could modify my algorithm and most of the others' to work with chunks as well if RAM really was an issue.

If you work your bugs out, I am interested in your method of iterating over the shorter strings for the smaller length matches though. I haven't looked at it closely yet, but I'd like to see how that scales.

-sauoq
"My two cents aren't worth a dime.";

In reply to Re^2: Question about speeding a regexp count by sauoq
in thread Question about speeding a regexp count by Commander Salamander

Title:
Use:  <p> text here (a paragraph) </p>
and:  <code> code here </code>
to format your post, it's "PerlMonks-approved HTML":



  • Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!
  • Titles consisting of a single word are discouraged, and in most cases are disallowed outright.
  • Read Where should I post X? if you're not absolutely sure you're posting in the right place.
  • Please read these before you post! —
  • Posts may use any of the Perl Monks Approved HTML tags:
    a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, details, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, summary, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr
  • You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)
            For:     Use:
    & &amp;
    < &lt;
    > &gt;
    [ &#91;
    ] &#93;
  • Link using PerlMonks shortcuts! What shortcuts can I use for linking?
  • See Writeup Formatting Tips and other pages linked from there for more info.