in reply to Searching for a word that may only exist in part

This seems to work. Run it on a terminal 100+ wide to see the output properly; the wrapping screws it up.

#! perl -slw use strict; sub search { my( $haystack, $needle, $pos ) = @_; my $l = length $needle; my( @res1, @res2 ); return( $pos-1, 0, $l ) if $pos = 1+index $haystack, $needle; substr( $haystack, 0, $_ ) eq substr( $needle, -$_ ) and @res1 = ( 0, $l-$_, $_ ), last for reverse 1 .. $l-1; substr( $haystack, $_-$l ) eq substr( $needle, 0, $l-$_ ) and @res2 = ( length( $haystack )-( $l-$_), -$_, $l-$_ ), las +t for 1 .. $l-1; return unless @res1 or @res2; return @res1 unless @res2; return @res2 unless @res1; return $res1[2] > $res2[2] ? @res1 : @res2; } my $haystack = 'GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAAT +GGTGACTGCCATTGATGGAGGGAGACACA'; my @needles = ( 'NOT FOUND', 'CTGGATAAGAATGTTTTAGCAATCTCTT', 'CTGCCATTGATGGAGGGAGACACAACGTACGT', 'CATGAATCCATGGCAGTGACCATACTAATGGTGACTG', 'AGGGAGACACAGAATGTTTTAG', 'AGGGAGACACAGAATGTTTTAGC', ); $, = ' '; for my $needle ( @needles ) { my( $hOffset, $nOffset, $l ) = search $haystack, $needle; print "No match found in '$haystack'\n for '$needle'\n" and next unless defined $hOffset; printf "%s%s\n%s%s\n\n", ' 'x$nOffset, $haystack, ' 'x$hOffset, $needle; } __END__ c:\test>579215 No match found in 'GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACT +AATGGTGACTGCCATTGATGGAGGGAGACACA' for 'NOT FOUND' GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAATGGTGACTG +CCATTGATGGAGGGAGACACA CTGGATAAGAATGTTTTAGCAATCTCTT GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAATGGTGACTGCCATTGAT +GGAGGGAGACACA CTGCCATTGAT +GGAGGGAGACACAACGTACGT GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAATGGTGACTGCCATTGAT +GGAGGGAGACACA CATGAATCCATGGCAGTGACCATACTAATGGTGACTG GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAATGGTGACTGCCATTGAT +GGAGGGAGACACA + AGGGAGACACAGAATGTTTTAG GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAATGGTGA +CTGCCATTGATGGAGGGAGACACA AGGGAGACACAGAATGTTTTAGC

Examine what is said, not who speaks -- Silence betokens consent -- Love the truth but pardon error.
Lingua non convalesco, consenesco et abolesco. -- Rule 1 has a caveat! -- Who broke the cabal?
"Science is about questioning the status quo. Questioning authority".
In the absence of evidence, opinion is indistinguishable from prejudice.

Replies are listed 'Best First'.
Re^2: Searching for a word that may only exist in part
by seaver (Pilgrim) on Nov 07, 2006 at 20:38 UTC
    Dear BrowserUk,

    I ran your search subroutine on some hard data, three 'needles', and three different 'haystacks'. I realised that you didn't set a lower limit of consecutive matches, as you'll see from the output below, it even matches when just the first letter of the needle matches the last letter of the haystack.

    I was going to try and edit your code myself, but the collection of return statements had me confused, so I was wondering, what would you do to your code to set that lower limit, which should probably be a higher number, like 6, now that I think about it.

    Cheers
    Sam

    TTGTCAGCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGACCGGGCAGCCTAAAGGC +TATCCTTAACCAGGGAGCTGATT GCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGAC TTGTCAGCGAAAAAAATTAAAGCG +CAAGATTGTTGGTTTTTGCGTGATGGTGACCGGGCAGCCTAAAGGCTATCCTTAACCAGGGAGCTGATT CGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTT GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTA +ATGGTGACTGCCATTGATGGAGGGAGACACAGTGCACTGGCAAACTCACAC CATTACATTGCTGGATAAGAATGTTTTAGCAATCTCTTTCTGTCATGA GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAATGGTGACTGCCATTGAT +GGAGGGAGACACAGTGCACTGGCAAACTCACAC + CGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCT +GAATAACGTTT TAATCAAAACCAATAAACACGAAATAATCCCCATGCCGGTGAAGAAGGGGCGTGACTTTAGCGAAATGTT +GCCGTCGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTTATGCTGAAAGCGGATG +AATAAGGAGATGCG + + GCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGAC TAATCAAAACCAATAAACACGAAATAATCCCCATGCCGGTGAAGAAGGGGCGTGACTTTAGCGAAATGTT +GCCGTCGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTTATGCTGAAAGCGGATG +AATAAGGAGATGCG + CGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTT
    Update: Removed redundant lines

      All that is needed is a test in each loop to quit when the partial match attempt is less than whatever minimum you want to specify. Rather than duplicate the test in both loops, I've amalgamated them into a single loop which allows for a single test, with the nice side-effect of being 25% more efficient than the original above.

      I've also added some explanitary comments and (I think) improved the variable names. Let me know if anything is unclear.

      I've posted the output of a run with the minimum match set to 3, to test the breakpoint, but set it to whatever value makes sense for your purposes.

      #! perl -slw use strict; use constant MINMATCH => 3; ## returns a list of 3 integers. ## offset into the haystack ## offset into the needle ( can be negative, subtract from length(need +le) ) ## length of match. sub search { my( $haystack, $needle, $pos ) = @_; my( @start, @end ); my $l = length $needle; ## A full match found return( $pos-1, 0, $l ) if $pos = 1+index $haystack, $needle; ## iterate, and attempt matches at both ends for( 1 .. $l-1 ) { my $r = $l - $_; ## reverse offset last if $r < MINMATCH; ## quit if we've reached the minimum ma +tch length ## try a partial match at the beginning substr( $haystack, $_-$l ) eq substr( $needle, 0, $l-$_ ) and @start = ( length( $haystack )-( $l-$_), -$_, $l-$_ ) +; ## try a partial match at the end substr( $haystack, 0, $r ) eq substr( $needle, -$r ) and @end = ( 0, $_, $r ); ## Quit if we got either. last if @start or @end; } return unless @start or @end; ## No match return @start unless @end; ## No partial at the end, st +art is best return @end unless @start; ## No partial at the start, +end is best return $start[2] >= $end[2] ? @start : @end; ## Got both, longest +(or start if equal) is best } our @haystacks = ( 'TTGTCAGCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGACCGGGCAGCCTA +AAGGCTATCCTTAACCAGGGAGCTGATT', 'GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTAATGGTGACTGCCA +TTGATGGAGGGAGACACAGTGCACTGGCAAACTCACAC', 'TAATCAAAACCAATAAACACGAAATAATCCCCATGCCGGTGAAGAAGGGGCGTGACTTTAGCGAA +ATGTTGCCGTCGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTTATGCTGAAAGC +GGATGAATAAGGAGATGCG', ); our @needles = ( 'GCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGAC', 'CATTACATTGCTGGATAAGAATGTTTTAGCAATCTCTTTCTGTCATGA', 'CGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTT', ); $, = ' '; for my $needle ( @needles ) { for my $haystack ( @haystacks ) { my( $hOffset, $nOffset, $l ) = search $haystack, $needle; if( defined $hOffset ) { print "$hOffset, $nOffset, $l"; printf "%s%s\n%s%s\n\n", ' 'x$nOffset, $haystack, ' 'x$hOffset, $needle; } else { print "No match found in '$haystack'\n for '$needle'\n"; } } } __END__ C:\test>579215-2 6, 0, 48 TTGTCAGCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGACCGGGCAGCCTAAAGGC +TATCCTTAACCAGGGAGCTGATT GCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGAC No match found in 'GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACT +AATGGTGACTGCCATTGATGGAGGGAGACACAGTGCACTGGCAAACTCACAC' for 'GCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGAC' 150, -45, 3 TAATCAAAACCAATAAACACGAAATAATCCCCATGCCGGTGAAGAAGGGGCGTGACTTTAGCGAAATGTT +GCCGTCGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTTATGCTGAAAGCGGATG +AATAAGGAGATGCG + + GCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGTGAC No match found in 'TTGTCAGCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGT +GACCGGGCAGCCTAAAGGCTATCCTTAACCAGGGAGCTGATT' for 'CATTACATTGCTGGATAAGAATGTTTTAGCAATCTCTTTCTGTCATGA' 0, 18, 30 GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACTA +ATGGTGACTGCCATTGATGGAGGGAGACACAGTGCACTGGCAAACTCACAC CATTACATTGCTGGATAAGAATGTTTTAGCAATCTCTTTCTGTCATGA No match found in 'TAATCAAAACCAATAAACACGAAATAATCCCCATGCCGGTGAAGAAGGGGC +GTGACTTTAGCGAAATGTTGCCGTCGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACG +TTTATGCTGAAAGCGGATGAATAAGGAGATGCG' for 'CATTACATTGCTGGATAAGAATGTTTTAGCAATCTCTTTCTGTCATGA' No match found in 'TTGTCAGCGAAAAAAATTAAAGCGCAAGATTGTTGGTTTTTGCGTGATGGT +GACCGGGCAGCCTAAAGGCTATCCTTAACCAGGGAGCTGATT' for 'CGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTT' No match found in 'GAATGTTTTAGCAATCTCTTTCTGTCATGAATCCATGGCAGTGACCATACT +AATGGTGACTGCCATTGATGGAGGGAGACACAGTGCACTGGCAAACTCACAC' for 'CGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTT' 75, 0, 48 TAATCAAAACCAATAAACACGAAATAATCCCCATGCCGGTGAAGAAGGGGCGTGACTTTAGCGAAATGTT +GCCGTCGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTTATGCTGAAAGCGGATG +AATAAGGAGATGCG + CGCGACAACCGGAATATGAAAGCAAAGCGCAGCGTCTGAATAACGTTT

      Examine what is said, not who speaks -- Silence betokens consent -- Love the truth but pardon error.
      Lingua non convalesco, consenesco et abolesco. -- Rule 1 has a caveat! -- Who broke the cabal?
      "Science is about questioning the status quo. Questioning authority".
      In the absence of evidence, opinion is indistinguishable from prejudice.
Re^2: Searching for a word that may only exist in part
by seaver (Pilgrim) on Oct 26, 2006 at 18:35 UTC
    Dear BrowserUK, Mreece and Grandfather,

    Many thanks for your wise offerings.

    Sam