comment on

OfficeLinebacker asked ( in Question why this Regex isn't matching ) for a regex to extract the link text from some (HTML) data and was correctly advised to use a module instead. In fact, /me was one of those with that advice ...but the OP stuck in my mind and today I got to playing around with some possibilities.

Naturally, I screwed up a few times... that's no great suprise -- "Typos is use," just for starters.

But, IMO, the results of some of those typos/tests raised interesting questions. Here's the base code, with comments indicating how the regex (and the output label -- 1, 2, or 3) varied in separate but otherwise identical scripts:

#!/usr/bin/perl
use strict;
use warnings;
use 5.012;

# 928860a

$\=undef;
my $string = <DATA>;

#1 capture paren pair contains the non-grouping parens
if ( $string =~ />((?:\s|\w)+)(?!<\td>)/m ) { 
    say "regex 1: $1";
}

# cf #1  />((?:\s|\w)+)(?!<\td>)/m 
# "</td>" erroneously written as "<\td>" (escaping the "t") captures p
+recisely as intended

# cf #2  />((?:\s|\w)+)(?!<\/td>)/m
# regex as intended fails to capture the terminal "g"

# cf #3  />(?:(\s|\w)+)(?!<\/td>)/m
# captures only the (penultimate?) "n"

__DATA__
/td><td class="stdViewCon">Moving Services for MSCD Student Success Bu
+ilding</td></tr></table></TD></TR>

<TR VALIGN=top><TD><table class="stdViewITbl"><tr><td class="stdViewLn
+kLbl">
[download]

Here's a transcript of execution and output from the three almost-idential tests:

C:\>928860a.pl
regex 1: Moving Services for MSCD Student Success Building

C:\>928860b.pl
regex 2: Moving Services for MSCD Student Success Buildin

C:\>928860c.pl
regex 3: n
[download]

Admittedly, I've re-read the regex perldocs only in a cursory manner and have reviewed only the index in Mastering Regular Expressions, Second Edition (the one in my lib; Third Edition is current), but does anyone have a pointer to documentation to explain what is for me, inexplicable: Why does regex 1, with the error, produce the desired output, while regex 2 fails to capture the terminal "g" in the desired link-text and regex3 fails almost entirely?

In reply to Why do these regex variants behave as they do? by ww

Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!

Titles consisting of a single word are discouraged, and in most cases are disallowed outright.

Read Where should I post X? if you're not absolutely sure you're posting in the right place.

Please read these before you post! —

Posts may use any of the Perl Monks Approved HTML tags:

a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, details, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, summary, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr

You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)

	For:		Use:
	&		`&`
	<		`<`
	>		`>`
	[		`[`
	]		`]`

Link using PerlMonks shortcuts! What shortcuts can I use for linking?

See Writeup Formatting Tips and other pages linked from there for more info.