Re: PERL Regular expression

This is a good reason not to try writing your own regex for complex, well known problems - using a std module is almost always better. your code formatting was botched, but I took a couple of guesses.

I used YAPE::Regex::Explain to analyze what you did. See below. as for your question, \3 means not "\ or 3" in the set context that you have, not the value of the thing that $3 is.

#!/usr/bin/perl -w
use strict;

use YAPE::Regex::Explain;

my $line = '<log dest="calllog"> directory <value expr="callRecDirPath
+"/>';

if ( $line =~ /<(\w(?:[^ \/>]+))(?:(?:(?:\s+)([^=>]+)\s*=\s*("|')([^\3
+]*)\3))\s*\/>/) 
{

print "1=" . $1. "\n2=" . $2 . "\n3=" . $3 . "\n4=" . $4 . "\n5=" . $5
+ . "\n";

}

my $REx = q{$line =~ /<(\w(?:[^ \/>]+))(?:(?:(?:\s+)([^=>]+)\s*=\s*("|
+')([^\3]*)\3))\s*\/>/};

my $exp = YAPE::Regex::Explain->new($REx)->explain;

print $exp;
[download]

Output of Analysis:

Use of uninitialized value $5 in concatenation (.) or string at line 1
+1.
1=log
2=dest
3="
4=calllog"> directory <value expr="callRecDirPath
5=
The regular expression:

(?-imsx:$line =~ /<(\w(?:[^ \/>]+))(?:(?:(?:\s+)([^=>]+)\s*=\s*("|')([
+^\3]*)\3))\s*\/>/)

matches as follows:
  
NODE                     EXPLANATION
----------------------------------------------------------------------
(?-imsx:                 group, but do not capture (case-sensitive)
                         (with ^ and $ matching normally) (with . not
                         matching \n) (matching whitespace and #
                         normally):
----------------------------------------------------------------------
  $                        before an optional \n, and the end of the
                           string
----------------------------------------------------------------------
  line =~ /<               'line =~ /<'
----------------------------------------------------------------------
  (                        group and capture to \1:
----------------------------------------------------------------------
    \w                       word characters (a-z, A-Z, 0-9, _)
----------------------------------------------------------------------
    (?:                      group, but do not capture:
----------------------------------------------------------------------
      [^ \/>]+                 any character except: ' ', '\/', '>'
                               (1 or more times (matching the most
                               amount possible))
----------------------------------------------------------------------
    )                        end of grouping
----------------------------------------------------------------------
  )                        end of \1
----------------------------------------------------------------------
  (?:                      group, but do not capture:
----------------------------------------------------------------------
    (?:                      group, but do not capture:
----------------------------------------------------------------------
      (?:                      group, but do not capture:
----------------------------------------------------------------------
        \s+                      whitespace (\n, \r, \t, \f, and " ")
                                 (1 or more times (matching the most
                                 amount possible))
----------------------------------------------------------------------
      )                        end of grouping
----------------------------------------------------------------------
      (                        group and capture to \2:
----------------------------------------------------------------------
        [^=>]+                   any character except: '=', '>' (1 or
                                 more times (matching the most amount
                                 possible))
----------------------------------------------------------------------
      )                        end of \2
----------------------------------------------------------------------
      \s*                      whitespace (\n, \r, \t, \f, and " ")
                               (0 or more times (matching the most
                               amount possible))
----------------------------------------------------------------------
      =                        '='
----------------------------------------------------------------------
      \s*                      whitespace (\n, \r, \t, \f, and " ")
                               (0 or more times (matching the most
                               amount possible))
----------------------------------------------------------------------
      (                        group and capture to \3:
----------------------------------------------------------------------
        "                        '"'
----------------------------------------------------------------------
       |                        OR
----------------------------------------------------------------------
        '                        '\''
----------------------------------------------------------------------
      )                        end of \3
----------------------------------------------------------------------
      (                        group and capture to \4:
----------------------------------------------------------------------
        [^\3]*                   any character except: '\3' (0 or
                                 more times (matching the most amount
                                 possible))
----------------------------------------------------------------------
      )                        end of \4
----------------------------------------------------------------------
      \3                       what was matched by capture \3
----------------------------------------------------------------------
    )                        end of grouping
----------------------------------------------------------------------
  )                        end of grouping
----------------------------------------------------------------------
  \s*                      whitespace (\n, \r, \t, \f, and " ") (0 or
                           more times (matching the most amount
                           possible))
----------------------------------------------------------------------
  \/                       '/'
----------------------------------------------------------------------
  >/                       '>/'
----------------------------------------------------------------------
)                        end of grouping
----------------------------------------------------------------------
[download]

Comment on Re: PERL Regular expression Select or Download Code

Replies are listed 'Best First'.
Re^2: PERL Regular expression by Anonymous Monk on Sep 03, 2010 at 19:53 UTC
Thank you very much for your explanation & timely help.	[reply]