comment on

>>its simple to do, simply don't select it to begin with then in that case, we need to list ALL those tags we're interested in. won't this endup in a long list? HTML::Parser has a method ignore_tags() which could be used to ignore tags. I used it as below & tried to get the text, but it returned many nested arrays. I could not figure out how to access to final extracted text from this "@array"

my @array;
my $p = HTML::Parser->new(api_version => 3,
                          handlers => { text => [\@array, "text"]});
$p->ignore_tags(qw(table img));
$p->parse($page);
print "Size of array=$#array\n";
foreach my $aline (@array)
{
   print $aline;
}
print "\n";
[download]

Meanwhile, I found an alternative, but seems it is quite slower than what we could have achieved with HTML::Parser.

my $link = 'somelinek';
my $page = get($link) or die $!;
my $stream = HTML::TokeParser->new(\$page);
my $doparse = 1; ## 0 means don't parse
while (my $token = $stream->get_token)
{
    if ($token->[0] eq 'S')
    {
       if ($token->[1] eq 'table')
       {
         $doparse = 0;
       }
       elsif ($token->[1] eq 'img')
       {
         ;;
       }
    }
    elsif ($token->[0] eq 'E' and $token->[1] eq 'table')
    {
       $doparse = 1;
    }
    elsif ($token->[0] eq 'C')
    {
       ;;
    }
    elsif ($token->[0] eq 'T' and $doparse eq 1)
    { # text process the text in $token->[1]
        # skip: empty lines, "&nbsp;"
        if (defined ($token->[1]))
        {
          $token->[1] =~ s/&nbsp;/ /ig;
          $token->[1] =~ s/&#146;/'/ig;
          $token->[1] =~ s/&#14[7-8];/"/ig;
          $token->[1] =~ s/&#151;//ig;
          $token->[1] =~ s/&amp;/&/ig;
          $token->[1] =~ s/-{2,}//ig;
          print "$token->[1]";
        }
    }
}
[download]

This above use of TokeParser gives lot of broken text. Which could be better way? Thanks

In reply to Re^2: Ignoring specific html tags before parsing by ganeshPerlStarter
in thread Ignoring specific html tags before parsing by ganeshPerlStarter

Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!

Titles consisting of a single word are discouraged, and in most cases are disallowed outright.

Read Where should I post X? if you're not absolutely sure you're posting in the right place.

Please read these before you post! —

Posts may use any of the Perl Monks Approved HTML tags:

a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, details, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, summary, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr

You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)

	For:		Use:
	&		`&`
	<		`<`
	>		`>`
	[		`[`
	]		`]`

Link using PerlMonks shortcuts! What shortcuts can I use for linking?

See Writeup Formatting Tips and other pages linked from there for more info.