meta tag extraction with TokeParser

Anonymous Monk has asked for the wisdom of the Perl Monks concerning the following question:

There was one node on here when I searched how to do this (code shown below) but it works with errors. It produces the right output but it has some 30 lines of errors saying an uninitialized value.

my %meta;
my $htm2 = HTML::TokeParser->new( \$src );
while (my $token = $htm2->get_token) {
        next if $token->[1] ne 'meta' && $token->[0] ne 'S';
        $meta{$token->[2]{name}} = $token->[2]{content};
}
[download]

Ideally, I want to collect all meta tags and store each of them into a hash with the meta name as the key. I'm already using TokeParser for scraping, so please don't suggest I also use TokeParser::Simple. I read through the docs and can't seem to find any information on what I am looking for.

Also, if a modified version of the code above works, can you explain line for line what it's doing? I'm having trouble piecing things together.

My last question is this. Can I extract different parts of an HTML document with TokeParser in one run? Or must I run them all separately?

I can extract the title tag just fine, but only when I make a new reference to TokeParser. It seems like a waste of resources to call the module AGAIN when the html dump is still in memory, right? Or does the data change after each time you loop over tokens?

Comment on meta tag extraction with TokeParser Download Code

Replies are listed 'Best First'.
What's wrong with HTML::TokeParser::Simple? by Ovid (Cardinal) on Mar 20, 2006 at 23:34 UTC
I'm already using TokeParser for scraping, so please don't suggest I also use TokeParser::Simple. I don't understand this. Perhaps you weren't aware of this, but a deliberate design decision of HTML::TokeParser::Simple was to be a drop in replacement of HTML::TokeParser. Using my module means changing this: `use HTML::TokeParser; my $parser = HTML::TokeParser->new( \$src );` [download] To this: `use HTML::TokeParser::Simple; my $parser = HTML::TokeParser::Simple->new( \$src );` [download] After that change, all of your code still works. I did that deliberately so that folks could take advantage of the features of my module without having to rewrite their code. Your snippet changes from: `my %meta; my $htm2 = HTML::TokeParser->new( \$src ); while (my $token = $htm2->get_token) { next if $token->[1] ne 'meta' && $token->[0] ne 'S'; $meta{$token->[2]{name}} = $token->[2]{content}; }` [download] To: `my %meta; my $htm2 = HTML::TokeParser::Simple->new( \$src ); while (my $token = $htm2->get_token) { next unless $token->is_start_tag('meta'); $meta{$token->get_attr('name')} = $token->get_attr('content'); }` [download] I think most would agree that not only is this far more readable (and therefor far more maintainable). So if you're still not convinced, that's OK, but please don't try to suggest to others that HTML::TokeParser::Simple is not a good alternative to HTML::TokeParser. It's far easier to understand and use. You gain a lot and lose nothing. Cheers, Ovid New address of my CGI Course.	[reply] [d/l] [select]
Re: meta tag extraction with TokeParser by Thelonius (Priest) on Mar 21, 2006 at 00:18 UTC
I think instead of this: `next if $token->[1] ne 'meta' && $token->[0] ne 'S';` [download] you meant: `next if $token[0] ne 'S' \|\| $token[1] ne 'meta';` [download] Note the order of evaluation and the \|\| instead of && But it might be even clearer if you say: `if ($token->[0] eq 'S' && $token->[1] eq 'meta' && $token->[2]{name}) +{ $meta{$token->[2]{name}} = $token->[2]{content}; }` [download]	[reply] [d/l] [select]
Re: meta tag extraction with TokeParser by saintmike (Vicar) on Mar 20, 2006 at 20:59 UTC
`next if $token->[1] ne 'meta' && $token->[0] ne 'S'; $meta{$token->[2]{name}} = $token->[2]{content};` [download] If you're seeing warnings about an 'uninitialized value', you should check if `$token->[x]` exists and is defined before comparing it to a different value or following the alleged reference.	[reply] [d/l] [select]
Re^2: meta tag extraction with TokeParser by Anonymous Monk on Mar 20, 2006 at 21:44 UTC
I tried adding the following but it didn't change `if ($token->[2] ne '') { $meta{$token->[2]{name}} = $token->[2]{content}; }` [download]	[reply] [d/l]
Re^3: meta tag extraction with TokeParser by saintmike (Vicar) on Mar 20, 2006 at 21:47 UTC
Again, you're not checking if it's defined. Use this instead: `if( defined $token->[2] ) { ... }` [download]	[reply] [d/l]
Re^4: meta tag extraction with TokeParser by Anonymous Monk on Mar 20, 2006 at 21:51 UTC
Re^5: meta tag extraction with TokeParser by Anonymous Monk on Mar 20, 2006 at 21:54 UTC