comment on

Mainly it's a HTML::Parser exercise, done during a heavy research moment :)

The idea is very simple: randomly putting together texts and images from your browser's cache, you could get a snapshot of your (or someone else's) behaviours. Like a statistician in your garbage :)

I grabbed the idea from an old DOMUS issue, but I've lost the original URL.

Update: I'm thinking to use it as a permanent installation in an Internet Cafe'. Two monitors: one used by surfer, and one connected to a box that automatically refresh a page generated by this script.

#!/usr/bin/perl

use strict;


# Digs in your browser's cache 
# like a statistician in your trashcan...

package Lurker;

use File::Find;

my $cache = {
    IMAGES => [],
    DOCS => [],
};

sub lurk
{
    my $dir = shift;
    
    print STDERR "Reading cache...";

    find(
     sub 
     {   
         for ( $File::Find::name ) {
         /\.gif$/  ||
         /\.png$/  || 
         /\.jpg$/  &&    push @{ $cache->{ IMAGES }}, $_;
         /\.html$/ &&    push @{ $cache->{ DOCS }}, $_;
         }
     }, $dir 
    );

    print STDERR "OK!\n";
}

sub pick_random
{
    my $what = shift; 
    
    my $n = scalar( @{$cache->{ $what }} );
    
    return ${$cache->{ $what }}[ rand $n ];
}




package My_HTML_Parser;

use base 'HTML::Parser';

sub start
{
    my $self = shift;
    my ($tag, $attr, $attrseq, $origtext) = @_;
    
    my ($orig_src, $new_src);
    
    if ($tag eq 'img') {
    $orig_src = $attr->{'src'};        
    $new_src = Lurker::pick_random( 'IMAGES' );
    $origtext =~ s/$orig_src/$new_src/;
    }
    print $origtext;
}

sub text
{
    my $self = shift;
    my ($text) = @_;
    
    print $text;
}

sub end
{
    my $self = shift;
    my ($tag) = @_;
    
    print "</$tag>";
}



package main;

my $cache_directory = '/home/stefano/.netscape/cache';

Lurker::lurk( $cache_directory );

my $doc = Lurker::pick_random('DOCS');

print STDERR "Now parsing $doc...\n";

my $a = new My_HTML_Parser;
$a->parse_file( $doc );
[download]

In reply to Statistician in my garbage... by larsen

Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!

Titles consisting of a single word are discouraged, and in most cases are disallowed outright.

Read Where should I post X? if you're not absolutely sure you're posting in the right place.

Please read these before you post! —

Posts may use any of the Perl Monks Approved HTML tags:

a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, details, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, summary, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr

You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)

	For:		Use:
	&		`&`
	<		`<`
	>		`>`
	[		`[`
	]		`]`

Link using PerlMonks shortcuts! What shortcuts can I use for linking?

See Writeup Formatting Tips and other pages linked from there for more info.