comment on

Your suspicions are confirmed. The optimization has sped my strings-based tokenizer up by 39%. Interestingly, with the expanded test string I'm using, my original slightly edged out the optimized version of your single-regex. Here are the results:

           Rate    lists oneregex  str_org  str_opt
lists    2347/s       --     -38%     -45%     -60%
oneregex 3807/s      62%       --     -10%     -35%
str_org  4237/s      81%      11%       --     -28%
str_opt  5882/s     151%      55%      39%       --
[download]

This optimization is definitely effective. Thanks very much.

Oh. And here's the expanded test string if you want to play with it:

my $msg = q{This, is, an, example. Keep $2.50, 1,500, and 192.168.1.1.
+ 
I want to work this thing out a LITTEL!!!!L BITH!!!!! MORE@@@@@@ with 
some,.unhapp.yword,combinations.and , a little .. bit of,, 
confusing,text hopefully @#@#@#@%#$57)#$*(#&)(*$ it will @#@][] work.}
+;
[download]

In reply to Re: Re: Re: Re: tokenize plain text messages by revdiablo
in thread tokenize plain text messages by revdiablo

Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!

Titles consisting of a single word are discouraged, and in most cases are disallowed outright.

Read Where should I post X? if you're not absolutely sure you're posting in the right place.

Please read these before you post! —

Posts may use any of the Perl Monks Approved HTML tags:

a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, details, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, summary, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr

You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)

	For:		Use:
	&		`&`
	<		`<`
	>		`>`
	[		`[`
	]		`]`

Link using PerlMonks shortcuts! What shortcuts can I use for linking?

See Writeup Formatting Tips and other pages linked from there for more info.