Re: length() miscounting UTF8 characters?

Having no experience with this, I thought I would explore a bit. So I did the following:

#!/usr/bin/perl
use strict;
use warnings;

use open IO => ':utf8';

while(<DATA>) {
  chomp;
  (my $nonenglish = $_) =~ s/[A-Za-z]//g;
  my @chars = split(//,$nonenglish);
  my $chars = scalar(@chars);
  print scalar(@chars), " $nonenglish\n";
}

__DATA__
æ
æð
æða
æðaber
æðahnútur
æðakölkun
æðardúnn
æðarfugl
æðarkolla
æðarkóngur
æðarvarp
æði
æðimargur
æðisgenginn
æðiskast
æðislegur
æðrast
æðri
æðrulaus
æðruleysi
æðruorð
æðrutónn
æðstur
æður
æfa
__END__
[download]

Seems split sees those letters as two chars also, which makes sense now that I think of it... . Guess I have some things to learn about UTF8!
Thanks for the opportunity! Sorry this is not all that helpful. Suppose one could take the character count and just divide by two ...

$chars = $chars / 2;
print "$chars $nonenglish\n";
[download]

...

Update: Might also take a look at CPAN Test UTF8 and related...

...the majority is always wrong, and always the last to know about it...
Insanity: Doing the same thing over and over again and expecting different results...

wjw

Comment on Re: length() miscounting UTF8 characters? Select or Download Code

Replies are listed 'Best First'.
Re^2: length() miscounting UTF8 characters? by AppleFritter (Vicar) on Apr 27, 2014 at 22:08 UTC
Yes, simply dividing by two would work here (and that's what I've been doing, mentally), but that's only because all the non-English characters encountered here are encoded as two bytes in UTF8. As soon as there'd be 3- or 4-byte characters, it'd not work anymore. Thanks for your help! I'll take a look at that module.	[reply]