comment on

G'day Bill,

"I did not know what unicode character the \x96 was meant to represent."

A quick way to determine this is via "Unicode Character Code Charts" — it has "Find chart by hex code:" near the top of the page.

[Aside: Although that's a standard URL, I noted, when checking it, that it has: "Unicode 15.0 Character Code Charts". I thought that I'd just mention that Perl does a pretty good job of supporting the latest Unicode versions. Perl v5.36.0 (released in May this year) supports Unicode 14.0 (the current version at the time); if you're desperate for 15.0 support, it was added in v5.37.5 (or just wait for 5.38.0 to be released in May next year, or thereabouts).]

That will give you the name, <control>, and the informative alias, START OF GUARDED AREA; you can use the latter in \N{}.

$ perl -E 'say sprintf "%x", ord("\N{START OF GUARDED AREA}")'
96
[download]

In a script or one-liner, you can use Unicode::UCD, but it's not always straightforward. Compare:

$ perl -MUnicode::UCD=charinfo -E 'say charinfo(0x34)->{name}'
DIGIT FOUR

$ perl -MUnicode::UCD=charinfo -E 'say charinfo(0x34)->{unicode10} || 
+"<blank>"'
<blank>

$ perl -MUnicode::UCD=charinfo -E 'say charinfo(0x96)->{name}'
<control>

$ perl -MUnicode::UCD=charinfo -E 'say charinfo(0x96)->{unicode10} || 
+"<blank>"'
START OF GUARDED AREA
[download]

— Ken

In reply to Re^5: Malformed UTF-8 character by kcott
in thread Malformed UTF-8 character by BillKSmith

Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!

Titles consisting of a single word are discouraged, and in most cases are disallowed outright.

Read Where should I post X? if you're not absolutely sure you're posting in the right place.

Please read these before you post! —

Posts may use any of the Perl Monks Approved HTML tags:

a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, details, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, summary, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr

You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)

	For:		Use:
	&		`&`
	<		`<`
	>		`>`
	[		`[`
	]		`]`

Link using PerlMonks shortcuts! What shortcuts can I use for linking?

See Writeup Formatting Tips and other pages linked from there for more info.