We are doing Natural Language Processing on a range of English language documents (mainly

Question

0

Asked: May 13, 20262026-05-13T06:36:58+00:00 2026-05-13T06:36:58+00:00

We are doing Natural Language Processing on a range of English language documents (mainly

0

We are doing Natural Language Processing on a range of English language documents (mainly scientific) and run into problems in carrying non-ANSI characters through the various components. The documents may be “ASCII”, UNICODE, PDF, or HTML. We cannot predict at this stage what tools will be in our chain or whether they will allow character encodings other than ANSI. Even ISO-Latin characters expressed in UNICODE will give problems (e.g. displaying incorrectly in browsers). We are likely to encounter a range of symbols including mathematical and Greek. We would like to “flatten” these into a text string which will survive multistep processing (including XML and regex tools) and then possibly reconstitute it in the last step (although it is the semantics rather than the typography we are concerned with so this is a minor concern).

I appreciate that there is no absolute answer – any escaping can clash in some cases – but I am looking for something allong the lines of XML’s <![CDATA[ ...]]> which will survive most non-recursive XML operations. Characters such as [ are bad as they are common in regexes. So I’m wondering if there is a generally adopted approach rather than inventing our own.

A typical example is the “degrees” symbol:

HTML Entity (decimal)   &#176;
HTML Entity (hex)   &#xb0;
HTML Entity (named)     &deg;
How to type in Microsoft Windows    Alt +00B0
Alt 0176
Alt 248
UTF-8 (hex)     0xC2 0xB0 (c2b0)
UTF-8 (binary)  11000010:10110000
UTF-16 (hex)    0x00B0 (00b0)
UTF-16 (decimal)    176
UTF-32 (hex)    0x000000B0 (00b0)
UTF-32 (decimal)    176
C/C++/Java source code  "\u00B0"
Python source code  u"\u00B0"

We are also likely to encounter TeX

$10\,^{\circ}{\rm C}$

or

\degree

so backslashes, curlies and dollars are a poor idea.

We could for example use markup like:

__deg__
__#176__

and this will probably work but I’d appreciate advice from those who have similar problems.

update I accept @MichaelB’s insistence that we use UTF-8 throughout. I am worried that some of our tools may not conform and if so I’ll revisit this. Note that my original question is not well worded – read his answer and the link in it.

Report

Leave an answer
Cancel reply

You must login to add an answer.

Need An Account,

1 Answer

Editorial Team · Answer 1 · 2026-05-13T06:36:58+00:00

Get someone to do this who really understands character encodings. It looks like you don’t, because you’re not using the terminology correctly. Alternatively, read this.
Do not brew up your own escape scheme – it will cause you more problems than it will solve. Instead, normalize the various source encodings to UTF-8 (which is really just one such escape scheme, except efficient and standardized) and handle character encodings correctly. Perhaps use UTF-7 if you’re really that scared of high bits.
In this day and age, not handling character encodings correctly is not acceptable. If a tool doesn’t, abandon it – it is most likely very bad quality code in many other ways as well and not worth the hassle using.

Sign Up

Sign In

Forgot Password

The Archive Base Latest Questions

We are doing Natural Language Processing on a range of English language documents (mainly

Leave an answerCancel reply

1 Answer

Leave an answer
Cancel reply