in reply to Re^9: Mixed Unicode and ANSI string comparisons?
in thread Mixed Unicode and ANSI string comparisons?
Thanks for taking the time to consider the implications of my question seriously.
Given that description, any sense of "sorting" seems pretty meaningless.
Hm. They want a unified index. Generally, they want similar things to be found roughly together; and given something specific to look for, a rough idea of where to start looking. 'Ordered'? 'Collated'? I'm not sure that any other term is much better?
Apart from that, if there's some desire to "classify" or "cluster" the non-ASCII, non-Unicode strings, statistics on byte ngrams can help a fair bit with that (but it remains a bit of a research task, with some training of models required for classification).
With enough time and knowledge and money, I've no doubt that something along those lines could be done, but they do not have the money to fund such a project. They were looking for a quick fix and I was basically thinking aloud when I asked my question. I didn't anticipate the hostility people would have to answering such a simple question.
|
|---|
| Replies are listed 'Best First'. | |
|---|---|
|
Re^11: Mixed Unicode and ANSI string comparisons?
by Anonymous Monk on Dec 16, 2015 at 12:35 UTC |