in reply to Re^9: Mixed Unicode and ANSI string comparisons?
in thread Mixed Unicode and ANSI string comparisons?

Thanks for taking the time to consider the implications of my question seriously.

Given that description, any sense of "sorting" seems pretty meaningless.

Hm. They want a unified index. Generally, they want similar things to be found roughly together; and given something specific to look for, a rough idea of where to start looking. 'Ordered'? 'Collated'? I'm not sure that any other term is much better?

Apart from that, if there's some desire to "classify" or "cluster" the non-ASCII, non-Unicode strings, statistics on byte ngrams can help a fair bit with that (but it remains a bit of a research task, with some training of models required for classification).

With enough time and knowledge and money, I've no doubt that something along those lines could be done, but they do not have the money to fund such a project. They were looking for a quick fix and I was basically thinking aloud when I asked my question. I didn't anticipate the hostility people would have to answering such a simple question.


With the rise and rise of 'Social' network sites: 'Computers are making people easier to use everyday'
Examine what is said, not who speaks -- Silence betokens consent -- Love the truth but pardon error.
"Science is about questioning the status quo. Questioning authority". I knew I was on the right track :)
In the absence of evidence, opinion is indistinguishable from prejudice.
  • Comment on Re^10: Mixed Unicode and ANSI string comparisons?

Replies are listed 'Best First'.
Re^11: Mixed Unicode and ANSI string comparisons?
by Anonymous Monk on Dec 16, 2015 at 12:35 UTC
    Its the end of days

    of the year

    an auspicious time to view questions as accusation and let inferiority complex flare