Home > mailing lists

Re: Question regarding UTF-8 data and "C" collation on definition of field of table - Mailing list pgsql-general

From	Tom Lane
Subject	Re: Question regarding UTF-8 data and "C" collation on definition of field of table
Date	February 6, 2023 00:19:01
Msg-id	2556580.1675642741@sss.pgh.pa.us Whole thread Raw
In response to	Re: Question regarding UTF-8 data and "C" collation on definition of field of table (Dionisis Kontominas <dkontominas@gmail.com>)
Responses	Re: Question regarding UTF-8 data and "C" collation on definition of field of table Re: Question regarding UTF-8 data and "C" collation on definition of field of table
List	pgsql-general

Tree view

Dionisis Kontominas <dkontominas@gmail.com> writes:
>    I suppose that affects the outcome of ORDER BY clauses on the field,
> along with the content of the indexes. Is this right?

Yeah.

>    Assuming that the requirement exists, to store UTF-8 characters on a
> field that can be from multiple languages, and the database default
> encoding is UTF8 which is the right thing I suppose (please verify), what
> do you think should be the values of the Collation and Ctype for the
> database to behave correctly?

Um ... so define "correct".  If you have a mishmash of languages in the
same column, it's likely that they have conflicting rules about sorting,
and there may be no ordering that's not surprising to somebody.

If there's a predominant language in the data, selecting a collation
matching that seems like your best bet.  Otherwise, maybe you should
just shrug your shoulders and stick with C collation.  It's likely
to be faster than any alternative.

            regards, tom lane

pgsql-general by date:

From: Dionisis Kontominas
Date: 05 February 2023, 23:36:54
Subject: Re: Question regarding UTF-8 data and "C" collation on definition of field of table

From: Dionisis Kontominas
Date: 06 February 2023, 00:48:15
Subject: Re: Question regarding UTF-8 data and "C" collation on definition of field of table

Re: Question regarding UTF-8 data and "C" collation on definition of field of table - Mailing list pgsql-general

Previous

Next