Home > mailing lists

Re: C11: should we use char32_t for unicode code points? - Mailing list pgsql-hackers

From	Tatsuo Ishii
Subject	Re: C11: should we use char32_t for unicode code points?
Date	October 28, 2025 11:36:13
Msg-id	20251028.173613.18179479132562731.ishii@postgresql.org Whole thread Raw
In response to	Re: C11: should we use char32_t for unicode code points? (Thomas Munro <thomas.munro@gmail.com>)
List	pgsql-hackers

Tree view

> The EUC family has direct encoding of 7-bit ASCII and then 3
> selectable character sets represented by sequences with the high bit
> set, with details varying between the Chinese (simplified Chinese),
> Taiwanese (traditional Chinese), Japanese (2 kinds) and Korean
> variants.  I don't know if the pg_wchar encoding we're producing in
> pg_euc*2wchar_with_len() has a name, but it doesn't appear to match
> the description of the standard "fixed" representation on the
> Wikipedia page for Extended Unix Code (it's too wide for starters,
> looking at the shift distances).

Yes. pg_euc*2wchar_with_len() creates "variable length" representation
of EUC, 1 byte to 4 bytes range per character. Then, expands each
character into pg_wchar. Also it can be converted back to the
multibyte representation easily.

Note that the standard "fixed" representation of EUC includes ASCII
range bytes in *non* ASCII characters, thus I think it is not easy to
use for backend safe encoding.

Best regards,
--
Tatsuo Ishii
SRA OSS K.K.
English: http://www.sraoss.co.jp/index_en/
Japanese:http://www.sraoss.co.jp

pgsql-hackers by date:

From: Bertrand Drouvot
Date: 28 October 2025, 11:13:06
Subject: Consistently use the XLogRecPtrIsInvalid() macro

From: torikoshia
Date: 28 October 2025, 11:43:49
Subject: Re: RFC: Allow EXPLAIN to Output Page Fault Information

Re: C11: should we use char32_t for unicode code points? - Mailing list pgsql-hackers

Previous

Next