On 2026-09-15 Tu 8:51 AM, Nazir Bilal Yavuz wrote:
> Hi,
>
> Thank you for working on this!
>
> On Tue, 15 Sept 2026 at 14:31, Andrew Dunstan <andrew@dunslane.net> wrote:
>> On 2026-09-15 Tu 12:22 AM, Chao Li wrote:
>>> The attached is my test script.
>>>
>> Great, thanks for the review and tests.
> What do you think about continuing from where text_ascii_check() is
> left? I wrote a patch for this and benchmarked with Chao's script.
>
> Timings are v1 vs v2, not master vs v2.
>
> # unicode_is_normalized()
>
> * all ascii: 89.115ms | 87.803ms
> * mixed: 249.239ms | 126.011ms -> improvement
> * non-ascii: 1243.496ms | 1244.052ms
> * late-non-ascii: 3632.410ms | 370.966ms -> improvement
>
> # unicode_normalize_func()
>
> * all ascii: 84.458ms | 84.636ms
> * mixed: 578.684ms | 226.093ms -> improvement
> * non-ascii: 3651.573ms | 3644.026ms
> * late-non-ascii: 10513.861ms | 942.470ms -> improvement
>
> # unicode_assigned()
>
> * all ascii: 61.147ms | 60.658ms
> * mixed: 124.193ms | 86.186ms -> improvement
> * non-ascii: 507.130ms | 510.247ms
> * late-non-ascii: 1169.617ms | 166.924ms -> improvement
>
> Do you think these results worth the additional complexity?
Yes, I think it is - it's not that much extra. This is neat, in effect
avoiding duplicating the work.
cheers
andrew
--
Andrew Dunstan
EDB: https://www.enterprisedb.com