Use shift based DFA for utf8 validation #41533
Comments
|
I guess there's also this: https://lemire.me/blog/2020/10/20/ridiculously-fast-unicode-utf-8-validation/, which may be implemented elsewhere in the ecosystem. |
|
I read through the Lemire paper for inspiration on vectorized string length that handles our way of counting invalid data as characters correctly. Can't find the issue now. |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
There's a post going around on Twitter for how to write a super fast DFA, e.g. for UTF8 validation: https://twitter.com/_rsc/status/1413843059972923394?s=19. We have a UTF8 validation function, so we should try this technique. Should be reasonably simple to do.
The text was updated successfully, but these errors were encountered: