Skip to content

Latest commit

 

History

272 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SimdUnicode

.NET

This is a fast C# library to validate UTF-8 strings and to make UTF-16 strings well formed.

Motivation

We seek to speed up the Utf8Utility.GetPointerToFirstInvalidByte function from the C# runtime library. The function is private in the Microsoft Runtime, but we can expose it manually. The C# runtime function is well optimized and it makes use of advanced CPU instructions. Nevertheless, we propose an alternative that can be several times faster.

Specifically, we provide the function SimdUnicode.UTF8.GetPointerToFirstInvalidByte which is a faster drop-in replacement:

// Returns &inputBuffer[inputLength] if the input buffer is valid.
/// <summary>
/// Given an input buffer <paramref name="pInputBuffer"/> of byte length <paramref name="inputLength"/>,
/// returns a pointer to where the first invalid data appears in <paramref name="pInputBuffer"/>.
/// The parameter <paramref name="Utf16CodeUnitCountAdjustment"/> is set according to the content of the valid UTF-8 characters encountered, counting -1 for each 2-byte character, -2 for each 3-byte and 4-byte characters.
/// The parameter <paramref name="ScalarCodeUnitCountAdjustment"/> is set according to the content of the valid UTF-8 characters encountered, counting -1 for each 4-byte character.
/// </summary>
/// <remarks>
/// Returns a pointer to the end of <paramref name="pInputBuffer"/> if the buffer is well-formed.
/// </remarks>
public unsafe static byte* GetPointerToFirstInvalidByte(byte* pInputBuffer, int inputLength, out int Utf16CodeUnitCountAdjustment, out int ScalarCodeUnitCountAdjustment);

The function uses advanced instructions (SIMD) on 64-bit ARM and x64 processors, but fallbacks on a conventional implementation on other systems. We provide extensive tests and benchmarks.

We apply the algorithm used by Node.js, Bun, Oracle GraalVM, by the PHP interpreter and other important systems. The algorithm has been described in the follow article:

Requirements

We recommend you install .NET 10 or better: https://dotnet.microsoft.com/en-us/download/dotnet/10.0

Running tests

dotnet test

To see which tests are running, we recommend setting the verbosity level:

dotnet test -v=normal

More details could be useful:

dotnet test -v d

To get a list of available tests, enter the command:

dotnet test --list-tests

To run specific tests, it is helpful to use the filter parameter:

dotnet test --filter TooShortErrorAvx2

Or to target specific categories:

dotnet test --filter "Category=scalar"

Running Benchmarks

To run the benchmarks, run the following command:

cd benchmark
dotnet run -c Release

To run just one benchmark, use a filter:

cd benchmark
dotnet run --configuration Release --filter "*Twitter*"
dotnet run --configuration Release --filter "*Lipsum*"

If you are under macOS or Linux, you may want to run the benchmarks in privileged mode:

cd benchmark
sudo dotnet run -c Release

Results (x64)

On an Intel Xeon Gold 6548N (.NET 10, AVX-512), our validation function is up to 14 times faster than the standard library on non-ASCII text. A realistic input is Twitter.json which is mostly ASCII with some Unicode content where we are 2.5 times faster.

data set SimdUnicode AVX-512 (GB/s) .NET speed (GB/s) speed up
Twitter.json 36 14 2.5 x
Arabic-Lipsum 15 3.3 4.5 x
Chinese-Lipsum 15 5.3 2.8 x
Emoji-Lipsum 15 1.1 14 x
Hebrew-Lipsum 15 3.2 4.5 x
Hindi-Lipsum 15 2.5 5.8 x
Japanese-Lipsum 15 4.5 3.2 x
Korean-Lipsum 15 1.6 9.1 x
Latin-Lipsum 72 107 ---
Russian-Lipsum 15 3.3 4.5 x

On pure ASCII (Latin-Lipsum), .NET 10's own UTF-8 scanner is faster. On mixed and multibyte input, SimdUnicode is substantially faster.

On x64 system, we offer several functions: a fallback function for legacy systems, a SSE4.2 function for older CPUs, an AVX2 function for current x64 systems and an AVX-512 function for the most recent processors (AMD Zen 4 or better, Intel Ice Lake, etc.).

Results (ARM)

On an Apple M2 system, our validation function is 1.5 to four times faster than the standard library.

data set SimdUnicode speed (GB/s) .NET speed (GB/s) speed up
Twitter.json 25 14 1.8 x
Arabic-Lipsum 7.4 3.5 2.1 x
Chinese-Lipsum 7.4 4.8 1.5 x
Emoji-Lipsum 7.4 2.5 3.0 x
Hebrew-Lipsum 7.4 3.5 2.1 x
Hindi-Lipsum 7.3 3.0 2.4 x
 Japanese-Lipsum 7.3 4.6  1.6 x
Korean-Lipsum 7.4 1.8 4.1 x
Latin-Lipsum 87 38 2.3 x
Russian-Lipsum 7.4 2.7 2.7 x

On a Graviton 3, our validation function is 1.2 to over five times faster than the standard library.

data set SimdUnicode speed (GB/s) .NET speed (GB/s) speed up
Twitter.json 19 11 1.7 x
Arabic-Lipsum 5.2 2.7 1.9 x
Chinese-Lipsum 5.2 4.5 1.2 x
Emoji-Lipsum 5.2 0.9 5.8 x
Hebrew-Lipsum 5.2 2.7 1.9 x
Hindi-Lipsum 5.2 2.4 2.2 x
 Japanese-Lipsum 5.2 3.9  1.3 x
Korean-Lipsum 5.2 1.5 3.5 x
Latin-Lipsum 57 26 2.2 x
Russian-Lipsum 5.2 2.8 1.9 x

On a Neoverse V1 (Graviton 3), our validation function is 1.3 to over five times faster than the standard library.

data set SimdUnicode speed (GB/s) .NET speed (GB/s) speed up
Twitter.json 14 8.7 1.4 x
Arabic-Lipsum 4.2 2.0 2.1 x
Chinese-Lipsum 4.2 2.6 1.6 x
Emoji-Lipsum 4.2 0.8 5.3 x
Hebrew-Lipsum 4.2 2.0 2.1 x
Hindi-Lipsum 4.2 1.6 2.6 x
 Japanese-Lipsum 4.2 2.4  1.8 x
Korean-Lipsum 4.2 1.3 3.2 x
Latin-Lipsum 42 17 2.5 x
Russian-Lipsum 4.2 0.95 4.4 x

On a Qualcomm 8cx gen3 (Windows Dev Kit 2023), we get roughly the same relative performance boost as the Neoverse V1.

data set SimdUnicode speed (GB/s) .NET speed (GB/s) speed up
Twitter.json 17 10 1.7 x
Arabic-Lipsum 5.0 2.3 2.2 x
Chinese-Lipsum 5.0 2.9 1.7 x
Emoji-Lipsum 5.0 0.9 5.5 x
Hebrew-Lipsum 5.0 2.3 2.2 x
Hindi-Lipsum 5.0 1.9 2.6 x
 Japanese-Lipsum 5.0 2.7  1.9 x
Korean-Lipsum 5.0 1.5 3.3 x
Latin-Lipsum 50 20 2.5 x
Russian-Lipsum 5.0 1.2 5.2 x

On a Neoverse N1 (Graviton 2), our validation function is 1.3 to over four times faster than the standard library.

data set SimdUnicode speed (GB/s) .NET speed (GB/s) speed up
Twitter.json 12 8.7 1.4 x
Arabic-Lipsum 3.4 2.0 1.7 x
Chinese-Lipsum 3.4 2.6 1.3 x
Emoji-Lipsum 3.4 0.8 4.3 x
Hebrew-Lipsum 3.4 2.0 1.7 x
Hindi-Lipsum 3.4 1.6 2.1 x
 Japanese-Lipsum 3.4 2.4  1.4 x
Korean-Lipsum 3.4 1.3 2.6 x
Latin-Lipsum 42 17 2.5 x
Russian-Lipsum 3.3 0.95 3.5 x

On a Neoverse N1 (Graviton 2), our validation function is up to over three times faster than the standard library.

data set SimdUnicode speed (GB/s) .NET speed (GB/s) speed up
Twitter.json 7.8 5.7 1.4 x
Arabic-Lipsum 2.5 0.9 2.8 x
Chinese-Lipsum 2.5 1.8 1.4 x
Emoji-Lipsum 2.5 0.7 3.6 x
Hebrew-Lipsum 2.5 0.9 2.7 x
Hindi-Lipsum 2.3 1.0 2.3 x
 Japanese-Lipsum 2.4 1.7  1.4 x
Korean-Lipsum 2.5 1.0 2.5 x
Latin-Lipsum 23 13 1.8 x
Russian-Lipsum 2.3 0.7 3.3 x

Well-formed UTF-16 strings

.NET strings may contain lone surrogates. SimdUnicode.UTF16.ToWellFormed replaces each lone surrogate by the replacement character U+FFFD, like JavaScript's String.prototype.toWellFormed(), and SimdUnicode.UTF16.IsWellFormed returns true when there is no lone surrogate (like isWellFormed()).

string s = UTF16.ToWellFormed("ab\uD800cd"); // "ab\uFFFDcd"
// Well-formed strings are returned as is, without allocation.
bool ok = UTF16.IsWellFormed(span);
// Buffer to buffer, or in place (same pointer for input and output).
UTF16.ToWellFormed(char* input, int length, char* output);
// Returns a pointer to the first lone surrogate, or input + length.
char* p = UTF16.GetPointerToFirstInvalidChar(char* input, int length);

We provide AVX-512, AVX2, SSE (SSE4.1) and ARM64 (NEON) kernels, selected at runtime, based on simdutf's algorithms described in the following article:

  • Robert Clausecker, Daniel Lemire, Fixing ill-formed UTF-16 strings with SIMD instructions, Software: Practice and Experience, 2026

Other systems fall back on the runtime's vectorized IndexOfAnyInRange.

We compare against the best approach we know that uses only public .NET APIs: return the input when it is well formed, otherwise copy it and fix the errors found with the vectorized IndexOfAnyInRange('\uD800', '\uDFFF'). It is fast on text without surrogates, but it stops at every surrogate pair, so it is slow on text such as emojis. The idiomatic string.Concat(s.EnumerateRunes()) runs at 0.2 GB/s to 0.4 GB/s, and a round trip through Encoding.Unicode at 1.3 GB/s to 7 GB/s.

To reproduce: dotnet run -c Release --filter "*UTF16WellFormed*" in the benchmark directory. All results are in GB/s of UTF-16 input, for well-formed inputs. "validate" is IsWellFormed (and ToWellFormed(string), which returns well-formed strings as is); "buffer" writes the output to a separate buffer; "short strings" cuts the input into strings of 1 to 64 code units and calls ToWellFormed(string) on each.

Intel Xeon Gold 6548N (.NET 10), AVX-512:

data set validate: SimdUnicode validate: IndexOfAnyInRange buffer: SimdUnicode buffer: copy + IndexOfAnyInRange plain copy short strings: SimdUnicode short strings: IndexOfAnyInRange
Twitter.json 69 31 26 15 28 9.2 7.1
Arabic-Lipsum 68 33 43 20 46 17 17
Chinese-Lipsum 115 39 43 20 46 17 16
Emoji-Lipsum 53 0.39 40 0.39 44 3.1 0.43
Hindi-Lipsum 68 33 43 21 45 17 17
Japanese-Lipsum 115 39 43 20 46 17 16
Korean-Lipsum 68 33 42 20 46 17 17
Latin-Lipsum 69 33 43 20 46 17 17
Russian-Lipsum 69 33 43 19 46 17 17

Same machine with AVX-512 disabled (DOTNET_EnableAVX512=0), AVX2 kernel (Haswell level):

data set validate: SimdUnicode validate: IndexOfAnyInRange buffer: SimdUnicode buffer: copy + IndexOfAnyInRange plain copy short strings: SimdUnicode short strings: IndexOfAnyInRange
Twitter.json 58 41 27 17 28 6.5 6.4
Arabic-Lipsum 58 42 43 23 47 17 17
Chinese-Lipsum 80 50 43 23 46 17 17
Emoji-Lipsum 35 0.50 28 0.49 45 3.4 0.53
Hindi-Lipsum 58 42 43 24 45 17 17
Japanese-Lipsum 77 50 43 24 46 17 17
Korean-Lipsum 58 42 43 23 46 17 17
Latin-Lipsum 58 42 43 24 46 17 17
Russian-Lipsum 58 42 43 22 47 17 17

Same machine with AVX disabled (DOTNET_EnableAVX=0), SSE kernel (Westmere level):

data set validate: SimdUnicode validate: IndexOfAnyInRange buffer: SimdUnicode buffer: copy + IndexOfAnyInRange plain copy short strings: SimdUnicode short strings: IndexOfAnyInRange
Twitter.json 50 28 26 15 26 6.1 6.1
Arabic-Lipsum 51 28 45 18 46 15 13
Chinese-Lipsum 63 28 43 18 46 15 13
Emoji-Lipsum 26 0.56 22 0.56 45 3.3 0.55
Hindi-Lipsum 50 28 43 17 45 15 13
Japanese-Lipsum 64 28 44 18 46 15 13
Korean-Lipsum 50 28 43 18 46 16 14
Latin-Lipsum 51 28 45 18 47 15 13
Russian-Lipsum 51 28 45 18 47 15 13

Apple M4 Max (.NET 10), NEON:

data set validate: SimdUnicode validate: IndexOfAnyInRange buffer: SimdUnicode buffer: copy + IndexOfAnyInRange plain copy short strings: SimdUnicode short strings: IndexOfAnyInRange
Twitter.json 106 62 71 35 84 23 21
Arabic-Lipsum 135 64 64 31 109 22 22
Chinese-Lipsum 134 64 66 40 75 22 22
Emoji-Lipsum 52 2.0 52 1.6 63 5.0 0.80
Hindi-Lipsum 135 64 73 38 85 23 23
Japanese-Lipsum 135 64 68 40 82 23 23
Korean-Lipsum 135 65 52 40 76 22 23
Latin-Lipsum 102 63 80 32 82 23 23
Russian-Lipsum 135 65 68 27 106 22 23
  • Validation is 1.4 to 3 times faster than IndexOfAnyInRange, and 25 to 140 times faster on emoji-heavy text.
  • Buffer to buffer, on x64 we are within 10% of the speed of a plain copy (except on emojis), and 1.3 to 2.5 times faster than copying and then scanning with IndexOfAnyInRange on all systems (30 to 110 times faster on emojis). On the M4 Max, we do not reach the speed of a plain copy.
  • On short strings, we are on par (within 5%) or faster, and 6 to 8 times faster on emojis.
  • When the string must be fixed (100 lone surrogates per million code units), both approaches are dominated by the allocation of the new string, except on emojis.

Building the library

cd src
dotnet build

Code format

We recommend you use dotnet format. E.g.,

dotnet format

Programming tips

You can print the content of a vector register like so:

        public static void ToString(Vector256<byte> v)
        {
            Span<byte> b = stackalloc byte[32];
            v.CopyTo(b);
            Console.WriteLine(Convert.ToHexString(b));
        }
        public static void ToString(Vector128<byte> v)
        {
            Span<byte> b = stackalloc byte[16];
            v.CopyTo(b);
            Console.WriteLine(Convert.ToHexString(b));
        }

Performance tips

  • Be careful: Vector128.Shuffle is not the same as Ssse3.Shuffle nor is Vector256.Shuffle the same as Avx2.Shuffle. Prefer the latter.
  • Similarly Vector128.Shuffle is not the same as AdvSimd.Arm64.VectorTableLookup, use the latter.
  • stackalloc arrays should probably not be used in class instances.
  • In C#, struct might be preferable to class instances as it makes it clear that the data is thread local.
  • You can ask for an asm dump: DOTNET_JitDisasm=NEON64HTMLScan dotnet run -c Release. See Viewing JIT disassembly and dumps.
  • You can get profiling data: dotnet run -c Release -- -p EP.

More reading

About

Fast SIMD-based UTF-8 Validation in C#

Resources

Stars

51 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages