Skip to content

3x performance, resolved memory allocation problem and correct handling of double-character codepoints - #16

Merged
thecoderok merged 10 commits into
thecoderok:masterfrom
csm101:array_instead_of_dictionary
Sep 5, 2023
Merged

thecoderok merged 10 commits into
thecoderok:masterfrom
csm101:array_instead_of_dictionary

Conversation

@csm101

@csm101 csm101 commented Jul 2, 2023

Copy link
Copy Markdown

I added a benchmark to the solution (using BenchmarkDotNet).
these are the benchmark results in my modified branch:

|               Method |      Mean |     Error |    StdDev |   Gen0 | Allocated |
|--------------------- |----------:|----------:|----------:|-------:|----------:|
|     UnidecodeRussian | 42.312 ns | 0.4374 ns | 0.4092 ns | 0.0038 |      64 B |
|       UnidecodeAscii | 15.640 ns | 0.0360 ns | 0.0319 ns |      - |         - |
| UnidecodeRussianChar |  3.132 ns | 0.0133 ns | 0.0124 ns |      - |         - |
|   UnidecodeAsciiChar |  2.673 ns | 0.0167 ns | 0.0156 ns |      - |         - |

and these are the result of the same banchmark when run on the original unmodified sources

|               Method |       Mean |     Error |    StdDev |   Gen0 | Allocated |
|--------------------- |-----------:|----------:|----------:|-------:|----------:|
|     UnidecodeRussian | 145.070 ns | 2.4058 ns | 2.2504 ns | 0.0148 |     248 B |
|       UnidecodeAscii |  71.189 ns | 0.8209 ns | 0.7679 ns | 0.0019 |      32 B |
| UnidecodeRussianChar |   5.627 ns | 0.0453 ns | 0.0423 ns |      - |         - |
|   UnidecodeAsciiChar |   8.355 ns | 0.0546 ns | 0.0456 ns | 0.0014 |      24 B |

I obtained this by implementing these modifications:

  1. I used a string[][] array instead of a Dictionary<int,string[]> for Unidecoder.Characters (i modified the source code python generator script accordingly).
  2. if the translation of the input string will surely generate a string shorter than 8192 characters, I use an optimized Unidecode() that uses a stack-allocated buffer (Span stackBuffer = stackalloc char[8192)]. instead of a stringbuilder. It was essential to apply the [SkipLocalsInit] attribute to such version to tell the compiler that i didn't need that buffer to be zeroed: zeroing 8192 bytes had a significative impact. Using this attribute requires the code to be compiled allowing compilation of unsafe code... but It is only for that attribute, I am not doing weird things with pointers.
  3. I optimized also the single char Unidecode() char extension in order to avoid allocating a new single character string every time it decodes a char whose codepoint is <0x80 (for all other cases it was already using a const string taken from the characters dictionary... the only case when it was creating a new string for each call was for the "easy" case of an ascii character.

I also added an helper method that does the translation from indexes found in the decoded string to indexes in the original undecoded string.

I needed this method because I use unidecode for searching records that match a search string, but I need to visually highlight the matching occourrences in the original not "unidecoded" text. I think it might be useful also to other people, so I added it to Unidecoder.

the only thing that needs to be checked if it compiles for all the platform you want to support. I cared/checked only for net 7, honestly.

@csm101 csm101 mentioned this pull request Jul 3, 2023
@csm101

csm101 commented Jul 3, 2023 •

Copy link
Copy Markdown
Author

I went on coding and fixed issues #14 and #15 .
for #14 i scrapped the Unidecode.Character.cs source file and initialized the characters by reading data from a new resource file linked in the assembly
for #15, since proper handling of all codepoints has a certain performance impact, I introduce a UnidecodeAlgorithm enum that lets you choose between the "fast" implementation and the "Complete", but slower implementation.

To be precise now you can do these things:

  "xyz".Unidecode() // uses one of the two algorithms depending on the static property Unidecode.Algorithm
  "xyz".Unidecode(UnidecodeAlgorithm.Fast); // uses the exact algorithm specified in the argument
  
  Unidecode.Algorithm = UnidecodeAlgorithm.Complete; //sets the algoritm for the parameterless version to "complete"

I also added a test case for the example provided in issue #15

these are the update benchmark results:

|                   Method |       Mean |     Error |    StdDev |   Gen0 | Allocated |
|------------------------- |-----------:|----------:|----------:|-------:|----------:|
|     FastUnidecodeRussian |  41.808 ns | 0.2353 ns | 0.2201 ns | 0.0038 |      64 B |
| CompleteUnidecodeRussian | 108.688 ns | 2.1999 ns | 2.5335 ns | 0.0148 |     248 B |
|       FastUnidecodeAscii |  13.860 ns | 0.1878 ns | 0.1756 ns |      - |         - |
|   CompleteUnidecodeAscii |  66.209 ns | 0.9575 ns | 0.8956 ns | 0.0019 |      32 B |
|     UnidecodeRussianChar |   2.400 ns | 0.0662 ns | 0.0619 ns |      - |         - |
|       UnidecodeAsciiChar |   1.338 ns | 0.0033 ns | 0.0031 ns |      - |         - |
|     UnidecodeRussianRune |   1.747 ns | 0.0134 ns | 0.0125 ns |      - |         - |
|       UnidecodeAsciiRune |   1.436 ns | 0.0168 ns | 0.0157 ns |      - |         - |

@csm101 csm101 changed the title 3x performance improvement 3x performance, resolved memory allocation problem and correct handling of double-character codepoints Jul 5, 2023
@phnx47

phnx47 commented Sep 5, 2023

Copy link
Copy Markdown
Contributor

@csm101 Hi! I don't support this library anymore. I removed myself from repository. Try reach out repository owner - @thecoderok.

@thecoderok
thecoderok merged commit e119903 into thecoderok:master Sep 5, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants