Major performance degradation introduced with 10.3.3 (UTF8 char fetch) #1272
Description
Activity
Thanks @pl752 for the detailed report. I dug into this and can confirm both the regression and that your PR #1252 fixes it correctly. Here is what I found, with the supporting measurements.
Root cause
Commit bd0db88 (the fix for #1213) replaced the zero-allocation
RuneCount()/Substringtruncation withEnumerateRunesToChars(), which allocates onechar[]per code point and is then materialized with.ToList()on the read hot path. This runs on everyCHAR(n)fetch — even when no truncation is needed, which is the common case. OnlyDbDataType.Chargoes through this path (VarCharreturns the string directly), which is why aVARCHAR-based benchmark doesn't surface it: the regression only shows withCHAR(n) CHARACTER SET UTF8.The affected read sites are
GdsStatement.ReadRawValue/ReadRawValueAsyncandDbField.SetValue; the equivalentEnumerateRunesToChars().Count()is also used on the write/validate paths inDbValueandGdsStatement.WriteRawParameter*(10 call sites in total, all covered by your PR).Proof 1 — synthetic micro-benchmark (no DB, .NET 10)
Isolating just the truncation/count operation per field value:
Scenario old ( EnumerateRunesToChars)new (PR #1252) CHAR(100), value + space padding, no truncation (the hot path) 691.7 ns / 4176 B 1.9 ns / 0 B 120 chars → truncate to 100 1464 ns / 5536 B 15 ns / 224 B 80 emoji → truncate to 60 1011 ns / 4016 B 69 ns / 264 B count-only (write/validate path) 332 ns / 3320 B 4.6 ns / 0 B Even when no truncation happens, the old code allocates ~4 KB per value while the new code allocates nothing.
Proof 2 — end-to-end against a live Firebird 3.0 server (.NET 10)
I built two trees, current
mastervsmaster+ PR #1252, and ran the fetch benchmarks against the same server. Running yourLargeFetchBenchmark(CHAR(255) CHARACTER SET UTF8, 100,000 rows):current master + PR #1252 Mean 1,445.6 ms 1,007 ms Allocated 4339.37 MB 369.8 MB That reproduces your exact figures and the 11.73× allocation ratio. A smaller
CHAR(100) UTF8/ 100-row fetch shows the same shape: 1.76 MB → 192 KB (~9.4× less).Proof 3 — correctness (does not reintroduce #1213)
CountRunesandTruncateStringToRuneCountmatch thestring.EnumerateRunes()reference across all edge cases plus a 200,000-string fuzz (0 mismatches), including lone/unpaired surrogates.- The Improper truncation when reading from UTF-8 CHAR(n) fields containing characters outside of the Basic Multilingual Plane #1213 scenarios behave correctly: a
CHAR(1)holding😊returns the full surrogate pair (one rune), not a truncated high surrogate; aCHAR(2)holding😊!returns😊!. - The existing server tests
HighLowSurrogatePassingTestandHighLowSurrogateReadingTestpass with the PR applied. - The PR compiles cleanly on current
master. Note thatExtensions.cson master now uses the newextension(...)member syntax; the PR's classicthis ReadOnlySpan<char>methods coexist with it without issue.
The only behavioral difference I could find between old and new is on invalid UTF-16 (a lone surrogate) and only when truncation actually occurs: the old code normalizes it to U+FFFD while the new code preserves the original char. That is not reachable for valid UTF8 data and is unrelated to #1213.
Verdict
PR #1252 is correct and restores the pre-regression allocation behavior (~11.7× less for the UTF8
CHARcase) without reintroducing #1213. Two minor, non-blocking cleanups: after the PR,EnumerateRunesToCharsis no longer used and could be removed; and the two new methods could be moved into anextension(...)block to match the current style ofExtensions.cs.I'm happy to share the benchmark and correctness harness, and the exact patch I used, if that would be helpful.
- added a commit that references this issue
on May 23, 2026 @pl752 I’ve updated #1203 if you’d like to take a look. Feedback and suggestions are very welcome.
@fdcastel Added char utf-8 coverage, of course, reproduces the issue successfully too
Hello, @cincuranet , I was fiddling with @fdcastel `s PR (#1203) when I have noticed that between used in his PR nuget 10.3.1 and master there was an update (10.3.3) which significantly affected performance and allocation rates when woring with UTF8 strings.
Interesting fact is that the update was fixing an issue around the same place where one of my old PRs (#1252) was implementing optimizations directly related to affected methods. Related issue resulting in problematic fix bd0db88: #1213
Table relaed to my PR vs 10.3.1
Can you, please, review my PR changes and check if they are correct (aka they won't reintroduce the fixed issue, tests are passing successfully meanwhile)?