Skip to content

TextEncoder.encodeInto() underfills the destination for some non-ASCII text #65994

Description

@koal44

There are two problems I found with TextEncoder.encodeInto(), both of which can cause encoding to stall even when the next char can fit the destination.

  1. A 2-byte char requires a 3-byte destination

    const encoder = new TextEncoder();
    const text = '\u0400'.repeat(33);
    
    console.log(encoder.encodeInto(text, new Uint8Array(2)));  // { read: 0, written: 0 }
    console.log(encoder.encodeInto(text, new Uint8Array(3)));  // { read: 1, written: 2 }

    The second call proves that '\u0400' should fit into a 2-byte array.

  2. Appending an unread character changes encoding progress

    const encoder = new TextEncoder();
    const text = 'é'.repeat(33);
    
    console.log(encoder.encodeInto(text, new Uint8Array(2)));  // { read: 0, written: 0 }
    console.log(encoder.encodeInto(text + '☺', new Uint8Array(2)));  // { read: 1, written: 2 }

    Appending ☺ should not change whether preceding chars can be read into the buffer, but there it is.

The bugs were introduced by the encodeInto() performance change in Node.js v25.4.0. The examples above use length 33 strings to exercise that optimized path(kSmallStringThreshold = 32). Unfortunately the current encodeInto.any.js WPT tests fail to expose the problems because:

  1. all input cases use 7 or fewer code units
  2. even then, the cases don't use chars between U+0400 and U+07FF, and
  3. their cases don't contain a narrow dst capacity to reveal the signed-byte problem.

Proposed fixes

src/encoding_binding.cc

  1. Incorrect cutoff in simpleUtfEncodingLength()

    -  if (c < 0x400) return 2;
    +  if (c < 0x800) return 2;

    (very likely a typo, given the comment immediately below it says "Code points < 0x800: 2 bytes")

  2. Signed-byte handling in findBestFit()

    -    size_t extra = simpleUtfEncodingLength(data[pos]);
    +    size_t extra = simpleUtfEncodingLength(UTF16 ? data[pos] : static_cast<uint8_t>(data[pos]));

Activity

  1. XadillaX commented on Sep 12, 2026

    @XadillaX
    Contributor

    Thanks for the detailed report and proposed changes

    I opened #65997 based on the proposal here, with tests for these cases. It also uses the replacement-aware UTF-16 API which is available now

    Please take a look when you have time

  2. koal44 commented on Sep 12, 2026

    @koal44
    Author

    thx, I can confirm the surrogate-pair underfill you found too. looks good to me.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions