BPE Tokenizer is a TypeScript implementation of the byte pair encoding (BPE) algorithm used by OpenAI models. The library exposes three operations. The first operation converts text into token IDs, and the second operation converts those token IDs back into text. A third operation counts the tokens contained in a given text. Rank data is downloaded once from OpenAI's public blob storage and then cached on the local filesystem. Every subsequent load is therefore served from that local cache, so no network access is required.
git clone https://github.com/NeaByteLab/BPE-Tokenizer.gitimport BPE from '@neabyte/bpe-tokenizer'
// Returns the names of every available encoding.
const encodings = BPE.list()
// Loads an encoding. Repeat loads are served from cache.
const enc = await BPE.load('o200k_base')
// Converts text into token IDs.
const tokens = enc.encode('hello world', { allowed: 'all' })
// Converts token IDs back into text.
const text = enc.decode(tokens) // 'hello world'
// Counts tokens in a string.
const count = enc.count('hello world') // 2encode and count accept a BpeOptions object that governs how special tokens are treated. A special token is encoded as ordinary text unless it is explicitly permitted:
// Throws a TypeError. No special token is allowed.
enc.encode('<|endoftext|>')
// Permits every special token defined by the encoding.
enc.encode('<|endoftext|>', { allowed: 'all' })
// Permits one specific token and rejects all others.
enc.encode('<|endoftext|>', { allowed: ['<|endoftext|>'] })import BPE from '@neabyte/bpe-tokenizer'
// Loads the encoding.
const enc = await BPE.load('o200k_base')
// Encodes then decodes the input.
const input = 'Byte pair encoding is a tokenization algorithm'
const tokens = enc.encode(input, { allowed: 'all' })
const decoded = enc.decode(tokens)
// Prints the round trip result.
console.log(`Input: ${input}`)
console.log(`Tokens: [${tokens.join(', ')}]`)
console.log(`Count: ${tokens.length}`)
console.log(`Decoded: ${decoded}`)
console.log(`Match: ${input === decoded}`)Input: Byte pair encoding is a tokenization algorithm
Tokens: [10704, 10610, 24072, 382, 261, 6602, 2860, 22184]
Count: 8
Decoded: Byte pair encoding is a tokenization algorithm
Match: true
BPE.list()- Returns:
string[], which is the names of all seven available encodings as a fresh array on every call.
BPE.load(name, signal?)name<BpeName>is one of the seven supported encoding names.signal<AbortSignal>is optional and cancels a remote rank download.- Returns:
Promise<Encoding>, which is the encoding cached after the first successful load. - Throws the abort reason when
signalhas already aborted. A non-error reason becomes anAbortErrorDOM exception. - Throws
TypeErrorwhennameis not a string, is empty, or is padded with whitespace. - Throws
RangeErrorwhennameis a string but not a known encoding.
enc.encode(text, options?)text<string>is the text to encode.options<BpeOptions>is optional and controls special token handling.allowed<'all' | Iterable<string>>sets the special tokens to treat as valid. It defaults to none.disallowed<'all'>must remain'all', because a custom list cannot disable the guard. It defaults to'all'.
- Returns:
number[], which is the token IDs. - Throws
TypeErrorwhen the text contains an unpaired surrogate, a byte order mark, or a disallowed special token.
enc.decode(tokens)tokens<Iterable<number>>is the collection of token IDs to decode.- Returns:
string, which is the decoded text. - Throws
TypeErrorwhen the decoded bytes are not valid UTF-8.
enc.count(text, options?)text<string>is the text to count.options<BpeOptions>is optional and is identical to the options accepted byencode.- Returns:
number, which is the token count.
interface BpeOptions {
allowed?: 'all' | Iterable<string>
disallowed?: 'all' | Iterable<string>
}
type BpeName =
| 'o200k_base'
| 'o200k_harmony'
| 'cl100k_base'
| 'p50k_base'
| 'p50k_edit'
| 'r50k_base'
| 'gpt2'deno task checkdeno task testThis project reimplements the BPE algorithm and uses encoding data from OpenAI's tiktoken.
This project is distributed under the MIT license. See the LICENSE file for the full text.