Skip to content

Latest commit

 

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

BPE Tokenizer

Token encoding and decoding library for OpenAI model vocabularies

Deno Module type: Deno/ESM License

What is BPE Tokenizer?

BPE Tokenizer is a TypeScript implementation of the byte pair encoding (BPE) algorithm used by OpenAI models. The library exposes three operations. The first operation converts text into token IDs, and the second operation converts those token IDs back into text. A third operation counts the tokens contained in a given text. Rank data is downloaded once from OpenAI's public blob storage and then cached on the local filesystem. Every subsequent load is therefore served from that local cache, so no network access is required.

Installation

git clone https://github.com/NeaByteLab/BPE-Tokenizer.git

Usage

import BPE from '@neabyte/bpe-tokenizer'

// Returns the names of every available encoding.
const encodings = BPE.list()

// Loads an encoding. Repeat loads are served from cache.
const enc = await BPE.load('o200k_base')

// Converts text into token IDs.
const tokens = enc.encode('hello world', { allowed: 'all' })

// Converts token IDs back into text.
const text = enc.decode(tokens) // 'hello world'

// Counts tokens in a string.
const count = enc.count('hello world') // 2

Special Token Handling

encode and count accept a BpeOptions object that governs how special tokens are treated. A special token is encoded as ordinary text unless it is explicitly permitted:

// Throws a TypeError. No special token is allowed.
enc.encode('<|endoftext|>')

// Permits every special token defined by the encoding.
enc.encode('<|endoftext|>', { allowed: 'all' })

// Permits one specific token and rejects all others.
enc.encode('<|endoftext|>', { allowed: ['<|endoftext|>'] })

Example

import BPE from '@neabyte/bpe-tokenizer'

// Loads the encoding.
const enc = await BPE.load('o200k_base')

// Encodes then decodes the input.
const input = 'Byte pair encoding is a tokenization algorithm'
const tokens = enc.encode(input, { allowed: 'all' })
const decoded = enc.decode(tokens)

// Prints the round trip result.
console.log(`Input:    ${input}`)
console.log(`Tokens:   [${tokens.join(', ')}]`)
console.log(`Count:    ${tokens.length}`)
console.log(`Decoded:  ${decoded}`)
console.log(`Match:    ${input === decoded}`)
Input:    Byte pair encoding is a tokenization algorithm
Tokens:   [10704, 10610, 24072, 382, 261, 6602, 2860, 22184]
Count:    8
Decoded:  Byte pair encoding is a tokenization algorithm
Match:    true

API

list

BPE.list()
  • Returns: string[], which is the names of all seven available encodings as a fresh array on every call.

load

BPE.load(name, signal?)
  • name <BpeName> is one of the seven supported encoding names.
  • signal <AbortSignal> is optional and cancels a remote rank download.
  • Returns: Promise<Encoding>, which is the encoding cached after the first successful load.
  • Throws the abort reason when signal has already aborted. A non-error reason becomes an AbortError DOM exception.
  • Throws TypeError when name is not a string, is empty, or is padded with whitespace.
  • Throws RangeError when name is a string but not a known encoding.

encode

enc.encode(text, options?)
  • text <string> is the text to encode.
  • options <BpeOptions> is optional and controls special token handling.
    • allowed <'all' | Iterable<string>> sets the special tokens to treat as valid. It defaults to none.
    • disallowed <'all'> must remain 'all', because a custom list cannot disable the guard. It defaults to 'all'.
  • Returns: number[], which is the token IDs.
  • Throws TypeError when the text contains an unpaired surrogate, a byte order mark, or a disallowed special token.

decode

enc.decode(tokens)
  • tokens <Iterable<number>> is the collection of token IDs to decode.
  • Returns: string, which is the decoded text.
  • Throws TypeError when the decoded bytes are not valid UTF-8.

count

enc.count(text, options?)
  • text <string> is the text to count.
  • options <BpeOptions> is optional and is identical to the options accepted by encode.
  • Returns: number, which is the token count.

Types

interface BpeOptions {
  allowed?: 'all' | Iterable<string>
  disallowed?: 'all' | Iterable<string>
}

type BpeName =
  | 'o200k_base'
  | 'o200k_harmony'
  | 'cl100k_base'
  | 'p50k_base'
  | 'p50k_edit'
  | 'r50k_base'
  | 'gpt2'

Build

deno task check

Testing

deno task test

Acknowledgements

This project reimplements the BPE algorithm and uses encoding data from OpenAI's tiktoken.

License

This project is distributed under the MIT license. See the LICENSE file for the full text.

Contributors

Languages