Skip to content

Tokenizer protocol, char tokenizer

ModuleL1.1 · build · Python · Pass 3 · 2 to 3 h
You buildpython/tinyllm/tok/base.py: the Tokenizer protocol, check_ids, ByteTokenizer; python/tinyllm/tok/char.py: CharTokenizer (train, encode, decode, save, load)
Contractcourse/contracts/py/tinyllm/tok/base.pyi · course/contracts/py/tinyllm/tok/char.pyi · formats: tokenizer.md, tokenizer-char.schema.json
Testscourse/tests/L1.1/ (what they check: section 4) · your own tests in python/tests/l1-1-tokenizers/, rung R2, graded by mutation (threshold 0.60)
NeedsM05.2 the GPT-2 byte map (bytes_to_unicode, unicode_to_bytes) · reading: L0.0 the tracer’s byte ids (or --ref-deps)
Used byL1.2 byte-level BPE, L1.3 WordPiece, and L1.4 Unigram implement the protocol and call check_ids · L1.6 measures any Tokenizer · later: L1.5 the Rust tokenizer behind the same protocol, L3.6 char-level language models
MilestoneMS-L1 (tokenizers, part of the pass gate MS-P3)
Optional depthThe Unicode Standard, ch. 2 “General Structure” and ch. 3 section 3.9 (UTF-8); Jurafsky and Martin, Speech and Language Processing (3rd ed. draft), ch. 2 “Regular Expressions, Tokenization, Edit Distance”
  • A tokenizer is a contract, not an algorithm: every tokenizer in your system (bytes, char, BPE, WordPiece, Unigram, and later Rust) has the same six members, so everything that consumes tokens is written once (test_both_satisfy_the_protocol).
  • The byte tokenizer of the tracer is already a tokenizer: 256 ids, no training, every text round-trips, and its token strings are the GPT-2 byte map, the alphabet BPE starts from (test_hand_example_bytes, test_byte_tokens_are_the_gpt2_byte_map).
  • A char tokenizer’s ids must not depend on the order of the training texts: <unk> is 0, the specials follow, then code points sorted by code point (test_hand_example_vocab).
  • Unknown input is data, not an exception: an unseen character becomes <unk>, and special-token texts typed by a user stay ordinary characters (test_hand_example_encode_decode, test_special_text_is_not_matched).
  • Decoding never trusts ids and loading never trusts files: -1, vocab_size, True, and a malformed tinyllm_char.json are all ValueError (test_decode_rejects_bad_ids, test_load_rejects_invalid_files).
Terminal window
ol start L1.1 # stubs base.py and char.py into your repo, contracts alongside
ol tests L1.1 # read the test catalog first
ol check L1.1 # exit code is the verdict; then grades your tests by mutation
ol check L1.1 --ref-deps # only if your M05.2 is not passing yet
ol mutate L1.1 # the full mutation grade of python/tests/l1-1-tokenizers/
ol diff L1.1 # after passing: your code against the reference

This is rung R2 of the testing ladder: you write the bodies of the test names listed at the end of section 4, in python/tests/l1-1-tokenizers/. ol check runs them against the reference with one planted bug at a time, and at least 60% of the planted bugs must make one of your tests fail.


Your tracer speaks bytes. L0.0 trained on raw UTF-8 bytes and L10.0 serves tl_tokenizer = "bytes", so é costs two tokens and the model spends capacity learning how to spell UTF-8. Pass 3 replaces that with trained vocabularies: byte-level BPE (L1.2), WordPiece (L1.3), and Unigram (L1.4), compared by L1.6 and chosen for the capstone. Four tokenizers with four call conventions would mean four versions of the corpus pipeline, the metrics, and the engine’s tokenizer factory. This module fixes the shape first: one Tokenizer protocol, the byte tokenizer you already have written against it, and the simplest trained tokenizer, one id per character, which also becomes the vocabulary of the char-level models in L3.6.

SymbolMeaningType / shape
Σ\Sigmathe alphabet of Unicode code points, U+0000 to U+10FFFFset
t∈Σ∗t \in \Sigma^*a text: a finite sequence of code pointsstr
VVvocabulary size: ids are 0,…,V−10, \dots, V - 1int
enc:Σ∗→{0..V−1}∗\mathrm{enc}: \Sigma^* \to \{0..V-1\}^*encode: text to idsstr -> list[int]
dec:{0..V−1}∗→Σ∗\mathrm{dec}: \{0..V-1\}^* \to \Sigma^*decode: ids to textlist[int] -> str
uuthe unknown id (unk_id), or none when every text encodesint or None
SSthe special ids: <unk>, <bos>, <eos>, <pad>dict[str, int]
β:{0..255}→Σ\beta: \{0..255\} \to \Sigmathe GPT-2 byte map of M05.2 (byte 0x20 is Ġ)bijection

Text is code points, storage is bytes. A Python str is a sequence of code points; len("é") == 1. On disk and on the wire it is UTF-8: one to four bytes per code point, é (U+00E9) is c3 a9, an emoji is four bytes. Two texts that render the same may be different sequences: é can be the single code point U+00E9 or e followed by the combining acute accent U+0301. A tokenizer that silently normalizes one into the other can never give the original text back.

The protocol. A tokenizer is anything with these members (a typing.Protocol, so no inheritance is needed):

  • vocab_size: ids live in [0,V)[0, V), and every id in that range decodes (no gaps).
  • special_ids: the special tokens’ texts mapped to their ids.
  • unk_id: where unknown input goes, or None when every text encodes (bytes, byte-level BPE).
  • encode(text, add_special=False) and decode(ids, skip_special=False).
  • token_to_id, id_to_token: the vocabulary as a two-way table.
  • save(dir) and the class method load(dir): a tokenizer lives in a model directory.

The two laws are dec(enc(t))=t\mathrm{dec}(\mathrm{enc}(t)) = t for every text the vocabulary covers, and enc\mathrm{enc} never raises on a text: unknown input becomes uu. add_special adds a template around the ids (BOS and EOS here, [CLS] and [SEP] for WordPiece) and never changes the ids of the text itself. skip_special drops the ids in SS when decoding.

Ids are checked once, in one place. Every decode starts with check_ids(ids, V): each id must be an integer (hasattr(i, "__index__"), so numpy integers pass) and not a bool (in Python True == 1, and a mask passed by mistake would decode as text), and 0≤i<V0 \le i < V. The range check matters more than it looks: vocab[-1] is valid Python and returns the last entry, so without it a model that emits -1 decodes to a plausible character.

Bytes as a tokenizer. ByteTokenizer is the tracer’s tokenizer (D32) under the protocol: V=256V = 256, enc(t)\mathrm{enc}(t) is the UTF-8 bytes of tt, dec\mathrm{dec} is UTF-8 decoding with replacement, so a byte sequence that is not valid UTF-8 becomes U+FFFD instead of an exception. A sampled model can emit any bytes, and the server must still answer. Its token strings are β(i)\beta(i), the M05.2 byte map: printable code points for every byte, so id_to_token(32) is Ġ and id_to_token(10) is Ċ. That is exactly the alphabet byte-level BPE (L1.2) starts from, so “bytes” is BPE with zero merges.

The char tokenizer. One id per code point. Training only collects the code points that occur, and the vocabulary is fixed by a rule that does not depend on the order of the texts:

  1. id 0 is <unk>, always;
  2. then the requested specials, each one of <bos>, <eos>, <pad>, in the order given, a repeat counted once;
  3. then every code point of the training texts, sorted by code point.

Encoding maps each code point to its id, or to u=0u = 0 when it was never seen. The texts of the specials are not matched in the input: if a user could type <eos> and get the end-of-sequence id, any prompt could stop generation. add_special puts <bos> first and <eos> last when those roles exist.

The file. save(dir) writes dir/tinyllm_char.json (tokenizer.md, the char tokenizer): {"type": "char", "vocab": [...], "specials": {"unk": 0, ...}}. specials maps roles to ids, not texts to ids, so a reader in any language finds <unk> without guessing its spelling. load checks everything the constructor checks: unique entries, unk present, every role one of the four, every special id inside the vocabulary, every other entry exactly one code point.

Train. CharTokenizer.train(["hello", "hi"], specials=("<bos>", "<eos>")). The code points seen are h e l o i. Sorted by code point: e (101), h (104), i (105), l (108), o (111). With <unk> first and the two specials next:

id01234567
text<unk><bos><eos>ehilo

Training on ["hi", "hello"] gives the same table: the order of the texts does not matter. First-appearance order would have given h id 3 and e id 4, and a tokenizer retrained on shuffled data would silently disagree with every saved model.

Encode. "hello" is h e l l o, so [4, 3, 6, 6, 7]. "hi!" is [4, 5, 0]: ! was never seen, so it is <unk>. With add_special=True, "hi" is [1, 4, 5, 2].

Decode. [4, 5, 0] decodes to "hi<unk>" (the special’s text), and with skip_special=True to "hi". [-1] and [8] are ValueError.

Bytes. ByteTokenizer().encode("hé"): h is 0x68 = 104, é is U+00E9, two UTF-8 bytes c3 a9 = 195, 169. So [104, 195, 169]. Decoding [195] alone is an incomplete sequence: one U+FFFD, "�". Decoding [226, 130] (the first two bytes of a three-byte character) is also one "�": one maximal ill-formed subpart, one replacement character.

These are the first cases in section 4: test_hand_example_vocab, test_hand_example_encode_decode, and test_hand_example_bytes.

python/tinyllm/tok/base.py
@runtime_checkable
class Tokenizer(Protocol):
vocab_size: int
special_ids: dict[str, int]
unk_id: Optional[int]
def encode(self, text: str, add_special: bool = False) -> list[int]: ...
def decode(self, ids: Sequence[int], skip_special: bool = False) -> str: ...
def token_to_id(self, s: str) -> Optional[int]: ...
def id_to_token(self, i: int) -> str: ...
def save(self, dir: str) -> None: ...
@classmethod
def load(cls, dir: str) -> "Tokenizer": ...
def check_ids(ids: Sequence[int], vocab_size: int) -> list[int]
class ByteTokenizer: ... # vocab_size 256, special_ids {}, unk_id None
# python/tinyllm/tok/char.py
class CharTokenizer(Tokenizer):
vocab: list[str]; roles: dict[str, int]
def __init__(self, vocab: Sequence[str], specials: Mapping[str, int]) -> None
@classmethod
def train(cls, texts: Iterable[str], specials: Sequence[str] = ()) -> "CharTokenizer"

The contracts carry the exact rules: which inputs are ValueError, that ByteTokenizer.save creates the directory and writes nothing (its config.json says tl_tokenizer = "bytes"), and that token_to_id of the byte tokenizer is unicode_to_bytes()[s] from M05.2.

TestKINDChecksWhy it matters downstream
test_hand_example_vocabunitthe section 3 table, id for id, in either text ordersaved models and retrained tokenizers agree
test_hand_example_encode_decodeunit"hello" is [4, 3, 6, 6, 7], "hi!" ends in <unk>, decode with and without specialsyou and the test agree on the definition
test_hand_example_bytesunit"hé" is [104, 195, 169]; a partial character decodes to one U+FFFD; id 256 is rejectedthe tracer engine and L10.0 decode the same way
test_add_special_wraps_bos_eosunit<bos> first, <eos> last, only when the role existsL3.6 trains on wrapped sequences
test_no_normalizationboundaryU+00E9 and e + U+0301 get different ids and both round-tripyour corpus pipeline owns normalization, not the tokenizer
test_roundtrip_unicode_propertyproperty200 random texts across scripts and planes decode back exactlyany training text survives
test_special_text_is_not_matchedboundarytyping <eos> gives five character idsa user cannot end generation from the prompt
test_decode_rejects_bad_idsboundary-1, V, 1.0, True, "1" are ValueErrora corrupt id never decodes to a plausible character
test_save_writes_the_char_formatunitthe exact tinyllm_char.json document, roles to ids, and load gives the same idsany reader of the format finds <unk>
test_load_rejects_invalid_filesboundaryno unk, a duplicate, a two-code-point entry, a fifth role, an id past the end, another typea broken file fails at load, not mid-training
test_train_rejects_unknown_specialboundary<mask> is refused; a repeated <pad> and an explicit <unk> are counted oncethe format has exactly four roles
test_both_satisfy_the_protocolunitboth classes are Tokenizer instances with every attributeL1.6 and the corpus pipeline take any tokenizer
test_byte_tokens_are_the_gpt2_byte_mapunitid_to_token(32) is Ġ, and the map is a bijection on 0 to 255bytes are BPE with no merges (L1.2)
test_byte_roundtrip_propertyproperty200 random texts survive bytes encode and decodethe tracer’s tokenizer loses nothing

Write these in python/tests/l1-1-tokenizers/test_tokenizers.py, importing only names from the two contracts. Each docstring says what the test must establish; the bodies are yours.

def test_vocab_is_unk_specials_then_sorted_code_points():
"""train(["hello", "hi"], ("<bos>", "<eos>")) has vocab <unk> <bos> <eos> e h i l o."""
def test_unseen_character_is_unk():
"""An unseen character encodes to unk_id, it does not raise."""
def test_skip_special_drops_unk_bos_eos():
"""decode(..., skip_special=True) drops every special id, <unk> included."""
def test_decomposed_accent_round_trips():
""""e" + U+0301 survives encode and decode unchanged (no normalization)."""
def test_special_text_in_input_is_plain_text():
"""Typing "<eos>" gives five character ids, never the eos id."""
def test_negative_and_out_of_range_ids_raise():
"""decode([-1]) and decode([vocab_size]) are ValueError."""
def test_save_load_round_trip():
"""load(save(t)) encodes like t, and the file's specials map roles to ids."""
def test_load_rejects_two_character_entry():
"""A vocab entry of two code points, or a role outside unk/bos/eos/pad, is ValueError."""
def test_byte_tokenizer_utf8_and_replacement():
"""ByteTokenizer encodes UTF-8 bytes and decodes a lone continuation byte to U+FFFD."""
PitfallSymptomCaught by
1. normalizing in encode (NFC)"café" typed with a combining accent comes back as a different string; byte counts in L1.6 changetest_no_normalization (mutant s06)
2. vocabulary in first-appearance orderids depend on the order of the training files; a retrained tokenizer disagrees with the saved modeltest_hand_example_vocab (mutant s01)
3. a dictionary lookup without a defaultan emoji in a prompt raises KeyError inside the servertest_hand_example_encode_decode (mutant s03)
4. trusting vocab[i]-1 decodes to the last character; True decodes as id 1test_decode_rejects_bad_ids (mutants s09, m01, m02)
5. matching special texts in the inputa prompt containing <eos> stops generationtest_special_text_is_not_matched (mutant s08)
6. a loader that trusts the filea two-character entry or a made-up role loads and encodes wrong ids without a wordtest_load_rejects_invalid_files (mutants s11, s12)
7. decoding bytes with errors="ignore" or as Latin-1a partial character vanishes, or é decodes as étest_hand_example_bytes (mutants s15, s17)
8. chr(i) as the byte token stringbyte 32 is " " instead of Ġ, and BPE’s alphabet no longer matchestest_byte_tokens_are_the_gpt2_byte_map (mutant s16)

| Forward | L3.6 | Registered call site uses this module. |

DirectionModuleHow it uses this
BackM05.2bytes_to_unicode and unicode_to_bytes give the byte tokenizer its token strings
BackL0.0the tracer’s byte ids, now behind the protocol (reading)
ForwardL1.2byte-level BPE implements the protocol, calls check_ids, and starts from the same 256 byte symbols
ForwardL1.3WordPiece implements the protocol; add_special becomes [CLS] ... [SEP]
ForwardL1.4Unigram implements the protocol; <unk> is id 0 there too
ForwardL1.6fertility, bytes per token, and the fallback rate take any Tokenizer

L1.5 implements a Rust tokenizer independently; Python and Rust compare tokenization through shared tokenizer.json files and fixtures. L3.6 (char-level language models) arrives in a later batch and uses this protocol. If you skip this module, ol check L1.2 stops with needs L1.1: build it, or rerun with --ref-deps.

Your pieceProduction equivalentWhat it addsWhere to look
Tokenizer protocolHugging Face tokenizers Tokenizera pipeline object (normalizer, pre-tokenizer, model, post-processor, decoder) with offsets and paddingtokenizers/src/tokenizer/mod.rs
ByteTokenizerByT5, and llama.cpp’s byte fallback tokens <0x41>byte models at scale; byte fallback inside a subword vocabularyXue et al., ByT5 (2021); llama-vocab.cpp
CharTokenizerKarpathy’s char-level nanoGPT on Shakespearethe same vocabulary rule, used to train a small GPTnanoGPT/data/shakespeare_char/prepare.py
check_idstokenizers decoderejects ids past the vocabulary with an error, never wrapstokenizers/src/tokenizer/mod.rs decode