Tokenization · Findings of EMNLP 2024

Tokenization falling short

On subword robustness in large language models

Yekun Chai1*, Yewei Fang2*, Qiwei Peng3, Xuhong Li1

1Baidu2ModelBest3University of Copenhagen*Equal contribution

Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1582–1599

3research questions on the curse of tokenization
7token-structure probes with 1,400 test items
4typo perturbations on GSM8K, MMLU, TruthfulQA and HumanEval
6LLMs, from Mistral-7B to GPT-4 Turbo
Two angle diagrams. The embedding of “assignment” and the sum of the embeddings of “assign” and “ment” are 78.16 degrees apart, cosine 0.21. “import” and “im” plus “port” are 82.47 degrees apart, cosine 0.13. “assignment” “assign” + “ment” 78.16° cosine 0.21 “import” “im” + “port” 82.47° cosine 0.13
A word’s embedding against the sum of the embeddings of its subword pieces (the paper’s Figure 1). Nearly orthogonal vectors mean the model does not compose a word’s surface form from its tokens.

Overview

Language models read text as subword identifiers from a fixed vocabulary. That makes them sensitive to typographical errors and length variation, and largely blind to the characters inside each token. We call these issues the curse of tokenization.

TKEval measures how much the curse still affects current LLMs through three questions: can models solve problems that need access to characters, do they know the internal structure of their own tokens, and how robust are they to typographical variation? A final experiment tests whether subword regularization during training helps.

Findings

  • Scale helps, but does not lift the curse.Larger models solve more character-level problems, yet every model struggles with characters counted from the end of a word and with non-contiguous patterns across tokens.
  • Typos still bias models.Models are much more sensitive to noise than to reordering, and more to character-level changes than to token-level ones, regardless of size.
  • Moderate subword regularization helps.Post-training with BPE-dropout at p = 0.2 often beats the baseline; high drop rates hurt.

Problem solving

Tasks whose answers depend on characters the tokenizer hides. All three come from BIG-bench.

Cycled letters

CL

Undo a rotation of a word’s letters.

In
remo
Out
more

Word unscrambling

WU

Recover a word from randomly scrambled letters.

In
nad
Out
and

Identify math theorems

IMT

Read a theorem written in LaTeX, decide whether it is true and, if not, pick the correct version. 54 problems in nine areas of mathematics; models choose by perplexity.

In
Let f ∈ L1(ℝ) be an integrable function. The span of {fa(x) = f(x + a) : a ∈ ℝ} is dense in L1(ℝ) if and only if f̂ …
A
… has no real roots. correct
C
… is irreducible over ℚ.
D
… has no repeated roots.
Exact match (%) on Identify Math Theorems by number of in-context examples (Table 1). GPT-3 rows are from Srivastava et al. (2022).
Model0-shot1-shot2-shot3-shot
GPT-3 6B33.9628.3033.9628.30
GPT-3 200B32.0830.1933.9630.19
Llama 2 7B37.7034.0035.8037.70
Llama 3 8B41.5145.2845.2835.85
Llama 3 70B62.2679.2569.8171.70
Mistral 7B47.2043.4037.7037.70
Mixtral 8×7B49.1056.6064.2062.30

Llama 3 70B leads at every shot count, but more examples do not help consistently, and GPT-3 200B is no better than GPT-3 6B. On the anagram tasks, GPT-4 Turbo outperforms every other model at every shot count, and all models struggle with longer words.

Token structure

Seven probes ask what a model knows about the characters inside its tokens and the substrings shared across them. Each has 200 test items, built from about 300 words.

Inside a token

Character count

CC
Q
Which character appears 3 times in the word ‘messrs’?
A
s

N-th character

NC
Q
What is the 4th character of the word ‘myron’?
A
o

N-th character from the end

NCR
Q
What is the 2nd character from the end of the word ‘pensioner’?
A
e

Case conversion

CCV

Convert the characters of a word to upper case, lower case or title case.

Across tokens

Common substrings

CS
Q
What are the common substrings of ‘critical’ and ‘conscious’?
A
i, c

Longest common substring

LCS
Q
What are the longest common substrings of ‘cow’ and ‘condition’?
A
co

Longest common subsequence

LCSeq
Q
What are the longest common subsequences of ‘illustrate’ and ‘critical’?
A
ita
Character count
0 → 81%
Llama 3 8B from zero to three examples. Demonstrations help most from zero to one shot, then level off.
N-th character
1 → 55%
Llama 3 70B from zero to three examples.
From the end
52%
The best one-shot score on reverse character lookup, by GPT-4 Turbo. Counting from the end is harder for every model.
Subsequences
1 → 4%
Llama 3 8B on longest common subsequence. Even the best models struggle with non-contiguous patterns, far below their scores on common substrings.

Typos

Four perturbations are applied to the questions of GSM8K, MMLU, TruthfulQA and HumanEval, over n-grams of 2, 3 and 5; answers stay intact.

Character levelPermutation and noise

Shuffle characters within word boundaries (50% per n-gram), or insert, delete and replace characters (10% each).

Token levelPermutation and noise

Reorder tokens within an n-gram (50%), or insert, delete and replace tokens (30%).

HumanEval · character level · 5-grams

          
GSM8K · token level · 3-grams

Examples from the paper’s Figures 9 and 10; highlighted characters and tokens differ from the original. Token IDs are decoded with the Llama 3 vocabulary.

  • Noise hurts more than reordering.Character-level noise degrades every model, regardless of size; reordering the same n-grams costs much less.
  • Characters hurt more than tokens.Models handle token-level permutations better than character-level shuffles and noise.
  • Longer n-grams are easier.As the n-gram grows from 2 to 5, scores tend to stabilize or recover. GPT-4 Turbo stays accurate under reordering at every n-gram size.

BPE-dropout

BPE-dropout randomly skips merges during tokenization, so the model sees the same word split in different ways. We post-train Mistral-7B with drop rates from 0 to 0.8 on about 111k synthesized token-structure examples.

  • p = 0.2 works best.A moderate rate frequently surpasses the baseline, and consistently on case conversion, n-th character and longest common substring.
  • High rates hurt.p = 0.6 and 0.8 score lower on most tasks, which the paper attributes to training for only five epochs.
  • Sensitivity varies by task.Common substrings and character count are robust to dropout; n-th character, longest common substring and reverse lookup are the most sensitive.
Model
Mistral-7B, drop rates p ∈ {0, 0.2, 0.4, 0.6, 0.8}
Training
AdamW (β = 0.9, 0.95), peak learning rate 5e-5 with 10% warmup and cosine decay, 5 epochs, batch 16, context 4,096
Evaluation
Exact match on the seven token-structure probes, zero to three shots

Get started

# download the data
git lfs install
git clone https://huggingface.co/datasets/floatai/TKEval data/TKEval

# complex problem solving: cycled letters
python -u src/eval_rq1.py \
  --model_path /path/to/Meta-Llama-3-8B \
  --data_path data/TKEval/complex_problem_solving/cycled_letters_all_data_0123_shots.json \
  --eval_metric generation \
  --output_path output/cycled_letters.llama3-8b.json

Citation

@inproceedings{chai-etal-2024-tokenization,
  title     = "Tokenization Falling Short: On Subword Robustness in Large Language Models",
  author    = "Chai, Yekun and Fang, Yewei and Peng, Qiwei and Li, Xuhong",
  booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
  month     = nov,
  year      = "2024",
  address   = "Miami, Florida, USA",
  publisher = "Association for Computational Linguistics",
  url       = "https://aclanthology.org/2024.findings-emnlp.86/",
  doi       = "10.18653/v1/2024.findings-emnlp.86",
  pages     = "1582--1599"
}