Tokenization · Findings of EMNLP 2024
Tokenization falling short
On subword robustness in large language models
1Baidu2ModelBest3University of Copenhagen*Equal contribution
Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1582–1599
Overview
Language models read text as subword identifiers from a fixed vocabulary. That makes them sensitive to typographical errors and length variation, and largely blind to the characters inside each token. We call these issues the curse of tokenization.
TKEval measures how much the curse still affects current LLMs through three questions: can models solve problems that need access to characters, do they know the internal structure of their own tokens, and how robust are they to typographical variation? A final experiment tests whether subword regularization during training helps.
Findings
- Scale helps, but does not lift the curse.Larger models solve more character-level problems, yet every model struggles with characters counted from the end of a word and with non-contiguous patterns across tokens.
- Typos still bias models.Models are much more sensitive to noise than to reordering, and more to character-level changes than to token-level ones, regardless of size.
- Moderate subword regularization helps.Post-training with BPE-dropout at p = 0.2 often beats the baseline; high drop rates hurt.
Problem solving
Tasks whose answers depend on characters the tokenizer hides. All three come from BIG-bench.
Cycled letters
CLUndo a rotation of a word’s letters.
- In
remo- Out
more
Word unscrambling
WURecover a word from randomly scrambled letters.
- In
nad- Out
and
Identify math theorems
IMTRead a theorem written in LaTeX, decide whether it is true and, if not, pick the correct version. 54 problems in nine areas of mathematics; models choose by perplexity.
- In
- Let f ∈ L1(ℝ) be an integrable function. The span of {fa(x) = f(x + a) : a ∈ ℝ} is dense in L1(ℝ) if and only if f̂ …
- A
- … has no real roots. correct
- C
- … is irreducible over ℚ.
- D
- … has no repeated roots.
| Model | 0-shot | 1-shot | 2-shot | 3-shot |
|---|---|---|---|---|
| GPT-3 6B | 33.96 | 28.30 | 33.96 | 28.30 |
| GPT-3 200B | 32.08 | 30.19 | 33.96 | 30.19 |
| Llama 2 7B | 37.70 | 34.00 | 35.80 | 37.70 |
| Llama 3 8B | 41.51 | 45.28 | 45.28 | 35.85 |
| Llama 3 70B | 62.26 | 79.25 | 69.81 | 71.70 |
| Mistral 7B | 47.20 | 43.40 | 37.70 | 37.70 |
| Mixtral 8×7B | 49.10 | 56.60 | 64.20 | 62.30 |
Llama 3 70B leads at every shot count, but more examples do not help consistently, and GPT-3 200B is no better than GPT-3 6B. On the anagram tasks, GPT-4 Turbo outperforms every other model at every shot count, and all models struggle with longer words.
Token structure
Seven probes ask what a model knows about the characters inside its tokens and the substrings shared across them. Each has 200 test items, built from about 300 words.
Inside a token
Character count
CC- Q
- Which character appears 3 times in the word ‘messrs’?
- A
s
N-th character
NC- Q
- What is the 4th character of the word ‘myron’?
- A
o
N-th character from the end
NCR- Q
- What is the 2nd character from the end of the word ‘pensioner’?
- A
e
Case conversion
CCVConvert the characters of a word to upper case, lower case or title case.
Across tokens
Common substrings
CS- Q
- What are the common substrings of ‘critical’ and ‘conscious’?
- A
i,c
Longest common substring
LCS- Q
- What are the longest common substrings of ‘cow’ and ‘condition’?
- A
co
Longest common subsequence
LCSeq- Q
- What are the longest common subsequences of ‘illustrate’ and ‘critical’?
- A
ita
- Character count
- 0 → 81%
- Llama 3 8B from zero to three examples. Demonstrations help most from zero to one shot, then level off.
- N-th character
- 1 → 55%
- Llama 3 70B from zero to three examples.
- From the end
- 52%
- The best one-shot score on reverse character lookup, by GPT-4 Turbo. Counting from the end is harder for every model.
- Subsequences
- 1 → 4%
- Llama 3 8B on longest common subsequence. Even the best models struggle with non-contiguous patterns, far below their scores on common substrings.
Typos
Four perturbations are applied to the questions of GSM8K, MMLU, TruthfulQA and HumanEval, over n-grams of 2, 3 and 5; answers stay intact.
Shuffle characters within word boundaries (50% per n-gram), or insert, delete and replace characters (10% each).
Reorder tokens within an n-gram (50%), or insert, delete and replace tokens (30%).
Examples from the paper’s Figures 9 and 10; highlighted characters and tokens differ from the original. Token IDs are decoded with the Llama 3 vocabulary.
- Noise hurts more than reordering.Character-level noise degrades every model, regardless of size; reordering the same n-grams costs much less.
- Characters hurt more than tokens.Models handle token-level permutations better than character-level shuffles and noise.
- Longer n-grams are easier.As the n-gram grows from 2 to 5, scores tend to stabilize or recover. GPT-4 Turbo stays accurate under reordering at every n-gram size.
BPE-dropout
BPE-dropout randomly skips merges during tokenization, so the model sees the same word split in different ways. We post-train Mistral-7B with drop rates from 0 to 0.8 on about 111k synthesized token-structure examples.
- p = 0.2 works best.A moderate rate frequently surpasses the baseline, and consistently on case conversion, n-th character and longest common substring.
- High rates hurt.p = 0.6 and 0.8 score lower on most tasks, which the paper attributes to training for only five epochs.
- Sensitivity varies by task.Common substrings and character count are robust to dropout; n-th character, longest common substring and reverse lookup are the most sensitive.
- Model
- Mistral-7B, drop rates p ∈ {0, 0.2, 0.4, 0.6, 0.8}
- Training
- AdamW (β = 0.9, 0.95), peak learning rate 5e-5 with 10% warmup and cosine decay, 5 epochs, batch 16, context 4,096
- Evaluation
- Exact match on the seven token-structure probes, zero to three shots
Get started
# download the data
git lfs install
git clone https://huggingface.co/datasets/floatai/TKEval data/TKEval
# complex problem solving: cycled letters
python -u src/eval_rq1.py \
--model_path /path/to/Meta-Llama-3-8B \
--data_path data/TKEval/complex_problem_solving/cycled_letters_all_data_0123_shots.json \
--eval_metric generation \
--output_path output/cycled_letters.llama3-8b.json
Citation
@inproceedings{chai-etal-2024-tokenization,
title = "Tokenization Falling Short: On Subword Robustness in Large Language Models",
author = "Chai, Yekun and Fang, Yewei and Peng, Qiwei and Li, Xuhong",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.findings-emnlp.86/",
doi = "10.18653/v1/2024.findings-emnlp.86",
pages = "1582--1599"
}