Pretraining · October 2026
What is a repeated token worth?
The scaling geometry of multi-epoch pretraining
1FloatAI2ETH Zurich3Independent Researcher
arXiv preprint 2610.05591
a A second epoch is nearly free
b Half value after four to five epochs
c Scale epochs with data, not size
Takeaways
- A second epoch is nearly free.Measured against one epoch on the same data, a repeated token in a second epoch is worth 0.87–1.02 fresh tokens at every budget. Each further epoch is worth less, and the drop is steepest at small budgets.
- Larger budgets per parameter tolerate more repetition.Repeats fall to half the value of fresh tokens after about four to five epochs up to 20 tokens per parameter, and about ten at 300.
- Transfer epoch counts with data, not size.With unique data fixed, larger models should take fewer epochs; when unique data grow with the model, the best epoch count barely changes. To reuse an epoch count on a larger model, scale the unique data with it.
- Grow model and epochs together, up to the critical epoch count.With unique data fixed, extra compute goes to both in roughly equal shares until loss stops improving. Ignoring the cost of repetition loses at most 0.002 bpb up to that point.
- Spread and shuffle repeats.Space repeats across epochs and samples, repeat lower-entropy sources less, and re-tokenize only under heavy repetition. Compare such choices at separately annealed horizons; intermediate checkpoints can misrank them.
Overview
As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters.
We answer them by pricing a repeated token against two references at the same model size. Against fresh data at equal compute, the cost of repetition follows a single variable: the number of extra epochs divided by the unique tokens per parameter. Against the same data seen once, a second epoch is worth nearly as much as a fresh one, and repeated tokens fall to half the value of fresh ones after a critical epoch count that grows with the training budget per parameter but hardly with model size.
The same variable explains size trends that appear to conflict in earlier work. Larger models tolerate fewer epochs when the corpus is fixed, but not when unique data grow with the model. And counts alone do not determine loss: the order, allocation and source of the repeats all change the result.
Two prices
Every repeated run of \(R\) epochs over \(U\) unique tokens, \(D = RU\) in total, is compared with two one-epoch runs of the same model.
Baseline: one epoch on the same \(U\) tokens. \(U_{\mathrm{eq}}\) is the fresh budget that reaches the same loss, so \(\eta\) is the fresh-token equivalent of each repeated token: 1 if repeats were as good as fresh data, 0 if worthless.
Baseline: one epoch on \(D\) fresh tokens at the same compute and schedule. \(\Delta L\) is the extra held-out loss, in bits per byte, paid for repeating instead of collecting new data.
Fitted laws
-
Cost of repetition §4.2
\[ \Delta L = 0.0203\, z^{1.07}, \qquad z = \frac{(R-1)\,N_{\mathrm{total}}}{U} \tag{3} \]Excess loss over fresh data grows with \(z\), the extra epochs per unique token per parameter. Fitted to the 113 of 140 repeated runs above the 0.002 bpb resolution, \(R^2 = 0.96\); no further size term is needed once embeddings are counted.
-
Critical epochs §4.3
\[ R_c \approx 2.7 \left(\frac{D}{N}\right)^{0.24} \tag{4} \]The epoch count at which a repeated token is worth half a fresh one \((\eta = \tfrac12)\). Fits 35 measured crossings to 8% RMS: 4.0–5.3 epochs up to 20 tokens per parameter, 6.9–8.3 at 100, 10.0–11.8 at 300.
-
Allocation at fixed \(U\) §4.4
\[ N_{\mathrm{opt}} = \sqrt{\frac{C}{6\tau}}, \qquad R_{\mathrm{opt}} = \frac{1}{U}\sqrt{\frac{\tau C}{6}} \tag{5} \]With unique data fixed, compute goes to model size and epochs in roughly equal shares (fitted \(N \propto C^{0.52\text{–}0.53}\), \(R \propto C^{0.47\text{–}0.48}\)), holding \(\tau = D/N\) at 16–24 until loss stops improving near \(R_c\).
-
Best epochs at fixed \(U\) §5.1
\[ R^{\star} \approx 4.5 \left(\frac{U}{N}\right)^{0.51} \]Loss-minimizing epoch count with the corpus fixed, fitted to the five sizes from 127M to 2B whose loss turns upward within 32 epochs (\(R^2_{\log} = 0.94\), App. G). Because it depends on \(U/N\), a fixed corpus moves the optimum to fewer epochs as models grow.
Range of validity. Eqs. 3 and 4 are fitted on CommonCrawl with GPT-2 byte-level BPE, 23M–361M non-embedding parameters, 5–300 training tokens per parameter and 2–16 epochs, without weight decay or dropout. The constants may shift under other data or regularization.
Calculator
Price a planned run with Eqs. 3 and 4. Pick a model from the paper’s grid or enter your own parameter counts.
Evidence
The cost of repetition across the grid
Excess loss over fresh data at equal compute, in \(10^{-3}\) bits per byte (Figure 12a). Switch to Eq. 3 to see how much of the grid one variable explains.
A second epoch is nearly free; the fourth is not
| Model | Two epochs | Four epochs |
|---|---|---|
| 59M | 0.95 [0.76, 1.05] | 0.81 [0.75, 0.83] |
| 252M | 0.93 [0.75, 1.03] | 0.72 [0.64, 0.75] |
| 511M | 1.01 [0.82, 1.12] | 0.72 [0.65, 0.75] |
| 1B | 0.94 [0.76, 1.05] | 0.59 [0.51, 0.63] |
With the corpus fixed, larger models peak sooner
Scale the data with the model, and the best epoch count holds
| Model | CommonCrawl | GitHub |
|---|---|---|
| 23M | 3.72 | 3.06 |
| 45M | 3.79 | 2.84 |
| 78M | 4.03 | 2.75 |
| 185M | 4.22 | 2.68 |
| 361M | 4.30 | 2.66 |
| 0.8B | 4.42 | 2.61 |
| 1.2B | 4.22 | 2.51 |
Beyond counts
At identical token counts, how the repeats are laid out still changes loss.
- Order
- up to +0.46 bpb
- Largest rise from replaying each shard R times in a row instead of once per epoch, for 59M–511M models on 1B CommonCrawl tokens. Shuffling within epochs changes loss by less than 0.01.
- Allocation
- +0.0039 bpb
- Concentrating repeats on fewer samples: a third of the data seen four times versus half seen twice, at 45M with the same budget and the same number of extra exposures.
- Source
- r = −0.89
- Correlation between a source’s token entropy and its loss rise at 16 epochs. GitHub code rises 0.16 bpb above its minimum; Books does not rise.
- Re-tokenization
- −0.154 bpb
- BPE-dropout (p = 0.05) on repeats at 16 epochs. It costs 0.021 bpb at one epoch and 0.010 at four, so it pays only under heavy repetition.
Partial repetition follows a different law. Drawing a fraction \(f\) of the budget from a repeated subset costs \(\Delta L \propto f^{1.63}\, z_{\mathrm{sub}}^{0.53}\) \((R^2 = 0.89)\); Eq. 3 scaled by \(f\) overstates it about fourfold, so it serves as a conservative bound.
Citation
@misc{chai2026repeated,
title = {What Is a Repeated Token Worth? The Scaling Geometry of
Multi-Epoch Pretraining},
author = {Yekun Chai and Haoyi Xiong},
year = {2026},
eprint = {2610.05591},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.05591}
}