Pretraining · October 2026

What is a repeated token worth?

The scaling geometry of multi-epoch pretraining

Yekun Chai1,2, Haoyi Xiong3

1FloatAI2ETH Zurich3Independent Researcher

arXiv preprint 2610.05591

14M–2Bparameters across all sweeps
175runs in the budget × epochs grid, 140 of them repeated
up to 32epochs over a fixed corpus, 64 for the smallest models
7SlimPajama sources, from code to books

a A second epoch is nearly free

Value of a repeated token against tokens per parameter, for 2, 4, 8 and 16 epochs. A second epoch stays near 1 at every budget; more epochs are worth less, most at small budgets.00.515102050100300tokens per parameter D/Nvalue of a repeat ηhalf a fresh token2×4×8×16×

b Half value after four to five epochs

Critical epoch count, where a repeated token is worth half a fresh one, against tokens per parameter. Points for five model sizes from 23M to 361M overlap; the fitted law rises from about 4 to about 11.468125102050100300tokens per parameter D/Nepochs to half value Rcworth more than halfworth less than halfRc ≈ 2.7 (D/N)0.24

c Scale epochs with data, not size

Loss-minimizing epoch count against model size when unique data grow with the model, U = 2.5N. CommonCrawl stays near four epochs from 23M to 1.2B; GitHub is lower and drifts down.234623M78M361M1.2Bmodel size, non-embeddingbest epoch count R*CommonCrawl · γ = +0.04GitHub · γ = −0.04
(a) Fresh-token value of a repeated token, mean over five sizes (Fig. 3). (b) Epochs at which it falls to half, measured for five sizes (points) with Eq. 4 (Fig. 4). (c) Loss-minimizing epochs when unique data grow with the model, U = 2.5N (Fig. 7a). Lines are log-linear fits in (a) and power laws in (b) and (c).

Takeaways

  • A second epoch is nearly free.Measured against one epoch on the same data, a repeated token in a second epoch is worth 0.87–1.02 fresh tokens at every budget. Each further epoch is worth less, and the drop is steepest at small budgets.
  • Larger budgets per parameter tolerate more repetition.Repeats fall to half the value of fresh tokens after about four to five epochs up to 20 tokens per parameter, and about ten at 300.
  • Transfer epoch counts with data, not size.With unique data fixed, larger models should take fewer epochs; when unique data grow with the model, the best epoch count barely changes. To reuse an epoch count on a larger model, scale the unique data with it.
  • Grow model and epochs together, up to the critical epoch count.With unique data fixed, extra compute goes to both in roughly equal shares until loss stops improving. Ignoring the cost of repetition loses at most 0.002 bpb up to that point.
  • Spread and shuffle repeats.Space repeats across epochs and samples, repeat lower-entropy sources less, and re-tokenize only under heavy repetition. Compare such choices at separately annealed horizons; intermediate checkpoints can misrank them.

Overview

As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters.

We answer them by pricing a repeated token against two references at the same model size. Against fresh data at equal compute, the cost of repetition follows a single variable: the number of extra epochs divided by the unique tokens per parameter. Against the same data seen once, a second epoch is worth nearly as much as a fresh one, and repeated tokens fall to half the value of fresh ones after a critical epoch count that grows with the training budget per parameter but hardly with model size.

The same variable explains size trends that appear to conflict in earlier work. Larger models tolerate fewer epochs when the corpus is fixed, but not when unique data grow with the model. And counts alone do not determine loss: the order, allocation and source of the repeats all change the result.

Two prices

Every repeated run of \(R\) epochs over \(U\) unique tokens, \(D = RU\) in total, is compared with two one-epoch runs of the same model.

Value · data-matched
\[ \eta = \frac{U_{\mathrm{eq}} - U}{D - U} \]

Baseline: one epoch on the same \(U\) tokens. \(U_{\mathrm{eq}}\) is the fresh budget that reaches the same loss, so \(\eta\) is the fresh-token equivalent of each repeated token: 1 if repeats were as good as fresh data, 0 if worthless.

Cost · compute-matched
\[ \Delta L = L(N, U, R) - L(N, D, 1) \]

Baseline: one epoch on \(D\) fresh tokens at the same compute and schedule. \(\Delta L\) is the extra held-out loss, in bits per byte, paid for repeating instead of collecting new data.

Fitted laws

  1. Cost of repetition §4.2

    \[ \Delta L = 0.0203\, z^{1.07}, \qquad z = \frac{(R-1)\,N_{\mathrm{total}}}{U} \tag{3} \]

    Excess loss over fresh data grows with \(z\), the extra epochs per unique token per parameter. Fitted to the 113 of 140 repeated runs above the 0.002 bpb resolution, \(R^2 = 0.96\); no further size term is needed once embeddings are counted.

  2. Critical epochs §4.3

    \[ R_c \approx 2.7 \left(\frac{D}{N}\right)^{0.24} \tag{4} \]

    The epoch count at which a repeated token is worth half a fresh one \((\eta = \tfrac12)\). Fits 35 measured crossings to 8% RMS: 4.0–5.3 epochs up to 20 tokens per parameter, 6.9–8.3 at 100, 10.0–11.8 at 300.

  3. Allocation at fixed \(U\) §4.4

    \[ N_{\mathrm{opt}} = \sqrt{\frac{C}{6\tau}}, \qquad R_{\mathrm{opt}} = \frac{1}{U}\sqrt{\frac{\tau C}{6}} \tag{5} \]

    With unique data fixed, compute goes to model size and epochs in roughly equal shares (fitted \(N \propto C^{0.52\text{–}0.53}\), \(R \propto C^{0.47\text{–}0.48}\)), holding \(\tau = D/N\) at 16–24 until loss stops improving near \(R_c\).

  4. Best epochs at fixed \(U\) §5.1

    \[ R^{\star} \approx 4.5 \left(\frac{U}{N}\right)^{0.51} \]

    Loss-minimizing epoch count with the corpus fixed, fitted to the five sizes from 127M to 2B whose loss turns upward within 32 epochs (\(R^2_{\log} = 0.94\), App. G). Because it depends on \(U/N\), a fixed corpus moves the optimum to fewer epochs as models grow.

Range of validity. Eqs. 3 and 4 are fitted on CommonCrawl with GPT-2 byte-level BPE, 23M–361M non-embedding parameters, 5–300 training tokens per parameter and 2–16 epochs, without weight decay or dropout. The constants may shift under other data or regularization.

Calculator

Price a planned run with Eqs. 3 and 4. Pick a model from the paper’s grid or enter your own parameter counts.

Model

Total parameters count tied embeddings once, as in the paper. For its GPT-2 vocabulary that adds 50,257 × width.

Tokens per parameter D/N
Load z = (R−1)Ntotal/U
Excess loss vs. fresh data · Eq. 3
Critical epochs Rc · Eq. 4

    Evidence

    The cost of repetition across the grid

    Excess loss over fresh data at equal compute, in \(10^{-3}\) bits per byte (Figure 12a). Switch to Eq. 3 to see how much of the grid one variable explains.

    Model
    24,000

    A second epoch is nearly free; the fourth is not

    Value of a repeated token η at U = 1B CommonCrawl tokens (Table 10). Brackets give the range under alternative exponents. Across the main grid, a second epoch is worth 0.87–1.02 fresh tokens at every budget.
    ModelTwo epochsFour epochs
    59M0.95 [0.76, 1.05]0.81 [0.75, 0.83]
    252M0.93 [0.75, 1.03]0.72 [0.64, 0.75]
    511M1.01 [0.82, 1.12]0.72 [0.65, 0.75]
    1B0.94 [0.76, 1.05]0.59 [0.51, 0.63]

    With the corpus fixed, larger models peak sooner

    Loss change from repeating one billion unique tokens, relative to one epoch, for eight sizes from 14M to 2B. Small models keep improving up to 16 or 32 epochs; the 1B and 2B models are best near four epochs and lose ground quickly after.−100−75−50−250+25+5012481632epochs over the same 1B unique tokensloss vs. one epoch (10⁻³ bpb)1B (+191)2B (+496)511M252M14M127M31M59M
    The same 1B unique CommonCrawl tokens, relative to each model’s own single epoch (lower is better; Table 9). Curves pass through the measured points (monotone cubic in log R); rings mark the loss minima from parabola fits, at 15.3, 7.9, 5.8, 4.1 and 3.6 epochs from 127M to 2B.

    Scale the data with the model, and the best epoch count holds

    Loss-minimizing epochs when unique data grow with the model, U = 2.5N (Table 8; quadratic in log R through all points of a size). On CommonCrawl the optimum stays near four epochs from 23M to 1.2B; on lower-entropy GitHub code it is lower and drifts down.
    ModelCommonCrawlGitHub
    23M3.723.06
    45M3.792.84
    78M4.032.75
    185M4.222.68
    361M4.302.66
    0.8B4.422.61
    1.2B4.222.51

    Beyond counts

    At identical token counts, how the repeats are laid out still changes loss.

    Order
    up to +0.46 bpb
    Largest rise from replaying each shard R times in a row instead of once per epoch, for 59M–511M models on 1B CommonCrawl tokens. Shuffling within epochs changes loss by less than 0.01.
    Allocation
    +0.0039 bpb
    Concentrating repeats on fewer samples: a third of the data seen four times versus half seen twice, at 45M with the same budget and the same number of extra exposures.
    Source
    r = −0.89
    Correlation between a source’s token entropy and its loss rise at 16 epochs. GitHub code rises 0.16 bpb above its minimum; Books does not rise.
    Re-tokenization
    −0.154 bpb
    BPE-dropout (p = 0.05) on repeats at 16 epochs. It costs 0.021 bpb at one epoch and 0.010 at four, so it pays only under heavy repetition.

    Partial repetition follows a different law. Drawing a fraction \(f\) of the budget from a repeated subset costs \(\Delta L \propto f^{1.63}\, z_{\mathrm{sub}}^{0.53}\) \((R^2 = 0.89)\); Eq. 3 scaled by \(f\) overstates it about fourfold, so it serves as a conservative bound.

    Citation

    @misc{chai2026repeated,
      title         = {What Is a Repeated Token Worth? The Scaling Geometry of
                       Multi-Epoch Pretraining},
      author        = {Yekun Chai and Haoyi Xiong},
      year          = {2026},
      eprint        = {2610.05591},
      archivePrefix = {arXiv},
      primaryClass  = {cs.CL},
      url           = {https://arxiv.org/abs/2610.05591}
    }