← floatai.github.io

FloatAI

An open research network studying how large language models learn, read and act.

Research overview · October 2026

Our mission

Measure what language models actually do.

Then release the data and code, so that anyone can check the answer.

Who we are

An open research network

Researchers across universities and labs, working in public on how language models learn, read and act.

Focused questions

Each project asks one question that an experiment can answer.

Careful measurement

Outcomes rather than proxies, and reliability beside averages.

Open release

Data and evaluation code ship with every paper.

Research

What we study

Pretraining

How should training reuse a finite corpus?

Repeated tokens · Aurora-M

Agents

Do agents finish what they start?

XiangqiBench

Tokenization

What do models know below the token?

TKEval · Subword compositionality

Multilingual

Does capability transfer across languages?

HumanEval-XL · CodeMixBench · Debiasing

Efficiency

How much KV cache does inference need?

EvolKV

Interpretability

Which inputs drive a generation?

GiLOT

At a glance

Since 2024

11

papers at EMNLP, ICML and COLING

4

open benchmarks

22,080

parallel coding prompts

175

controlled pretraining runs

Selected work

  1. What is a repeated token worth?Pretraining · 2026
  2. Finding the move is not winning the gameAgents · 2026
  3. Tokenization falling shortTokenization · EMNLP 2024
  4. HumanEval-XLMultilingual · LREC-COLING 2024

01 What is a repeated token worth?

A repeated token is worth half a fresh one after four to five epochs

00.515102050100300tokens per parameter D/Nvalue of a repeat, ηhalf a fresh token2×4×8×16×
A second epoch is nearly free.η = 0.87–1.02 at every budget, then falls with each further epoch.
468125102050100300tokens per parameter D/Nepochs to half value Rcworth more than halfworth less than halfRc ≈ 2.7 (D/N)0.24
Half value at four to five epochs.Rising slowly with budget, and nearly the same for every model size.
234623M78M361M1.2Bmodel size (non-embedding)best epoch count R*CommonCrawl · γ = +0.04GitHub · γ = -0.04
Scale epochs with data, not size.With unique data at 2.5N, the best epoch count stays near four from 23M to 1.2B.

Points: measured. Lines: log-linear fits (left), Rc ≈ 2.7 (D/N)0.24 (centre), power laws (right). Chai and Xiong, arXiv 2610.05591, Figs. 3, 4 and 7.

02 Finding the move is not winning the game

Agents that find the right move rarely finish the game

pass@3 · won at least oncepass^3 · won every trial
Gemini 3.1 Pro38.7 / 5.9
GPT-5.517.6 / 2.5
Qwen 3.7 Max13.4 / 0.8
Claude Opus 4.710.9 / 0.8
DeepSeek V4 Pro5.0 / 1.7
Seed 2.0 Pro2.5 / 0.0
Kimi K2.51.7 / 0.0
DeepSeek R11.7 / 0.0
Checkmate rate (%) on 119 xiangqi endgames against an engine, three trials each. Chai, Peng and Xiong, arXiv 2610.02425.
  1. The right move is not a win.26.1% open correctly; 13.9% of those mate.
  2. Coverage is not reliability.Best model: 38.7% at least once, 5.9% every time.
  3. Self-simulation misleads.49.3% of defender replies leave the agent’s plan.

03 Tokenization falling short

Models know little about the characters inside their tokens

Two angle diagrams. The embedding of “assignment” and the sum of the embeddings of “assign” and “ment” are 78.16 degrees apart, cosine 0.21. “import” and “im” plus “port” are 82.47 degrees apart, cosine 0.13. “assignment” “assign” + “ment” 78.16° cosine 0.21 “import” “im” + “port” 82.47° cosine 0.13
A word’s embedding is nearly orthogonal to the sum of its subword embeddings. Chai, Fang, Peng and Li, Findings of EMNLP 2024.
  1. Scale helps, but does not close the gap.Counting from the end peaks at 52%; subsequences near 4%.
  2. Typos bias every model.Character noise hurts more than token-level changes.
  3. Moderate BPE-dropout helps.p = 0.2 often beats the baseline.

04 HumanEval-XL

With the task fixed, the prompt’s language still moves the score

EN RU ZH DE ES FR IT PT EL HU NL FI ID TR AR VI BG FA MS HE ET TL AF Python Java JavaScript TypeScript C# Go Kotlin PHP Ruby Perl Swift Scala
GPT-4 pass@11083.75The same 80 problems in 12 programming × 23 natural languages. Peng, Chai and Li, LREC-COLING 2024.
  1. Language matters.Scala: 10 to 41, depending on the prompt’s language.
  2. One model leads everywhere.GPT-4 is best in all 276 language pairs.
  3. Code pretraining beats size.CodeGen2 16B beats GPT-3.5 in 9 of 12.

Get involved

Run the benchmarks

Evaluate your models and share what you find.

Contribute

Fix a bug, or add a model or a language.

Propose a project

Bring a question and find collaborators.

Measure what models do.
Share it openly.