An open research network studying how large language models learn, read and act.
Research overview · October 2026
Our mission
Then release the data and code, so that anyone can check the answer.
Who we are
Researchers across universities and labs, working in public on how language models learn, read and act.
Each project asks one question that an experiment can answer.
Outcomes rather than proxies, and reliability beside averages.
Data and evaluation code ship with every paper.
Research
Pretraining
Repeated tokens · Aurora-M
Agents
XiangqiBench
Tokenization
TKEval · Subword compositionality
Multilingual
HumanEval-XL · CodeMixBench · Debiasing
Efficiency
EvolKV
Interpretability
GiLOT
At a glance
11
papers at EMNLP, ICML and COLING
4
open benchmarks
22,080
parallel coding prompts
175
controlled pretraining runs
01 What is a repeated token worth?
Points: measured. Lines: log-linear fits (left), Rc ≈ 2.7 (D/N)0.24 (centre), power laws (right). Chai and Xiong, arXiv 2610.05591, Figs. 3, 4 and 7.
02 Finding the move is not winning the game
03 Tokenization falling short
04 HumanEval-XL
Evaluate your models and share what you find.
Fix a bug, or add a model or a language.
Bring a question and find collaborators.