What is a repeated token worth?
Public text grows more slowly than compute. We price a repeated token against fresh data and find the epoch count past which it is worth half as much, across models from 14M to 2B parameters.
FloatAI Research
An open research network studying how large language models learn, read and act, and how to evaluate them. Our benchmarks and datasets are public.
Research
Public text grows more slowly than compute. We price a repeated token against fresh data and find the epoch count past which it is worth half as much, across models from 14M to 2B parameters.
LLM agents must deliver checkmate in xiangqi endgames against an engine. Across 12 frontier models, the right first move, pass@k and self-simulated lines each overstate how often an agent actually wins.
Models see subwords, not characters. We probe what LLMs know about token structure, and how much accuracy they lose under character- and token-level typos.
Most code benchmarks are English-in, code-out. We pose the same 80 problems in 23 natural languages and 12 programming languages, so any drop reflects the language, not the task.
Network
FloatAI brings together researchers in Beijing, Copenhagen, Zurich, London, Edinburgh and beyond. The work lives in public repositories, open to anyone who wants to contribute.
Run our harnesses on your models and report what you find. Negative results are welcome: open an issue with your config and logs.
Open an issue, fix a bug, add a model backend or a new language. Every merged contribution is credited.
Have a focused question about LLMs and want collaborators? Start a thread in Discussions.