FloatAI Research

Measure what models actually do.
Then share it openly.

An open research network studying how large language models learn, read and act, and how to evaluate them. Our benchmarks and datasets are public.

Research

Questions we keep coming back to

All repositories

PretrainingPreprint 2026

What is a repeated token worth?

Public text grows more slowly than compute. We price a repeated token against fresh data and find the epoch count past which it is worth half as much, across models from 14M to 2B parameters.

AgentsPreprint 2026

Can an agent finish the game, not just find the move?

LLM agents must deliver checkmate in xiangqi endgames against an engine. Across 12 frontier models, the right first move, pass@k and self-simulated lines each overstate how often an agent actually wins.

TokenizationFindings of EMNLP 2024

What do models know about the letters inside their tokens?

Models see subwords, not characters. We probe what LLMs know about token structure, and how much accuracy they lose under character- and token-level typos.

MultilingualLREC-COLING 2024

Does code generation transfer beyond English prompts?

Most code benchmarks are English-in, code-out. We pose the same 80 problems in 23 natural languages and 12 programming languages, so any drop reflects the language, not the task.

Publications

Selected papers

    Network

    An open network for LLM research

    FloatAI brings together researchers in Beijing, Copenhagen, Zurich, London, Edinburgh and beyond. The work lives in public repositories, open to anyone who wants to contribute.

    1. 01

      Use the benchmarks

      Run our harnesses on your models and report what you find. Negative results are welcome: open an issue with your config and logs.

    2. 02

      Contribute

      Open an issue, fix a bug, add a model backend or a new language. Every merged contribution is credited.

    3. 03

      Propose a project

      Have a focused question about LLMs and want collaborators? Start a thread in Discussions.