Agents · October 2026
Finding the move is not winning the game
XiangqiBench for closed-loop evaluation of LLM agents
1FloatAI2University of Copenhagen3Independent Researcher
arXiv preprint 2610.02425
Overview
Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. XiangqiBench measures that gap in xiangqi, or Chinese chess.
Starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender. An interactive REPL separates real moves, state queries and forward simulation. We record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols.
Three signals that look like competence each overstate closed-loop success. Agent evaluations should score the outcome of the whole game, and report reliability alongside coverage.
Findings
- Conversion gap
- 26.1% → 13.9%
- Models play the reference first move in 26.1% of sighted trials; only 13.9% of those trials end in a win.
- Consistency gap
- 38.7% vs 5.9%
- The leading model reaches 38.7% pass@3 but only 5.9% pass^3, winning all three trials on 7 of the 46 positions it ever wins.
- Simulation gap
- 32.3%
- of accepted simulation calls stop on an illegal move. Self-authored rollouts check legality but cannot anticipate the opponent.
- Opponent mismatch
- 49.3%
- of comparable defender replies depart from the line the agent simulated for them.
Leaderboard
Success rate (%) over 119 positions with three trials each. pass@3 counts positions won in at least one trial; pass^3, positions won in all three. The gap between the two bars is the consistency gap.
| Model | pass@1 | pass@3 | pass^3 | |||
|---|---|---|---|---|---|---|
| S | R | S | R | S | R | |
| Gemini 3.1 Pro | 22.1 | 14.0 | 38.7 | 32.8 | 5.9 | 1.7 |
| GPT-5.5 | 9.0 | 3.4 | 17.6 | 6.7 | 2.5 | 0.8 |
| Qwen 3.7 Max | 6.2 | 2.8 | 13.4 | 5.9 | 0.8 | 0.8 |
| Claude Opus 4.7 | 5.3 | 0.8 | 10.9 | 1.7 | 0.8 | 0.0 |
| DeepSeek V4 Pro | 2.8 | 0.0 | 5.0 | 0.0 | 1.7 | 0.0 |
| Seed 2.0 Pro | 0.8 | 0.6 | 2.5 | 1.7 | 0.0 | 0.0 |
| Kimi K2.5 | 0.8 | 0.3 | 1.7 | 0.8 | 0.0 | 0.0 |
| Mistral Large 3 | 0.8 | 0.0 | 0.8 | 0.0 | 0.8 | 0.0 |
| MiniMax M2.5 | 0.8 | 0.0 | 0.8 | 0.0 | 0.8 | 0.0 |
| DeepSeek R1 | 0.6 | 0.0 | 1.7 | 0.0 | 0.0 | 0.0 |
| GPT-OSS 120B | 0.3 | 0.0 | 0.8 | 0.0 | 0.0 | 0.0 |
| DeepSeek V3.2 | 0.3 | 0.0 | 0.8 | 0.0 | 0.0 | 0.0 |
Bootstrap intervals and k = 2 are in the paper. To add a model, run the benchmark with the standard settings and open an issue with the scored records.
Per position
Each column is one position, grouped by mate distance; each cell counts the wins of one model in three trials. Wins concentrate on a minority of positions: 56 of the 119 are never won by any model under sighted observation, and 73 under restricted observation.
Loading…
Replays
Four games from the paper, ply by ply: what the agent wrote, simulated and played on each turn, and how the defender answered. Use the arrow keys to step.
Loading…
Setup
- Positions
- 119 tactical endgames with forced mates in 3–11 plies, verified by engine or checks-only search: 116 from Shi Qing Ya Qu (1570), 3 from Jianghu
- Defender
- Pikafish, depth 18, one thread, 256 MB hash
- Interface
- REPL with real moves, state queries and forward simulation
- Protocols
- Sighted (board, FEN and legal moves each ply) and restricted (move diffs only)
- Scoring
- pass@k and pass^k with 95% case-bootstrap intervals
Get started
pip install xiangqibench
# build the Pikafish defender
git clone https://github.com/official-pikafish/Pikafish && make -C Pikafish/src -j build
export PIKAFISH_PATH=$PWD/Pikafish/src/pikafish
xiangqibench doctor
# run five positions and score them
xiangqibench run --model gpt-5.5 --mode sighted --limit 5
xiangqibench score runs/
Citation
@misc{chai2026xiangqibench,
title = {Finding the Move Is Not Winning the Game:
{XiangqiBench} for Closed-Loop Evaluation of {LLM} Agents},
author = {Yekun Chai and Qiwei Peng and Haoyi Xiong},
year = {2026},
eprint = {2610.02425},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.02425}
}