Agents · October 2026

Finding the move is not winning the game

XiangqiBench for closed-loop evaluation of LLM agents

Yekun Chai1, Qiwei Peng1,2, Haoyi Xiong3

1FloatAI2University of Copenhagen3Independent Researcher

arXiv preprint 2610.02425

119tactical endgames with forced mates in 3–11 plies
12frontier LLMs, each playing Red against an engine
8,568multi-turn trajectories, three trials per position
2observation protocols, sighted and restricted
Illustrative position. Red (shown in blue) plays R9+3: the rook reaches the back rank with check, and the horse covers the king’s only other flight square. Checkmate.

Overview

Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. XiangqiBench measures that gap in xiangqi, or Chinese chess.

Starting from 119 tactical endgames with forced mates, an LLM agent must deliver checkmate against an engine defender. An interactive REPL separates real moves, state queries and forward simulation. We record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols.

Three signals that look like competence each overstate closed-loop success. Agent evaluations should score the outcome of the whole game, and report reliability alongside coverage.

Findings

Conversion gap
26.1% → 13.9%
Models play the reference first move in 26.1% of sighted trials; only 13.9% of those trials end in a win.
Consistency gap
38.7% vs 5.9%
The leading model reaches 38.7% pass@3 but only 5.9% pass^3, winning all three trials on 7 of the 46 positions it ever wins.
Simulation gap
32.3%
of accepted simulation calls stop on an illegal move. Self-authored rollouts check legality but cannot anticipate the opponent.
Opponent mismatch
49.3%
of comparable defender replies depart from the line the agent simulated for them.

Leaderboard

Success rate (%) over 119 positions with three trials each. pass@3 counts positions won in at least one trial; pass^3, positions won in all three. The gap between the two bars is the consistency gap.

Success rate (%) over 119 positions, three trials each, ranked by sighted pass@1. S: sighted; R: restricted. pass@3 counts positions won in at least one of three trials, pass^3 positions won in all three.
Model pass@1 pass@3 pass^3
SR SR SR
Gemini 3.1 Pro22.114.038.732.85.91.7
GPT-5.59.03.417.66.72.50.8
Qwen 3.7 Max6.22.813.45.90.80.8
Claude Opus 4.75.30.810.91.70.80.0
DeepSeek V4 Pro2.80.05.00.01.70.0
Seed 2.0 Pro0.80.62.51.70.00.0
Kimi K2.50.80.31.70.80.00.0
Mistral Large 30.80.00.80.00.80.0
MiniMax M2.50.80.00.80.00.80.0
DeepSeek R10.60.01.70.00.00.0
GPT-OSS 120B0.30.00.80.00.00.0
DeepSeek V3.20.30.00.80.00.00.0

Bootstrap intervals and k = 2 are in the paper. To add a model, run the benchmark with the standard settings and open an issue with the scored records.

Per position

Each column is one position, grouped by mate distance; each cell counts the wins of one model in three trials. Wins concentrate on a minority of positions: 56 of the 119 are never won by any model under sighted observation, and 73 under restricted observation.

Loading…

Replays

Four games from the paper, ply by ply: what the agent wrote, simulated and played on each turn, and how the defender answered. Use the arrow keys to step.

Loading…

Setup

Positions
119 tactical endgames with forced mates in 3–11 plies, verified by engine or checks-only search: 116 from Shi Qing Ya Qu (1570), 3 from Jianghu
Defender
Pikafish, depth 18, one thread, 256 MB hash
Interface
REPL with real moves, state queries and forward simulation
Protocols
Sighted (board, FEN and legal moves each ply) and restricted (move diffs only)
Scoring
pass@k and pass^k with 95% case-bootstrap intervals

Get started

pip install xiangqibench
# build the Pikafish defender
git clone https://github.com/official-pikafish/Pikafish && make -C Pikafish/src -j build
export PIKAFISH_PATH=$PWD/Pikafish/src/pikafish
xiangqibench doctor

# run five positions and score them
xiangqibench run --model gpt-5.5 --mode sighted --limit 5
xiangqibench score runs/

Citation

@misc{chai2026xiangqibench,
  title         = {Finding the Move Is Not Winning the Game:
                   {XiangqiBench} for Closed-Loop Evaluation of {LLM} Agents},
  author        = {Yekun Chai and Qiwei Peng and Haoyi Xiong},
  year          = {2026},
  eprint        = {2610.02425},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2610.02425}
}