Multilingual · LREC-COLING 2024

HumanEval-XL

A multilingual code generation benchmark for cross-lingual natural language generalization

Qiwei Peng1*, Yekun Chai2*, Xuhong Li2

1University of Copenhagen2Baidu*Equal contribution

Proceedings of LREC-COLING 2024, pages 8383–8394

22,080prompts, each with executable tests
23natural languages in 11 families
12programming languages, from Python to Scala
80problems, parallel across every language pair
Programming language
Natural language
def digits(n):
    """Given a positive integer n, return the product of the odd digits.
    Return 0 if all digits are even.
    For example:
    digits(1)  == 1
    digits(4)  == 0
    digits(235) == 15
    """
Problem 58 · English · Python
One of the 80 problems, as the benchmark poses it. The task and tests stay fixed; only the natural language of the description and the programming language of the stub change.

Overview

Code generation benchmarks have mostly translated English prompts into many programming languages, or covered very few natural languages. That leaves the most common real case untested: a developer who describes the task in their own language and wants code in the language of their project.

HumanEval-XL connects 23 natural languages with 12 programming languages through the same 80 problems. Because every prompt is parallel, a drop in pass@1 from English to Finnish, or from Python to Scala, measures the language itself rather than a change of task.

Benchmark

How it compares

Execution-based code generation benchmarks (Table 2). Parallel: the same problems in every natural language.
BenchmarkSamplesTests per sampleSourcePLsNLsParallel
HumanEval1647.7Hand-written11—
MBPP9743.0Hand-written11—
APPS10,00013.2Competitions11—
DSP1,1192.1GitHub notebooks11—
MTPB1155.0Hand-written11—
DS-10001,0001.6StackOverflow11—
Multilingual HumanEval1,9357.8Hand-written121—
ODEX9451.8StackOverflow14—
HumanEval-XL22,0808.3Hand-written1223Yes

How it was built

  1. ExtractTake the natural-language part of each English prompt from Multilingual HumanEval and wrap it for translation.
  2. TranslateGPT-4 translates it into every target language, then back-translates each version into English.
  3. CheckKeep a translation only if the BERTScore between back-translation and original exceeds 0.95; retry up to three times, then discard.
  4. ReviewHeuristic checks, plus manual review and correction of randomly sampled examples.

Languages

Germanic
English5, German5, Dutch4, Afrikaans3
Romance
Spanish5, French5, Portuguese4, Italian4
Slavic
Russian4, Bulgarian3
Uralic
Finnish4, Hungarian4, Estonian3
Afro-Asiatic
Arabic5, Hebrew3
Austronesian
Indonesian3, Malay3, Tagalog3
Sino-Tibetan
Chinese5
Austro-Asiatic
Vietnamese4
Greek
Greek3
Iranian
Persian4
Turkic
Turkish3

Superscripts give the resource class of Joshi et al. (2020), from 5 (most resourced) to 3. Programming languages: Python, Java, JavaScript, TypeScript, C#, Go, Kotlin, PHP, Ruby, Perl, Swift and Scala.

Leaderboard

pass@1 for nine models on every language pair, from the paper’s appendix tables. Pick a programming language to see all 23 natural languages.

Loading results…
085

Findings

Strongest model
276 / 276
Language pairs where GPT-4 has the highest pass@1. Its mean lead over the next model ranges from 15 points in Python to 58 in PHP.
Language matters
10–41.25
GPT-4 pass@1 on the same 80 Scala problems, from Arabic to Turkish prompts. In Python it spans 71.25–83.75, from Greek to Hungarian.
Code pretraining
9 of 12
Programming languages where CodeGen2 16B outscores the much larger GPT-3.5 on average; GPT-3.5 leads only in Python, Perl and Scala.
Hardest language
Scala
No CodeT5+ or CodeGen2 model solves a single Scala problem in any natural language. For GPT-4, Go is next hardest at 44.8 mean pass@1.

Scale helps within a family: CodeGen2 rises from 9.6 to 19.7 mean pass@1 in Python from 1B to 16B parameters. Models of the same family also rank languages alike, with Pearson correlations of 0.80 (CodeGen2 3.7B and 16B) and 0.87 (GPT-3.5 and GPT-4) across resource levels, while models from different families do not.

Setup

Models
CodeT5+ (220M, 770M, 2B; encoder-decoder), CodeGen2 (1B, 3.7B, 7B, 16B), GPT-3.5 and GPT-4
Metric
pass@1 by executing the tests; 80 problems per pair, so scores move in steps of 1.25
Decoding
Top-p 0.95, temperature 0.2; only the first generated function is kept
Source
Multilingual HumanEval (Athiwaratkun et al., 2023), translated with GPT-4 and filtered by BERTScore

Get started

# the dataset ships a loading script, which datasets 4.x no longer runs
pip install "datasets<4"
from datasets import load_dataset

# one config per programming language, one split per natural language
dataset = load_dataset("floatai/HumanEval-XL", "python", trust_remote_code=True)
print(dataset["Chinese"][0]["prompt"])

The same files are also in the GitHub repository as data/<language>/<Natural language>.jsonl, with the evaluation harness under mxeval/.

Citation

@inproceedings{peng-etal-2024-humaneval,
  title     = "{H}uman{E}val-{XL}: A Multilingual Code Generation Benchmark for
               Cross-lingual Natural Language Generalization",
  author    = "Peng, Qiwei and Chai, Yekun and Li, Xuhong",
  booktitle = "Proceedings of the 2024 Joint International Conference on Computational
               Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
  month     = may,
  year      = "2024",
  address   = "Torino, Italia",
  publisher = "ELRA and ICCL",
  url       = "https://aclanthology.org/2024.lrec-main.735/",
  pages     = "8383--8394"
}