Multilingual · LREC-COLING 2024
HumanEval-XL
A multilingual code generation benchmark for cross-lingual natural language generalization
1University of Copenhagen2Baidu*Equal contribution
Proceedings of LREC-COLING 2024, pages 8383–8394
def digits(n):
"""Given a positive integer n, return the product of the odd digits.
Return 0 if all digits are even.
For example:
digits(1) == 1
digits(4) == 0
digits(235) == 15
"""
Overview
Code generation benchmarks have mostly translated English prompts into many programming languages, or covered very few natural languages. That leaves the most common real case untested: a developer who describes the task in their own language and wants code in the language of their project.
HumanEval-XL connects 23 natural languages with 12 programming languages through the same 80 problems. Because every prompt is parallel, a drop in pass@1 from English to Finnish, or from Python to Scala, measures the language itself rather than a change of task.
Benchmark
How it compares
| Benchmark | Samples | Tests per sample | Source | PLs | NLs | Parallel |
|---|---|---|---|---|---|---|
| HumanEval | 164 | 7.7 | Hand-written | 1 | 1 | — |
| MBPP | 974 | 3.0 | Hand-written | 1 | 1 | — |
| APPS | 10,000 | 13.2 | Competitions | 1 | 1 | — |
| DSP | 1,119 | 2.1 | GitHub notebooks | 1 | 1 | — |
| MTPB | 115 | 5.0 | Hand-written | 1 | 1 | — |
| DS-1000 | 1,000 | 1.6 | StackOverflow | 1 | 1 | — |
| Multilingual HumanEval | 1,935 | 7.8 | Hand-written | 12 | 1 | — |
| ODEX | 945 | 1.8 | StackOverflow | 1 | 4 | — |
| HumanEval-XL | 22,080 | 8.3 | Hand-written | 12 | 23 | Yes |
How it was built
- ExtractTake the natural-language part of each English prompt from Multilingual HumanEval and wrap it for translation.
- TranslateGPT-4 translates it into every target language, then back-translates each version into English.
- CheckKeep a translation only if the BERTScore between back-translation and original exceeds 0.95; retry up to three times, then discard.
- ReviewHeuristic checks, plus manual review and correction of randomly sampled examples.
Languages
- Germanic
- English5, German5, Dutch4, Afrikaans3
- Romance
- Spanish5, French5, Portuguese4, Italian4
- Slavic
- Russian4, Bulgarian3
- Uralic
- Finnish4, Hungarian4, Estonian3
- Afro-Asiatic
- Arabic5, Hebrew3
- Austronesian
- Indonesian3, Malay3, Tagalog3
- Sino-Tibetan
- Chinese5
- Austro-Asiatic
- Vietnamese4
- Greek
- Greek3
- Iranian
- Persian4
- Turkic
- Turkish3
Superscripts give the resource class of Joshi et al. (2020), from 5 (most resourced) to 3. Programming languages: Python, Java, JavaScript, TypeScript, C#, Go, Kotlin, PHP, Ruby, Perl, Swift and Scala.
Leaderboard
pass@1 for nine models on every language pair, from the paper’s appendix tables. Pick a programming language to see all 23 natural languages.
Findings
- Strongest model
- 276 / 276
- Language pairs where GPT-4 has the highest pass@1. Its mean lead over the next model ranges from 15 points in Python to 58 in PHP.
- Language matters
- 10–41.25
- GPT-4 pass@1 on the same 80 Scala problems, from Arabic to Turkish prompts. In Python it spans 71.25–83.75, from Greek to Hungarian.
- Code pretraining
- 9 of 12
- Programming languages where CodeGen2 16B outscores the much larger GPT-3.5 on average; GPT-3.5 leads only in Python, Perl and Scala.
- Hardest language
- Scala
- No CodeT5+ or CodeGen2 model solves a single Scala problem in any natural language. For GPT-4, Go is next hardest at 44.8 mean pass@1.
Scale helps within a family: CodeGen2 rises from 9.6 to 19.7 mean pass@1 in Python from 1B to 16B parameters. Models of the same family also rank languages alike, with Pearson correlations of 0.80 (CodeGen2 3.7B and 16B) and 0.87 (GPT-3.5 and GPT-4) across resource levels, while models from different families do not.
Setup
- Models
- CodeT5+ (220M, 770M, 2B; encoder-decoder), CodeGen2 (1B, 3.7B, 7B, 16B), GPT-3.5 and GPT-4
- Metric
- pass@1 by executing the tests; 80 problems per pair, so scores move in steps of 1.25
- Decoding
- Top-p 0.95, temperature 0.2; only the first generated function is kept
- Source
- Multilingual HumanEval (Athiwaratkun et al., 2023), translated with GPT-4 and filtered by BERTScore
Get started
# the dataset ships a loading script, which datasets 4.x no longer runs
pip install "datasets<4"
from datasets import load_dataset
# one config per programming language, one split per natural language
dataset = load_dataset("floatai/HumanEval-XL", "python", trust_remote_code=True)
print(dataset["Chinese"][0]["prompt"])
The same files are also in the GitHub repository as data/<language>/<Natural language>.jsonl, with the evaluation harness under mxeval/.
Citation
@inproceedings{peng-etal-2024-humaneval,
title = "{H}uman{E}val-{XL}: A Multilingual Code Generation Benchmark for
Cross-lingual Natural Language Generalization",
author = "Peng, Qiwei and Chai, Yekun and Li, Xuhong",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational
Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.lrec-main.735/",
pages = "8383--8394"
}