{"slug": "terminal-bench-lilt-multilingual-agentic-coding-benchmark-grounded-in-language", "title": "Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture", "summary": "Researchers introduced Terminal-Bench-LILT, a benchmark of 300 authentic coding tasks in ten languages (Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese) designed to evaluate coding agents on non-English software development issues. Evaluation of six frontier models showed the strongest model achieved only a 63.1% pass rate, with many tasks unsolved by any model, indicating that multilingual coding competence is a distinct capability not reflected in general coding benchmarks.", "body_md": "arXiv:2608.28641v1 Announce Type: new\nAbstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese. Each task targets issues specific to non-English software development that have no direct English equivalent, e.g., internationalization, encoding, text normalization, and cultural conventions. All tasks are authored by native-speaker programmers and validated through a multi-stage quality control pipeline. Evaluation of six frontier models reveals that even the strongest model reaches only 63.1\\% pass rate, with many tasks unsolved by any model. Performance varies substantially by language and does not track general coding benchmark rankings, highlighting that multilingual coding competence is a distinct and underexplored capability axis. Sample tasks are available at https://github.com/lilt/terminal-bench-lilt", "url": "https://wpnews.pro/news/terminal-bench-lilt-multilingual-agentic-coding-benchmark-grounded-in-language", "canonical_source": "https://arxiv.org/abs/2608.28641", "published_at": "2026-09-01 04:00:00+00:00", "updated_at": "2026-09-01 04:25:00.483025+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-agents"], "entities": ["Terminal-Bench-LILT", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/terminal-bench-lilt-multilingual-agentic-coding-benchmark-grounded-in-language", "markdown": "https://wpnews.pro/news/terminal-bench-lilt-multilingual-agentic-coding-benchmark-grounded-in-language.md", "text": "https://wpnews.pro/news/terminal-bench-lilt-multilingual-agentic-coding-benchmark-grounded-in-language.txt", "jsonld": "https://wpnews.pro/news/terminal-bench-lilt-multilingual-agentic-coding-benchmark-grounded-in-language.jsonld"}}