cd /news/artificial-intelligence/terminal-bench-lilt-multilingual-age… · home topics artificial-intelligence article
[ARTICLE · art-117325] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

Researchers introduced Terminal-Bench-LILT, a benchmark of 300 authentic coding tasks in ten languages (Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese) designed to evaluate coding agents on non-English software development issues. Evaluation of six frontier models showed the strongest model achieved only a 63.1% pass rate, with many tasks unsolved by any model, indicating that multilingual coding competence is a distinct capability not reflected in general coding benchmarks.

read1 min views4 publishedSep 1, 2026

arXiv:2608.28641v1 Announce Type: new Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese. Each task targets issues specific to non-English software development that have no direct English equivalent, e.g., internationalization, encoding, text normalization, and cultural conventions. All tasks are authored by native-speaker programmers and validated through a multi-stage quality control pipeline. Evaluation of six frontier models reveals that even the strongest model reaches only 63.1% pass rate, with many tasks unsolved by any model. Performance varies substantially by language and does not track general coding benchmark rankings, highlighting that multilingual coding competence is a distinct and underexplored capability axis. Sample tasks are available at https://github.com/lilt/terminal-bench-lilt

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @terminal-bench-lilt 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/terminal-bench-lilt-…] indexed:0 read:1min 2026-09-01 ·