Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
Researchers introduced Terminal-Bench-LILT, a benchmark of 300 authentic coding tasks in ten languages (Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese) designed to evaluate coding …