This is a submission for the Kaggle Benchmarking Challenge I build AI tools for developers and learners in Kenya, so I kept asking one question: if you brief a coding agent in a language it barely saw in training, does it still do the job?
I tested six languages: English, Chinese, French, Spanish, Swahili, and Maragoli (Lulogooli), a Luhya language spoken in western Kenya. I wrote the Maragoli prompts myself.
Task 1: Coding. 5 small Python problems, each described in all 6 languages (30 prompts per model). Function names, signatures and unit tests stay identical, so only the instruction language changes. Score = unit-test pass rate.
Task 2: Planning. A good agent plans before it codes, and planning means asking the user clarifying questions. I gave 3 deliberately vague requests ("build me a small app to track my expenses", "write a script that cleans up my files", "make a website for my shop") in all 6 languages, each ending with: make a short plan, ask clarifying questions, and do not write code yet. I check whether the model asks a question, follows "no code yet", and replies in the language it was spoken to.
Task 1 is a control: the English code scaffolding makes the language barrier easy to cross. Task 2 removes that scaffolding.
16 models ran Task 1, and 7 also completed Task 2: Claude Haiku 5.5 and Sonnet 5.5, GPT-6.1 Sol, Gemini 3.1 Pro Preview , a Gemini 3.8 Flash model, GLM-5 and gpt-oss-120b. I picked a spread of providers and sizes, including Chinese-lab models (GLM, Qwen, DeepSeek). Opus 5.5 and GPT-6 Astra scored 100% on Task 1 but could not run Task 2 (free quota), so the "0" on their leaderboard Overall reflects a failed run, not their ability. Grok was not available on the platform.
Short answer: for code, mostly yes. For a conversation in Maragoli, no.
Task 1. 11 of 16 models scored 100%. The rest: gpt-oss-120b, GLM-5 and Gemma 4 31B at 96.7%, gpt-oss-20b at 93.3%, and Qwen3 Next 80B Thinking at 86.7%. The language barrier barely mattered, because the code scaffolding was in English.
Task 2, Maragoli prompts (7 models x 3 prompts = 21 replies):
| Model | Replied in Maragoli* | English or Swahili | Another language | Followed "no code yet" |
|---|---|---|---|---|
| Claude Haiku 5.5 | 0/3 | 3 | 0 | 2/3 |
| Claude Sonnet 5.5 | 0/3 | 2 | 1 | 3/3 |
| GPT-6.1 Sol | 0/3 | 3 | 0 | 3/3 |
| Gemini 3.1 Pro Preview | 0/3 | 3 | 0 | 3/3 |
| Gemini 3.8 Flash | 0/3 | 2 | 1 | 3/3 |
| GLM-5 | 0/3 | 3 | 0 | 1/3 |
| gpt-oss-120b | 0/3 | 3 | 0 | 2/3 | *Keyword heuristic, see limits.
In my earlier local test runs (Gemini 3.8 Flash Preview), the failures were concrete: it read the negation word dave ("not") as a person's name and greeted me as "Dave", built a data-usage tracker instead of an expense tracker, and answered one prompt in Shona and another in Kirundi. Models may be matching words like riduka ("shop") to the nearest language they know, but that is a guess from a tiny sample.
What surprised me: the code task hid the problem completely. A model that passes every unit test can still be unable to talk to the person asking.
What I would measure next: more and harder problems, multi-turn agent tasks where the user answers back, a proper language identifier validated by native speakers, and more Luhya varieties and other Kenyan languages.
[https://www.kaggle.com/benchmarks/eveliaveldrine/language-barrier-benchmark](https://www.kaggle.com/benchmarks/eveliaveldrine/language-barrier-benchmark)
If you speak Maragoli or another Luhya variety and see something wrong in my prompts, tell me in the comments.