cd /news/large-language-models/does-your-coding-agent-understand-yo… · home › topics › large-language-models › article
[ARTICLE · art-148933] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Does your coding agent understand you in your native language?

A developer in Kenya benchmarked 16 coding models across six languages — English, Chinese, French, Spanish, Swahili and the low-resource Luhya language Maragoli — finding that 11 of 16 scored 100% on unit-test-based Python coding tasks regardless of prompt language, but that none of the seven models tested on planning tasks replied in Maragoli. In the planning task, all 21 Maragoli replies came back in English, Swahili or another language, and one model misread the Maragoli negation word "dave" as a person's name and greeted the user as "Dave". The developer concluded that English code scaffolding masks the language barrier, noting "a model that passes every unit test can still be unable to talk to the person asking.

by read3 min views1 publishedOct 10, 2026

This is a submission for the Kaggle Benchmarking Challenge I build AI tools for developers and learners in Kenya, so I kept asking one question: if you brief a coding agent in a language it barely saw in training, does it still do the job?

I tested six languages: English, Chinese, French, Spanish, Swahili, and Maragoli (Lulogooli), a Luhya language spoken in western Kenya. I wrote the Maragoli prompts myself.

Task 1: Coding. 5 small Python problems, each described in all 6 languages (30 prompts per model). Function names, signatures and unit tests stay identical, so only the instruction language changes. Score = unit-test pass rate.

Task 2: Planning. A good agent plans before it codes, and planning means asking the user clarifying questions. I gave 3 deliberately vague requests ("build me a small app to track my expenses", "write a script that cleans up my files", "make a website for my shop") in all 6 languages, each ending with: make a short plan, ask clarifying questions, and do not write code yet. I check whether the model asks a question, follows "no code yet", and replies in the language it was spoken to.

Task 1 is a control: the English code scaffolding makes the language barrier easy to cross. Task 2 removes that scaffolding.

16 models ran Task 1, and 7 also completed Task 2: Claude Haiku 5.5 and Sonnet 5.5, GPT-6.1 Sol, Gemini 3.1 Pro Preview , a Gemini 3.8 Flash model, GLM-5 and gpt-oss-120b. I picked a spread of providers and sizes, including Chinese-lab models (GLM, Qwen, DeepSeek). Opus 5.5 and GPT-6 Astra scored 100% on Task 1 but could not run Task 2 (free quota), so the "0" on their leaderboard Overall reflects a failed run, not their ability. Grok was not available on the platform.

Short answer: for code, mostly yes. For a conversation in Maragoli, no.

Task 1. 11 of 16 models scored 100%. The rest: gpt-oss-120b, GLM-5 and Gemma 4 31B at 96.7%, gpt-oss-20b at 93.3%, and Qwen3 Next 80B Thinking at 86.7%. The language barrier barely mattered, because the code scaffolding was in English.

Task 2, Maragoli prompts (7 models x 3 prompts = 21 replies):

Model Replied in Maragoli* English or Swahili Another language Followed "no code yet"
Claude Haiku 5.5 0/3 3 0 2/3
Claude Sonnet 5.5 0/3 2 1 3/3
GPT-6.1 Sol 0/3 3 0 3/3
Gemini 3.1 Pro Preview 0/3 3 0 3/3
Gemini 3.8 Flash 0/3 2 1 3/3
GLM-5 0/3 3 0 1/3

| gpt-oss-120b | 0/3 | 3 | 0 | 2/3 | *Keyword heuristic, see limits.

In my earlier local test runs (Gemini 3.8 Flash Preview), the failures were concrete: it read the negation word dave ("not") as a person's name and greeted me as "Dave", built a data-usage tracker instead of an expense tracker, and answered one prompt in Shona and another in Kirundi. Models may be matching words like riduka ("shop") to the nearest language they know, but that is a guess from a tiny sample.

What surprised me: the code task hid the problem completely. A model that passes every unit test can still be unable to talk to the person asking.

What I would measure next: more and harder problems, multi-turn agent tasks where the user answers back, a proper language identifier validated by native speakers, and more Luhya varieties and other Kenyan languages.

[https://www.kaggle.com/benchmarks/eveliaveldrine/language-barrier-benchmark](https://www.kaggle.com/benchmarks/eveliaveldrine/language-barrier-benchmark)

If you speak Maragoli or another Luhya variety and see something wrong in my prompts, tell me in the comments.
── more in #large-language-models 4 stories · sorted by recency
── more on @maragoli 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/does-your-coding-age…] indexed:0 read:3min 2026-10-10 · —