cd /news/large-language-models/multilingual-in-name-only-cultural-a… · home topics large-language-models article
[ARTICLE · art-126499] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

A study of 93 Urdu stories generated by three contemporary LLMs — GPT-5.1, Qwen-3-Max, and DeepSeek-3.1 — found that the models frequently make basic grammar and semantic errors, produce incoherent and repetitively unnatural text, and show pervasive cultural shallowness, according to the arXiv paper 2609.10758v1. The researchers manually annotated the generated corpus under a nine-label linguistic, semantic, and cultural taxonomy and reported that few-shot prompting left the cultural and context errors largely unresolved. The findings, which use Urdu as a representative low-resource language, highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.

by read1 min views1 publishedSep 11, 2026

arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.

── more in #large-language-models 4 stories · sorted by recency
── more on @gpt-5.1 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/multilingual-in-name…] indexed:0 read:1min 2026-09-11 ·