Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu
A study of 93 Urdu stories generated by three contemporary LLMs — GPT-5.1, Qwen-3-Max, and DeepSeek-3.1 — found that the models frequently make basic grammar and semantic errors, produce incoherent and repetitively unnat…