cd /news/large-language-models/cross-lingual-transfer-in-tulu-legal… · home topics large-language-models article
[ARTICLE · art-117326] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict

A new arXiv study (2608.28645v1) found that three large language models—Llama3, Hex-1, and Sarvam—showed script-dependent improvement in classifying legal complaints in the low-resource Dravidian language Tulu when queries were transliterated across Dravidian scripts, with Kannada producing the strongest positive trend. Retrieval-augmented generation (RAG) using Kannada legal papers yielded mixed results, with failures attributed to fact substitution and confabulation, indicating that reasoning failures stem from the models' parsing and reasoning rather than corpus contents.

read1 min views1 publishedSep 1, 2026

arXiv:2608.28645v1 Announce Type: new Abstract: Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, there is a need to develop rigorous evaluation methods to identify when models fail in cross lingual low-resource environments. Using the legal domain as a backdrop, three models (Llama3, Hex-1, Sarvam) were tested on the ability to classify legal complaints written in a low resource Dravidian language (Tulu). Transliterating queries across Dravidian scripts allowed models to gain a preliminary understanding of speakers' complaints without the use of wide scale training, though the level of comprehension was heavily script dependent (with Kannada - another relatively low-resource language - producing the strongest positive trend). Retrieving from a corpus of Kannada legal papers across a RAG framework caused mixed results. Some models had a weak positive trend in comprehension under certain conditions, but when models failed, it was often across two axes: fact substitution (fixating on specific passage excerpts that skewed reasoning) and confabulation (hallucination that had no basis in either query or corpus). Within low resource domains, results identify the model's parsing of information and subsequent reasoning as the source of reasoning failure, rather than corpus contents. Script-dependent comprehension and RAG robustness also seem to travel together. This is further supported by the reasoning-trace analysis and a statistical-honesty framework deployed - techniques that are more broadly applicable to low-resource multilingual RAG evaluation.

── more in #large-language-models 4 stories · sorted by recency
── more on @llama3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cross-lingual-transf…] indexed:0 read:1min 2026-09-01 ·