cd /news/large-language-models/vakyarth-evaluating-pragmatic-compet… · home topics large-language-models article
[ARTICLE · art-119800] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

Researchers introduced VakyArth, the first pragmatic benchmark for Indic languages, evaluating large language models (LLMs) across Hindi, Punjabi, Tamil, and Malayalam. Testing five pragmatic phenomena, the study found consistent model failures on culturally rooted meanings, with MCQ accuracy exceeding NLI accuracy in all model-language combinations and translation performance not reliably tracking pragmatic understanding.

read1 min views1 publishedSep 3, 2026

arXiv:2609.01788v1 Announce Type: new Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.

── more in #large-language-models 4 stories · sorted by recency
── more on @vakyarth 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/vakyarth-evaluating-…] indexed:0 read:1min 2026-09-03 ·