cd /news/large-language-models/how-robust-are-llms-to-vietnamese-di… · home topics large-language-models article
[ARTICLE · art-93008] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

How Robust Are LLMs to Vietnamese Dialects?

Researchers introduced VialectBench, the first systematic benchmark for evaluating LLM robustness to Vietnamese dialects, containing 400 Standard Vietnamese instances and 2,400 human-written dialectal rewrites across six dialect groups. Testing ten instruction-tuned models, they found dialectal inputs reduce average performance by 2.82%, with QA showing the largest degradation and no model fully dialect-invariant. The Central dialect group caused the highest average harmful-flip rate at 6.54%, while PNB slightly improved performance by 0.42%.

read1 min views1 publishedAug 12, 2026

arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalization instead of measuring how the model fails under Vietnamese dialectal inputs. To address this gap, we present the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks, quantifying performance degradation and failure patterns. We introduce VialectBench (Vietnamese Dialects Benchmarking), a controlled benchmark for testing whether model decisions remain stable across six Vietnamese dialect groups. VialectBench contains 400 Standard Vietnamese source instances and 2,400 human-written dialectal rewrites spanning emotion recognition (ER), natural language inference (NLI), question answering (QA), and multiple-choice question answering (MCQA). Dataset evaluation with a fixed reference language model shows that the dialectal rewrites induce a measurable model-relative likelihood shift while remaining nearly equal in length to their Standard counterparts. Across ten instruction-tuned models, dialectal inputs reduce average performance by 2.82%, and no evaluated model is fully dialect-invariant. All four tasks are affected, with QA showing the largest average degradation. Robustness also varies substantially across dialect groups: PNT3 and PNT2 cause the largest average performance drops, at 6.17% and 4.73%, respectively, whereas PNB slightly improves average performance by 0.42%. The Central dialect group (PNT1-PNT4) also yields the highest average harmful-flip rate across all models, at 6.54%. These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation.

── more in #large-language-models 4 stories · sorted by recency
── more on @vialectbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-robust-are-llms-…] indexed:0 read:1min 2026-08-12 ·