Are you speaking my languages? On spoken language adherence in multimodal LLMs

wpnews.pro

cd /news/large-language-models/are-you-speaking-my-languages-on-spo… · home › topics › large-language-models › article

[ARTICLE · art-30529] src=arxiv.org ↗ pub=2026-06-17T04:00Z topic=large-language-models verified=true sentiment=· neutral

Are you speaking my languages? On spoken language adherence in multimodal LLMs

Researchers at arXiv propose a soft prompting approach to improve language adherence in multimodal LLMs for ASR, introducing a metric to quantify violations and evaluating zero-shot prompting, supervised fine-tuning, and Chain-of-Thought reasoning across multiple languages.

read1 min views1 publishedJun 17, 2026

arXiv:2606.17281v1 Announce Type: new Abstract: While Large Language Model (LLM) based Automatic Speech Recognition (ASR) enables seamless multilingual use, models often misidentify the output language, compromising transcription fidelity and downstream application quality. To preserve flexibility and code-switching capabilities, we propose a soft prompting approach that hints at potential spoken languages without strictly constraining the output. We formally define this challenge as a lack of language adherence, introduce a novel metric to quantify violations, and evaluate three mitigation strategies: (1) zero-shot prompting for robust guidance under uncertainty, (2) supervised fine-tuning (SFT) to improve prompt adherence, and (3) Chain-of-Thought (CoT) reasoning to enforce adherence during decoding. We present a comparative analysis of these methods across multiple languages, evaluating effectiveness in reducing the language violation while maintaining overall ASR performance. Finally, we discuss trade-offs to guide strategy selection under various compute constraints.

source & further reading

arxiv.org — original article

~/api · this article 200

$curl api.wpnews.pro/v1/news/are-you-speaking-my-lang…

Read original on arxiv.org → arxiv.org/abs/2606.17281

mentioned entities

arXiv

metadata

slugare-you-speaking-my-languages-on-spoken-language-adherence-in-multimodal-llms

topic#large-language-models

secondary2 topics

sentimentneutral

canonicalarxiv.org

navigation

← prevRay Data LLM enables 2x throughp…

next →Claude Agent SDK Permissions: An…

── more in #large-language-models 4 stories · sorted by recency

arxiv.org · 17 Jun · #large-language-models

Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty

arxiv.org · 17 Jun · #large-language-models

MemTrace: Probing What Final Accuracy Misses in Long-Term Memory

arxiv.org · 17 Jun · #large-language-models

MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors

arxiv.org · 17 Jun · #large-language-models

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

── more on @arxiv 3 stories trending now

wpnews · 16 Jun · #ai-agents

The LLM Is Not the Final Authority: Building Trust Infrastructure for AI Agents

wpnews · 16 Jun · #artificial-intelligence

Most Businesses Lose Leads at Night — So I Built This

wpnews · 16 Jun · #ai-safety

Researchers propose causal framework to audit synthetic data

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required