cd /news/large-language-models/evaluating-prompt-scope-and-demonstr… · home topics large-language-models article
[ARTICLE · art-79721] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation

A new arXiv study evaluating local LLM machine translation finds that dedicated MT systems from OPUS-MT and NLLB-200 remain stronger overall than three instruction-tuned LLMs (llama3.2:3b, mistral:latest, qwen2.5:14b) on English-to-Romance and English-to-Germanic translation across nine EU languages. Few-shot prompting helps mistral:latest and qwen2.5:14b but hurts llama3.2:3b, while family-scope prompting is feasible for stronger models but exposes failures in smaller ones.

read1 min views1 publishedJul 30, 2026

arXiv:2607.26286v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt shape: translate one source sentence into one target language. In practice, users may ask for one target language, for several related languages at once, or for translations conditioned on examples. This paper studies prompt scope and demonstration selection as experimental variables for local LLM machine translation. We evaluate English-to-Romance and English-to-Germanic translation on the full FLORES devtest split for nine official European Union languages. We compare three local instruction-tuned LLMs, llama3.2:3b, mistral:latest, and qwen2.5:14b, against dedicated MT baselines from OPUS-MT and NLLB-200. We test zero-shot prompting and k=5 few-shot prompting with random, lexical-similarity, and embedding-similarity demonstration selection. We also compare single-target prompts with JSON-formatted family-scope prompts that request all languages in a family at once. Results show that dedicated MT systems remain strongest overall, especially for Germanic languages. Few-shot prompting helps mistral:latest and qwen2.5:14b, but hurts llama3.2:3b; embedding retrieval is best on average for the stronger LLMs, but its advantage over random and lexical examples is modest. Family-scope prompting is feasible for stronger local LLMs but exposes structured-output failures in smaller models. These findings motivate evaluating LLM translation not only by language pair and metric, but also by prompt scope, retrieval strategy, and multi-target compliance.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-prompt-sc…] indexed:0 read:1min 2026-07-30 ·