cd /news/natural-language-processing/talkfa-a-unified-benchmark-for-farsi… · home topics natural-language-processing article
[ARTICLE · art-119802] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

Researchers introduced TALKFA, a unified benchmark for Farsi dialogue generation and understanding, comprising three datasets: WIKI-FADIAL (4.2K Wikipedia-grounded dialogues), DAILYDIALOG-FA (6.6K dialogues with dialogue acts and emotions), and PLAYDIAL-FA (2.1K theatrical dialogues with sentiment labels). Experiments with six LLAMA and MISTRAL models showed that LoRA improves dialogue generation while requiring only 25-50% of training data to recover over 90% of performance gains, and that automatic metrics overestimate dialogue quality compared to GPT-4.1 as a judge.

read1 min views1 publishedSep 3, 2026

arXiv:2609.01810v1 Announce Type: new Abstract: Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialogues annotated for dialogue acts and emotions; and (3) PLAYDIAL-FA, 2.1K theatrical dialogues with sentiment labels. While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialogues are released. Experiments with six LLAMA and MISTRAL models show that LoRA substantially improves dialogue generation while requiring only 25-50% of the training data to recover over 90% of the final performance gains. Across classification tasks, FABERT achieves the best dialogue-act performance, LORA-MISTRAL-7B performs best on emotion recognition, and MISTRAL-24B achieves the highest sentiment score. Human evaluation and independent external validation demonstrate the reliability of the benchmark, while comparisons with GPT-4.1 as an LLM judge reveal that automatic metrics substantially overestimate dialogue quality. Zero-shot evaluation with frontier LLMs further shows that TalkFa remains a challenging benchmark. We will release all datasets, annotation guidelines, code, and checkpoints.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @talkfa 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/talkfa-a-unified-ben…] indexed:0 read:1min 2026-09-03 ·