{"slug": "talkfa-a-unified-benchmark-for-farsi-dialogue-generation-and-understanding", "title": "TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding", "summary": "Researchers introduced TALKFA, a unified benchmark for Farsi dialogue generation and understanding, comprising three datasets: WIKI-FADIAL (4.2K Wikipedia-grounded dialogues), DAILYDIALOG-FA (6.6K dialogues with dialogue acts and emotions), and PLAYDIAL-FA (2.1K theatrical dialogues with sentiment labels). Experiments with six LLAMA and MISTRAL models showed that LoRA improves dialogue generation while requiring only 25-50% of training data to recover over 90% of performance gains, and that automatic metrics overestimate dialogue quality compared to GPT-4.1 as a judge.", "body_md": "arXiv:2609.01810v1 Announce Type: new\nAbstract: Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialogues annotated for dialogue acts and emotions; and (3) PLAYDIAL-FA, 2.1K theatrical dialogues with sentiment labels. While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialogues are released. Experiments with six LLAMA and MISTRAL models show that LoRA substantially improves dialogue generation while requiring only 25-50% of the training data to recover over 90% of the final performance gains. Across classification tasks, FABERT achieves the best dialogue-act performance, LORA-MISTRAL-7B performs best on emotion recognition, and MISTRAL-24B achieves the highest sentiment score. Human evaluation and independent external validation demonstrate the reliability of the benchmark, while comparisons with GPT-4.1 as an LLM judge reveal that automatic metrics substantially overestimate dialogue quality. Zero-shot evaluation with frontier LLMs further shows that TalkFa remains a challenging benchmark. We will release all datasets, annotation guidelines, code, and checkpoints.", "url": "https://wpnews.pro/news/talkfa-a-unified-benchmark-for-farsi-dialogue-generation-and-understanding", "canonical_source": "https://arxiv.org/abs/2609.01810", "published_at": "2026-09-03 04:00:00+00:00", "updated_at": "2026-09-03 04:25:34.069707+00:00", "lang": "en", "topics": ["natural-language-processing", "large-language-models", "ai-research"], "entities": ["TALKFA", "WIKI-FADIAL", "DAILYDIALOG-FA", "PLAYDIAL-FA", "LLAMA", "MISTRAL", "FABERT", "GPT-4.1"], "alternates": {"html": "https://wpnews.pro/news/talkfa-a-unified-benchmark-for-farsi-dialogue-generation-and-understanding", "markdown": "https://wpnews.pro/news/talkfa-a-unified-benchmark-for-farsi-dialogue-generation-and-understanding.md", "text": "https://wpnews.pro/news/talkfa-a-unified-benchmark-for-farsi-dialogue-generation-and-understanding.txt", "jsonld": "https://wpnews.pro/news/talkfa-a-unified-benchmark-for-farsi-dialogue-generation-and-understanding.jsonld"}}