{"slug": "on-improving-faithfulness-of-podcasts-from-documents", "title": "On Improving Faithfulness of Podcasts from Documents", "summary": "A new study from arXiv finds that large language models (LLMs) frequently generate ungrounded content when producing podcast transcripts from documents, with even GPT-4o showing significant unfaithfulness. The researchers introduce a dataset of over 1500 documents and a turn-level LLM-as-a-judge evaluation framework, along with a model-agnostic 'catch-n-repair' method that improves faithfulness across domains.", "body_md": "arXiv:2607.21961v1 Announce Type: new\nAbstract: Large language models (LLMs) are increasingly used to generate long-form conversational content such as podcasts from textual sources. While these systems produce fluent and engaging narratives, they often introduce ungrounded information. In this work, we present the first systematic study of faithfulness in document-grounded podcast generation, where grounding must be maintained across conversational turns in long-form, multi-speaker transcripts. We construct a dataset of over 1500 documents spanning five domains and generate podcast transcripts using multiple LLMs. We introduce a turn-level LLM-as-a-judge framework for evaluating whether conversational turns are supported by the source document, and validate its reliability through human studies. Our analysis shows that even state-of-the-art models, including GPT-4o, frequently generate ungrounded content. To mitigate this issue, we propose catch-n-repair, a model-agnostic framework that detects and rewrites unfaithful conversational turns while preserving conversational flow. Experiments demonstrate consistent improvements in faithfulness across both in-domain and out-of-domain settings.", "url": "https://wpnews.pro/news/on-improving-faithfulness-of-podcasts-from-documents", "canonical_source": "https://arxiv.org/abs/2607.21961", "published_at": "2026-07-27 04:00:00+00:00", "updated_at": "2026-07-27 04:24:52.048937+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-research", "natural-language-processing"], "entities": ["arXiv", "GPT-4o"], "alternates": {"html": "https://wpnews.pro/news/on-improving-faithfulness-of-podcasts-from-documents", "markdown": "https://wpnews.pro/news/on-improving-faithfulness-of-podcasts-from-documents.md", "text": "https://wpnews.pro/news/on-improving-faithfulness-of-podcasts-from-documents.txt", "jsonld": "https://wpnews.pro/news/on-improving-faithfulness-of-podcasts-from-documents.jsonld"}}