Improving Factual Accuracy in Neural Data-to-Text Generation through Input Quality and Scalable Evaluation Barkavi Sundararajan's PhD research, presented at the 1st Workshop for Young Researchers in Natural Language Generation in Hanoi, Vietnam, in October 2025, proposes two approaches to reduce hallucinations in neural data-to-text generation: analyzing input quality and structure, and developing an LLM-as-Judge framework for scalable evaluation. The work aims to improve factual accuracy in large language models (LLMs) for applications where faithfulness to input data is critical. Abstract Neural Language Models have become central to Natural Language Generation NLG research and can produce fluent and coherent text. However, when models generate text from complex, structured or long-form data such as tables or event logs, they often hallucinate and introduce factual errors. These hallucinations limit the practical deployment of large language models LLMs in applications where factual accuracy is critical. In my research, factual accuracy refers to the faithfulness of the generated text to the given input data. My PhD focuses on reducing hallucinations and improving factual accuracy in data-to-text generation, which I address through two core approaches: i analysing how input quality and structure improve factual accuracy, and ii developing a manual error annotation protocol and extending it into an LLM-as-Judge framework. This work aims to assess when automatic evaluation can complement human annotation and enable larger-scale evaluation.- Anthology ID: - 2025.ynlg-main.3 - Volume: Proceedings of the 1st Workshop for Young Researchers in Natural Language Generation /volumes/2025.ynlg-main/ - Month: - October - Year: - 2025 - Address: - Hanoi, Vietnam - Editors: Alyssa Allen /people/alyssa-allen/unverified/ , Nils Feldhus /people/nils-feldhus/ , Rudali Huidrom /people/rudali-huidrom/unverified/ , Michela Lorandi /people/michela-lorandi/ , Adarsa Sivaprasad /people/adarsa-sivaprasad/ , Patrícia Schmidtová /people/patricia-schmidtova/ - Venue: YNLG /venues/ynlg/ - SIG: SIGGEN /sigs/siggen/ - Publisher: - Association for Computational Linguistics - Note: - Pages: - 10–16 - Language: - URL: https://aclanthology.org/2025.ynlg-main.3/ https://aclanthology.org/2025.ynlg-main.3/ - DOI: - Cite ACL : - Barkavi Sundararajan. 2025. Improving Factual Accuracy in Neural Data-to-Text Generation through Input Quality and Scalable Evaluation https://aclanthology.org/2025.ynlg-main.3/ . In Proceedings of the 1st Workshop for Young Researchers in Natural Language Generation , pages 10–16, Hanoi, Vietnam. Association for Computational Linguistics. - Cite Informal : Improving Factual Accuracy in Neural Data-to-Text Generation through Input Quality and Scalable Evaluation https://aclanthology.org/2025.ynlg-main.3/ Sundararajan, YNLG 2025 - PDF: https://aclanthology.org/2025.ynlg-main.3.pdf https://aclanthology.org/2025.ynlg-main.3.pdf