A few shades of supervision for discourse segmentation: experiments on a French conversational corpus Laurent Prevot and Philippe Muller, in a study published in Dialogue Discourse Volume 16, found that classical fine-tuning remains the most effective approach for segmenting Elementary Discourse Units in spontaneous French speech, outperforming weakly supervised and in-context learning methods. Using an 8-hour French conversational corpus, they showed that only a reasonable amount of gold-annotated data is needed for best performance, and they identified prosodic integration and handling of pauses with disfluencies as key areas for improvement. Abstract Elementary Discourse Units EDUs constitutes the interface between language grammar and lan- guage use. On the one hand, they result from compositional semantic processes that combines individual word meanings into proposition-level representations. On the other hand, EDUs form the building blocks of most text, discourse, and dialogue frameworks. In written genres, where punctuation is available and reliable, segmenting EDUs is sometimes seen as a nearly solved problem, as least for high-resource languages. However, this is not the case for spontaneous speech transcripts. In this paper, we use a significant 8-hour French corpus, manually segmented into EDUs, to evaluate several large language model LLM -based approaches for this task. We compare various fine-tuning strategies, including those relying on weakly supervised labels, in relation to the amount of ”gold” manual annotations that can be available. We also experiment with in-context learning, where example instances are provided to condition a generative model few-shots learning or in a purely generative approach zero-shot . Our findings indicate that classical fine-tuning is still the most effective approach, requiring only a reasonable amount of gold-annotated data to achieve the best performance in our experiments. Beyond traditional quantitative evaluation, we conducted a systematic qualitative analysis, identifying directions for further improvement. These include integrating prosodic considerations while handling pauses when they co-occur with disfluencies or complex discourse markers uses. Finally, we argue for the significance of this task and the resulting units, compared to acoustic and syntactic proxies, especially for quantitative linguistics focusing on spontaneous speech. - Anthology ID: - 2025.dnd-16.5 - Volume: - Dialogue Discourse Volume 16 https://aclanthology.org/volumes/2025.dnd-16/ - Month: - 12 - Year: - 2025 - Address: - Editors: - Hendrik Buschmeier https://aclanthology.org/people/hendrik-buschmeier/ , Barbara Di Eugenio https://aclanthology.org/people/barbara-di-eugenio/ , Patrick Healey https://aclanthology.org/people/patrick-healey/ , Casey Kennington https://aclanthology.org/people/casey-kennington/ , David Schlangen https://aclanthology.org/people/david-schlangen/ , Manfred Stede https://aclanthology.org/people/manfred-stede/ , Amir Zeldes https://aclanthology.org/people/amir-zeldes/ - Venue: - DND https://aclanthology.org/venues/dnd/ - SIG: - SIGDIAL https://aclanthology.org/sigs/sigdial/ - Publisher: - Note: - Pages: - 35–73 - Language: - URL: - https://aclanthology.org/2025.dnd-16.5/ https://aclanthology.org/2025.dnd-16.5/ - DOI: - 10.5210/dad.2025.202 https://doi.org/10.5210/dad.2025.202 - Cite ACL : - Laurent Prevot and Philippe Muller. 2025. A few shades of supervision for discourse segmentation: experiments on a French conversational corpus https://aclanthology.org/2025.dnd-16.5/ . In Dialogue Discourse Volume 16 , pages 35–73. - Cite Informal : - A few shades of supervision for discourse segmentation: experiments on a French conversational corpus https://aclanthology.org/2025.dnd-16.5/ Prevot & Muller, DND 2025 - PDF: - https://aclanthology.org/2025.dnd-16.5.pdf https://aclanthology.org/2025.dnd-16.5.pdf