{"slug": "robust-summarization-of-doctor-patient-conversations-taltech-systems-for-the", "title": "Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge", "summary": "TalTech systems ranked first in both tracks of the Beyond Transcription Challenge (BeTraC), generating SOAP notes directly from long doctor-patient conversation recordings without intermediate transcription. The team adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning and DAPO reinforcement learning using the Open Medical Concept F1 metric as reward, achieving the lowest hallucination rate among all submissions. The findings show that reinforcement learning against a concept-matching metric need not compromise factual reliability and that fine-tuning on text transcripts transfers well to speech input.", "body_md": "arXiv:2607.17230v1 Announce Type: new\nAbstract: This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recordings, without intermediate transcription. After screening open-weight speech LLMs for long-audio robustness, we adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning followed by DAPO reinforcement learning that uses the challenge metric, Open Medical Concept F1, as its reward. Our systems ranked first in both tracks, and an independent LLM-as-a-judge evaluation showed the lowest hallucination rate among all submissions, indicating that reinforcement learning against a concept-matching metric need not compromise factual reliability. We also find that fine-tuning on text transcripts transfers well to speech input and appears to improve robustness on out-of-domain real recordings.", "url": "https://wpnews.pro/news/robust-summarization-of-doctor-patient-conversations-taltech-systems-for-the", "canonical_source": "https://www.machinebrief.com/news/robust-summarization-of-doctor-patient-conversations-taltech-pug0", "published_at": "2026-07-21 04:00:00+00:00", "updated_at": "2026-07-21 04:37:16.137695+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "natural-language-processing", "ai-research"], "entities": ["TalTech", "Voxtral Mini", "Voxtral Small", "Beyond Transcription Challenge", "BeTraC", "Open Medical Concept F1", "LoRA", "DAPO"], "alternates": {"html": "https://wpnews.pro/news/robust-summarization-of-doctor-patient-conversations-taltech-systems-for-the", "markdown": "https://wpnews.pro/news/robust-summarization-of-doctor-patient-conversations-taltech-systems-for-the.md", "text": "https://wpnews.pro/news/robust-summarization-of-doctor-patient-conversations-taltech-systems-for-the.txt", "jsonld": "https://wpnews.pro/news/robust-summarization-of-doctor-patient-conversations-taltech-systems-for-the.jsonld"}}