Voice of Reason: Reinforcement Learning for Spoken Math Applying reinforcement learning with verifiable rewards to the GLM-4-Voice speech model raised free-form accuracy on the GSM8K math benchmark to 74.8%, a new state of the art for speech-native models, according to an arXiv paper (2609.18677v1). The authors first adapted GLM-4-Voice with supervised fine-tuning on synthesized spoken question-answering data, then showed RL improved GSM8K accuracy even without extra reasoning tokens, exceeding prior speech-model results that relied on supplementary reasoning traces. Combining RL with existing streaming reasoning techniques produced the further gain to 74.8% free-form accuracy. arXiv:2609.18677v1 Announce Type: new Abstract: Speech language models enable richer spoken interactions between humans and machines than cascaded systems, allowing access to paralinguistic information and lower latency. However, their accuracy on mathematical reasoning benchmarks has lagged behind those of text models. Reinforcement learning RL with verifiable rewards has been instrumental in extending text models' capabilities for solving complex problems and limiting hallucinations. In this work, we explore applying RL to the GLM-4-Voice speech model Zeng et al., 2024 to bridge the gap between textual and spoken mathematical problem solving. We first adapt the model to the domain using supervised fine-tuning on synthesized spoken question-answering data. We then show that, even without extra reasoning tokens, RL improves the accuracy on GSM8K beyond levels previously achieved for speech models only with supplementary reasoning traces. When combined with existing streaming reasoning techniques, we show further gains to 74.8% free-form accuracy. This establishes a new state-of-the-art for mathematical spoken abilities with speech-native models.