arXiv:2607.10092v1 Announce Type: new Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken-query interaction. We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore's four official languages, combining LoRA fine-tuning, a surrogate text-QA dataset that guards against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech. We also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form. The resulting HT-Moonstone (5B) matches or outperforms SLMs up to 7x its size on most tasks, attains the best accent and gender recognition among all models evaluated, and loses under 2% of its original speech QA ability.
Efficiently Adapting Spoken Language Models for the Singaporean Context
Researchers from the Singaporean Home Team adapted an open-source spoken language model (SLM) to handle five speech tasks in Singapore's four official languages, creating HT-Moonstone (5B). The model matches or outperforms SLMs up to 7x its size on most tasks, achieves best accent and gender recognition among evaluated models, and loses under 2% of its original speech QA ability. The team built the HTD-multilingual-QA dataset with 504,853 samples and used LoRA fine-tuning with a surrogate text-QA dataset and multi-task objective to guard against catastrophic forgetting.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.