A comprehensive survey paper titled "Recent Advances in Speech Language Models: A Survey," authored by a Chinese University of Hong Kong team, has been accepted by the ACL 2025 main conference, marking the first systematic review in this field. The survey positions Speech Large Models (SpeechLM) as the next frontier in AI, aiming to overcome the limitations of traditional speech interaction, such as information loss, severe latency, and error accumulation, by enabling end-to-end natural human-machine speech dialogue. The paper meticulously details SpeechLM's technical architecture, comprising speech tokenizers, language models, and vocoders, along with its training strategies, including pre-training and instruction tuning. It explores advanced interaction paradigms like full-duplex communication, highlighting SpeechLM's broad applications in natural dialogue, personalized assistants, and emotional speech generation. Furthermore, the survey discusses evaluation methodologies and future challenges, such as optimization, real-time processing, safety, and support for low-resource languages. This seminal work anticipates that SpeechLM will usher in a new era of speech AI, fundamentally transforming human-computer interaction by enabling more natural and intuitive communication.
Speech Large ModelsSpeech InteractionHuman-Computer InteractionSurveyMultimodalLarge Language ModelNatural Language ProcessingMultimodal