Soundwave: Less is More for Speech-Text Alignment in LLMs

Yuhao Zhang, Chihang Lau, Fan Bu, Ruiyu Zhang, Wang, Haizhou Li

Soundwave: Less is More for Speech-Text Alignment in LLMs: 78 upvotes on Hugging Face Daily Papers, #1 of 33 papers on 2025-02-19. Day-by-day upvote history. It lost 8 votes when the Hub removed votes in bulk.

Existing end-to-end speech large language models (LLMs) usually rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth. We focus on two fundamental problems between speech and text: the representation space gap and sequence length inconsistency. We propose Soundwave, which utilizes an efficient training strategy and a novel architecture to address these issues. Results show that Soundwave outperforms the advanced Qwen2-Audio in speech translation and AIR-Bench speech tasks, using only one-fiftieth of the training data. Further analysis shows that Soundwave still retains its intelligence during conversation. The project is available at https://github.com/FreedomIntelligence/Soundwave.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.