Xu Tan
Xu Tan on Hugging Face Daily Papers: 24 papers, 9 in the top 3 of their day, 630 upvotes.
- Chain-of-Model Learning for Language Model 107 upvotes, #1 of 2025-05-20
- Kimi-Audio Technical Report 14 upvotes, #7 of 2025-04-28
- The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation 7 upvotes, #15 of 2025-03-07
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis 22 upvotes, #6 of 2025-02-07
- Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey 49 upvotes, #3 of 2024-12-30
- Foundation Models for Music: A Survey 35 upvotes, #3 of 2024-08-27
- E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS 18 upvotes, #12 of 2024-07-02
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers 11 upvotes, #5 of 2024-06-11
- FlashSpeech: Efficient Zero-Shot Speech Synthesis 27 upvotes, #4 of 2024-04-24
- MuPT: A Generative Symbolic Music Pretrained Transformer 14 upvotes, #5 of 2024-04-10
- RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis 7 upvotes, #10 of 2024-04-05
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models 28 upvotes, #3 of 2024-03-06
- Beyond Language Models: Byte Models are Digital World Simulators 52 upvotes, #3 of 2024-03-01
- CoMoSVC: Consistency Model-based Singing Voice Conversion 10 upvotes, #8 of 2024-01-04
- Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis 34 upvotes, #2 of 2023-12-07
- MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models 24 upvotes, #2 of 2023-10-19
- UniAudio: An Audio Foundation Model Toward Universal Audio Generation 20 upvotes, #7 of 2023-10-06
- Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers 53 upvotes, #1 of 2023-09-18
- PromptTTS 2: Describing and Generating Voices with Text Prompt 15 upvotes, #6 of 2023-09-06
- EmoGen: Eliminating Subjective Bias in Emotional Music Generation 5 upvotes, #13 of 2023-07-06
- MuseCoco: Generating Symbolic Music from Text 2 upvotes, #14 of 2023-06-02
- Deliberate then Generate: Enhanced Prompting Framework for Text Generation 1 upvotes, #8 of 2023-06-01
- GETMusic: Generating Any Music Tracks with a Unified Representation and Diffusion Framework 2 upvotes, #12 of 2023-05-19
- CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency Model 7 upvotes, #1 of 2023-05-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.