Xu Tan

Xu Tan on Hugging Face Daily Papers: 24 papers, 9 in the top 3 of their day, 630 upvotes.

  1. Chain-of-Model Learning for Language Model 107 upvotes, #1 of 2025-05-20
  2. Kimi-Audio Technical Report 14 upvotes, #7 of 2025-04-28
  3. The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation 7 upvotes, #15 of 2025-03-07
  4. Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis 22 upvotes, #6 of 2025-02-07
  5. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey 49 upvotes, #3 of 2024-12-30
  6. Foundation Models for Music: A Survey 35 upvotes, #3 of 2024-08-27
  7. E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS 18 upvotes, #12 of 2024-07-02
  8. VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers 11 upvotes, #5 of 2024-06-11
  9. FlashSpeech: Efficient Zero-Shot Speech Synthesis 27 upvotes, #4 of 2024-04-24
  10. MuPT: A Generative Symbolic Music Pretrained Transformer 14 upvotes, #5 of 2024-04-10
  11. RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis 7 upvotes, #10 of 2024-04-05
  12. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models 28 upvotes, #3 of 2024-03-06
  13. Beyond Language Models: Byte Models are Digital World Simulators 52 upvotes, #3 of 2024-03-01
  14. CoMoSVC: Consistency Model-based Singing Voice Conversion 10 upvotes, #8 of 2024-01-04
  15. Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis 34 upvotes, #2 of 2023-12-07
  16. MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models 24 upvotes, #2 of 2023-10-19
  17. UniAudio: An Audio Foundation Model Toward Universal Audio Generation 20 upvotes, #7 of 2023-10-06
  18. Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers 53 upvotes, #1 of 2023-09-18
  19. PromptTTS 2: Describing and Generating Voices with Text Prompt 15 upvotes, #6 of 2023-09-06
  20. EmoGen: Eliminating Subjective Bias in Emotional Music Generation 5 upvotes, #13 of 2023-07-06
  21. MuseCoco: Generating Symbolic Music from Text 2 upvotes, #14 of 2023-06-02
  22. Deliberate then Generate: Enhanced Prompting Framework for Text Generation 1 upvotes, #8 of 2023-06-01
  23. GETMusic: Generating Any Music Tracks with a Unified Representation and Diffusion Framework 2 upvotes, #12 of 2023-05-19
  24. CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency Model 7 upvotes, #1 of 2023-05-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.