Ghosh

Ghosh on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 262 upvotes.

  1. Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 11 upvotes, #19 of 2026-07-20
  2. Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music 28 upvotes, #12 of 2026-04-14
  3. Do Audio-Visual Large Language Models Really See and Hear? 7 upvotes, #32 of 2026-04-07
  4. MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos 13 upvotes, #18 of 2026-03-17
  5. Music Flamingo: Scaling Music Understanding in Audio Language Models 9 upvotes, #10 of 2025-11-14
  6. OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM 79 upvotes, #2 of 2025-10-20
  7. MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence 4 upvotes, #18 of 2025-08-20
  8. Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models 9 upvotes, #12 of 2025-07-14
  9. MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark 19 upvotes, #5 of 2024-10-28
  10. Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation 8 upvotes, #20 of 2024-10-18
  11. ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds 10 upvotes, #7 of 2024-09-17
  12. GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities 18 upvotes, #9 of 2024-06-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.