Ghosh
Ghosh on Hugging Face Daily Papers: 12 papers, 1 in the top 3 of their day, 262 upvotes.
- Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 11 upvotes, #19 of 2026-07-20
- Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music 28 upvotes, #12 of 2026-04-14
- Do Audio-Visual Large Language Models Really See and Hear? 7 upvotes, #32 of 2026-04-07
- MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos 13 upvotes, #18 of 2026-03-17
- Music Flamingo: Scaling Music Understanding in Audio Language Models 9 upvotes, #10 of 2025-11-14
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM 79 upvotes, #2 of 2025-10-20
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence 4 upvotes, #18 of 2025-08-20
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models 9 upvotes, #12 of 2025-07-14
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark 19 upvotes, #5 of 2024-10-28
- Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation 8 upvotes, #20 of 2024-10-18
- ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds 10 upvotes, #7 of 2024-09-17
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities 18 upvotes, #9 of 2024-06-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.