Dinesh Manocha

Dinesh Manocha on Hugging Face Daily Papers: 8 papers, 1 in the top 3 of their day, 147 upvotes.

  1. MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos 13 upvotes, #18 of 2026-03-17
  2. Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities 22 upvotes, #6 of 2025-03-07
  3. MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark 19 upvotes, #5 of 2024-10-28
  4. Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation 8 upvotes, #20 of 2024-10-18
  5. Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data 4 upvotes, #25 of 2024-10-04
  6. ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds 10 upvotes, #7 of 2024-09-17
  7. GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities 18 upvotes, #9 of 2024-06-18
  8. HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models 27 upvotes, #2 of 2023-10-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.