Daily Papers of 2024-08-07

  1. MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models 56 upvotes, #1 of 2024-08-07
  2. LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
  3. An Object is Worth 64x64 Pixels: Generating 3D Object via Image Diffusion 33 upvotes, #3 of 2024-08-07
  4. MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine 23 upvotes, #4 of 2024-08-07
  5. IPAdapter-Instruct: Resolving Ambiguity in Image-based Conditioning using Instruct Prompts 19 upvotes, #5 of 2024-08-07
  6. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters 15 upvotes, #6 of 2024-08-07
  7. Diffusion Models as Data Mining Tools 12 upvotes, #7 of 2024-08-07
  8. CoverBench: A Challenging Benchmark for Complex Claim Verification 11 upvotes, #8 of 2024-08-07
  9. StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation 8 upvotes, #9 of 2024-08-07
  10. ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer 8 upvotes, #9 of 2024-08-07
  11. Synthesizing Text-to-SQL Data from Weak and Strong LLMs 6 upvotes, #11 of 2024-08-07
  12. AVESFormer: Efficient Transformer Design for Real-Time Audio-Visual Segmentation 2 upvotes, #12 of 2024-08-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.