Daily Papers of 2024-08-07
- MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models 56 upvotes, #1 of 2024-08-07
- LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
- An Object is Worth 64x64 Pixels: Generating 3D Object via Image Diffusion 33 upvotes, #3 of 2024-08-07
- MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine 23 upvotes, #4 of 2024-08-07
- IPAdapter-Instruct: Resolving Ambiguity in Image-based Conditioning using Instruct Prompts 19 upvotes, #5 of 2024-08-07
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters 15 upvotes, #6 of 2024-08-07
- Diffusion Models as Data Mining Tools 12 upvotes, #7 of 2024-08-07
- CoverBench: A Challenging Benchmark for Complex Claim Verification 11 upvotes, #8 of 2024-08-07
- StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation 8 upvotes, #9 of 2024-08-07
- ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer 8 upvotes, #9 of 2024-08-07
- Synthesizing Text-to-SQL Data from Weak and Strong LLMs 6 upvotes, #11 of 2024-08-07
- AVESFormer: Efficient Transformer Design for Real-Time Audio-Visual Segmentation 2 upvotes, #12 of 2024-08-07
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.