Daily Papers of 2024-10-15
- LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models 53 upvotes, #1 of 2024-10-15
- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models 48 upvotes, #2 of 2024-10-15
- Animate-X: Universal Character Image Animation with Enhanced Motion Representation 46 upvotes, #3 of 2024-10-15
- Toward General Instruction-Following Alignment for Retrieval-Augmented Generation 43 upvotes, #4 of 2024-10-15
- MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks 34 upvotes, #5 of 2024-10-15
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 26 upvotes, #6 of 2024-10-15
- Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations 26 upvotes, #6 of 2024-10-15
- LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content 25 upvotes, #8 of 2024-10-15
- Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention 23 upvotes, #9 of 2024-10-15
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents 21 upvotes, #10 of 2024-10-15
- Rethinking Data Selection at Scale: Random Selection is Almost All You Need 14 upvotes, #11 of 2024-10-15
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models 14 upvotes, #11 of 2024-10-15
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory 9 upvotes, #13 of 2024-10-15
- Tree of Problems: Improving structured problem solving with compositionality 8 upvotes, #14 of 2024-10-15
- MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models 7 upvotes, #15 of 2024-10-15
- Thinking LLMs: General Instruction Following with Thought Generation 7 upvotes, #15 of 2024-10-15
- Generalizable Humanoid Manipulation with Improved 3D Diffusion Policies 6 upvotes, #17 of 2024-10-15
- TVBench: Redesigning Video-Language Evaluation 5 upvotes, #18 of 2024-10-15
- The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling 5 upvotes, #18 of 2024-10-15
- DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads 5 upvotes, #18 of 2024-10-15
- ReLU's Revival: On the Entropic Overload in Normalization-Free Large Language Models 3 upvotes, #21 of 2024-10-15
- Latent Action Pretraining from Videos 2 upvotes, #22 of 2024-10-15
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.