Daily Papers of 2024-10-15

  1. LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models 53 upvotes, #1 of 2024-10-15
  2. MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models 48 upvotes, #2 of 2024-10-15
  3. Animate-X: Universal Character Image Animation with Enhanced Motion Representation 46 upvotes, #3 of 2024-10-15
  4. Toward General Instruction-Following Alignment for Retrieval-Augmented Generation 43 upvotes, #4 of 2024-10-15
  5. MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks 34 upvotes, #5 of 2024-10-15
  6. Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 26 upvotes, #6 of 2024-10-15
  7. Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations 26 upvotes, #6 of 2024-10-15
  8. LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content 25 upvotes, #8 of 2024-10-15
  9. Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention 23 upvotes, #9 of 2024-10-15
  10. VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents 21 upvotes, #10 of 2024-10-15
  11. Rethinking Data Selection at Scale: Random Selection is Almost All You Need 14 upvotes, #11 of 2024-10-15
  12. TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models 14 upvotes, #11 of 2024-10-15
  13. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory 9 upvotes, #13 of 2024-10-15
  14. Tree of Problems: Improving structured problem solving with compositionality 8 upvotes, #14 of 2024-10-15
  15. MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models 7 upvotes, #15 of 2024-10-15
  16. Thinking LLMs: General Instruction Following with Thought Generation 7 upvotes, #15 of 2024-10-15
  17. Generalizable Humanoid Manipulation with Improved 3D Diffusion Policies 6 upvotes, #17 of 2024-10-15
  18. TVBench: Redesigning Video-Language Evaluation 5 upvotes, #18 of 2024-10-15
  19. The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling 5 upvotes, #18 of 2024-10-15
  20. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads 5 upvotes, #18 of 2024-10-15
  21. ReLU's Revival: On the Entropic Overload in Normalization-Free Large Language Models 3 upvotes, #21 of 2024-10-15
  22. Latent Action Pretraining from Videos 2 upvotes, #22 of 2024-10-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.