Daily Papers of 2024-06-14

  1. Depth Anything V2 83 upvotes, #1 of 2024-06-14
  2. An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels 46 upvotes, #2 of 2024-06-14
  3. Transformers meet Neural Algorithmic Reasoners 41 upvotes, #3 of 2024-06-14
  4. Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling 33 upvotes, #4 of 2024-06-14
  5. OpenVLA: An Open-Source Vision-Language-Action Model 28 upvotes, #5 of 2024-06-14
  6. Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models 27 upvotes, #6 of 2024-06-14
  7. Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning 22 upvotes, #7 of 2024-06-14
  8. DiTFastAttn: Attention Compression for Diffusion Transformer Models 19 upvotes, #8 of 2024-06-14
  9. Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models 17 upvotes, #9 of 2024-06-14
  10. MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding 17 upvotes, #9 of 2024-06-14
  11. Interpreting the Weight Space of Customized Diffusion Models 17 upvotes, #9 of 2024-06-14
  12. CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery 14 upvotes, #12 of 2024-06-14
  13. mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus 14 upvotes, #12 of 2024-06-14
  14. HelpSteer2: Open-source dataset for training top-performing reward models 13 upvotes, #14 of 2024-06-14
  15. EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts 12 upvotes, #15 of 2024-06-14
  16. Explore the Limits of Omni-modal Pretraining at Scale 10 upvotes, #16 of 2024-06-14
  17. Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs 9 upvotes, #17 of 2024-06-14
  18. Cognitively Inspired Energy-Based World Models 9 upvotes, #17 of 2024-06-14
  19. 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities 8 upvotes, #19 of 2024-06-14
  20. Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? 7 upvotes, #20 of 2024-06-14
  21. TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation 7 upvotes, #20 of 2024-06-14
  22. Real3D: Scaling Up Large Reconstruction Models with Real-World Images 6 upvotes, #22 of 2024-06-14
  23. CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark 5 upvotes, #23 of 2024-06-14
  24. Language Model Council: Benchmarking Foundation Models on Highly Subjective Tasks by Consensus 5 upvotes, #23 of 2024-06-14
  25. Estimating the Hallucination Rate of Generative AI 4 upvotes, #25 of 2024-06-14
  26. MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding 4 upvotes, #25 of 2024-06-14
  27. Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation 4 upvotes, #25 of 2024-06-14
  28. CMC-Bench: Towards a New Paradigm of Visual Signal Compression 4 upvotes, #25 of 2024-06-14
  29. Understanding Hallucinations in Diffusion Models through Mode Interpolation 4 upvotes, #25 of 2024-06-14
  30. LRM-Zero: Training Large Reconstruction Models with Synthesized Data 3 upvotes, #30 of 2024-06-14

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.