Daily Papers of 2024-06-14
- Depth Anything V2 83 upvotes, #1 of 2024-06-14
- An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels 46 upvotes, #2 of 2024-06-14
- Transformers meet Neural Algorithmic Reasoners 41 upvotes, #3 of 2024-06-14
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling 33 upvotes, #4 of 2024-06-14
- OpenVLA: An Open-Source Vision-Language-Action Model 28 upvotes, #5 of 2024-06-14
- Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models 27 upvotes, #6 of 2024-06-14
- Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning 22 upvotes, #7 of 2024-06-14
- DiTFastAttn: Attention Compression for Diffusion Transformer Models 19 upvotes, #8 of 2024-06-14
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models 17 upvotes, #9 of 2024-06-14
- MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding 17 upvotes, #9 of 2024-06-14
- Interpreting the Weight Space of Customized Diffusion Models 17 upvotes, #9 of 2024-06-14
- CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery 14 upvotes, #12 of 2024-06-14
- mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus 14 upvotes, #12 of 2024-06-14
- HelpSteer2: Open-source dataset for training top-performing reward models 13 upvotes, #14 of 2024-06-14
- EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts 12 upvotes, #15 of 2024-06-14
- Explore the Limits of Omni-modal Pretraining at Scale 10 upvotes, #16 of 2024-06-14
- Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs 9 upvotes, #17 of 2024-06-14
- Cognitively Inspired Energy-Based World Models 9 upvotes, #17 of 2024-06-14
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities 8 upvotes, #19 of 2024-06-14
- Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? 7 upvotes, #20 of 2024-06-14
- TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation 7 upvotes, #20 of 2024-06-14
- Real3D: Scaling Up Large Reconstruction Models with Real-World Images 6 upvotes, #22 of 2024-06-14
- CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark 5 upvotes, #23 of 2024-06-14
- Language Model Council: Benchmarking Foundation Models on Highly Subjective Tasks by Consensus 5 upvotes, #23 of 2024-06-14
- Estimating the Hallucination Rate of Generative AI 4 upvotes, #25 of 2024-06-14
- MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding 4 upvotes, #25 of 2024-06-14
- Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation 4 upvotes, #25 of 2024-06-14
- CMC-Bench: Towards a New Paradigm of Visual Signal Compression 4 upvotes, #25 of 2024-06-14
- Understanding Hallucinations in Diffusion Models through Mode Interpolation 4 upvotes, #25 of 2024-06-14
- LRM-Zero: Training Large Reconstruction Models with Synthesized Data 3 upvotes, #30 of 2024-06-14
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.