Daily Papers of 2024-06-18
- MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs 52 upvotes, #1 of 2024-06-18
- DataComp-LM: In search of the next generation of training sets for language models 33 upvotes, #2 of 2024-06-18
- mDPO: Conditional Preference Optimization for Multimodal Large Language Models 33 upvotes, #2 of 2024-06-18
- THEANINE: Revisiting Memory Management in Long-term Conversations with Timeline-augmented Response Generation 31 upvotes, #4 of 2024-06-18
- How Do Large Language Models Acquire Factual Knowledge During Pretraining? 25 upvotes, #5 of 2024-06-18
- MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers 21 upvotes, #6 of 2024-06-18
- A Simple and Effective L_2 Norm-Based Strategy for KV Cache Compression 21 upvotes, #6 of 2024-06-18
- VideoLLM-online: Online Video Large Language Model for Streaming Video 19 upvotes, #8 of 2024-06-18
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities 18 upvotes, #9 of 2024-06-18
- From Pixels to Prose: A Large Dataset of Dense Image Captions 16 upvotes, #10 of 2024-06-18
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models 16 upvotes, #10 of 2024-06-18
- LLaNA: Large Language and NeRF Assistant 16 upvotes, #10 of 2024-06-18
- In-Context Editing: Learning Knowledge from Self-Induced Distributions 14 upvotes, #13 of 2024-06-18
- WPO: Enhancing RLHF with Weighted Preference Optimization 13 upvotes, #14 of 2024-06-18
- Pandora: Towards General World Model with Natural Language Actions and Video States 12 upvotes, #15 of 2024-06-18
- L4GM: Large 4D Gaussian Reconstruction Model 11 upvotes, #16 of 2024-06-18
- WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences 11 upvotes, #16 of 2024-06-18
- MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 10 upvotes, #18 of 2024-06-18
- Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion 8 upvotes, #19 of 2024-06-18
- Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning 7 upvotes, #20 of 2024-06-18
- Task Me Anything 7 upvotes, #20 of 2024-06-18
- Unifying Multimodal Retrieval via Document Screenshot Embedding 6 upvotes, #22 of 2024-06-18
- Just How Flexible are Neural Networks in Practice? 6 upvotes, #22 of 2024-06-18
- Evaluating Open Language Models Across Task Types, Application Domains, and Reasoning Types: An In-Depth Experimental Analysis 5 upvotes, #24 of 2024-06-18
- CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training 4 upvotes, #25 of 2024-06-18
- HiddenTables & PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies 4 upvotes, #25 of 2024-06-18
- Breaking the Attention Bottleneck 4 upvotes, #25 of 2024-06-18
- Consistency^2: Consistent and Fast 3D Painting with Latent Consistency Models 3 upvotes, #28 of 2024-06-18
- Deep Bayesian Active Learning for Preference Modeling in Large Language Models 2 upvotes, #29 of 2024-06-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.