Daily Papers of 2024-06-18

  1. MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs 52 upvotes, #1 of 2024-06-18
  2. DataComp-LM: In search of the next generation of training sets for language models 33 upvotes, #2 of 2024-06-18
  3. mDPO: Conditional Preference Optimization for Multimodal Large Language Models 33 upvotes, #2 of 2024-06-18
  4. THEANINE: Revisiting Memory Management in Long-term Conversations with Timeline-augmented Response Generation 31 upvotes, #4 of 2024-06-18
  5. How Do Large Language Models Acquire Factual Knowledge During Pretraining? 25 upvotes, #5 of 2024-06-18
  6. MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers 21 upvotes, #6 of 2024-06-18
  7. A Simple and Effective L_2 Norm-Based Strategy for KV Cache Compression 21 upvotes, #6 of 2024-06-18
  8. VideoLLM-online: Online Video Large Language Model for Streaming Video 19 upvotes, #8 of 2024-06-18
  9. GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities 18 upvotes, #9 of 2024-06-18
  10. From Pixels to Prose: A Large Dataset of Dense Image Captions 16 upvotes, #10 of 2024-06-18
  11. Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models 16 upvotes, #10 of 2024-06-18
  12. LLaNA: Large Language and NeRF Assistant 16 upvotes, #10 of 2024-06-18
  13. In-Context Editing: Learning Knowledge from Self-Induced Distributions 14 upvotes, #13 of 2024-06-18
  14. WPO: Enhancing RLHF with Weighted Preference Optimization 13 upvotes, #14 of 2024-06-18
  15. Pandora: Towards General World Model with Natural Language Actions and Video States 12 upvotes, #15 of 2024-06-18
  16. L4GM: Large 4D Gaussian Reconstruction Model 11 upvotes, #16 of 2024-06-18
  17. WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences 11 upvotes, #16 of 2024-06-18
  18. MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 10 upvotes, #18 of 2024-06-18
  19. Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion 8 upvotes, #19 of 2024-06-18
  20. Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning 7 upvotes, #20 of 2024-06-18
  21. Task Me Anything 7 upvotes, #20 of 2024-06-18
  22. Unifying Multimodal Retrieval via Document Screenshot Embedding 6 upvotes, #22 of 2024-06-18
  23. Just How Flexible are Neural Networks in Practice? 6 upvotes, #22 of 2024-06-18
  24. Evaluating Open Language Models Across Task Types, Application Domains, and Reasoning Types: An In-Depth Experimental Analysis 5 upvotes, #24 of 2024-06-18
  25. CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training 4 upvotes, #25 of 2024-06-18
  26. HiddenTables & PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies 4 upvotes, #25 of 2024-06-18
  27. Breaking the Attention Bottleneck 4 upvotes, #25 of 2024-06-18
  28. Consistency^2: Consistent and Fast 3D Painting with Latent Consistency Models 3 upvotes, #28 of 2024-06-18
  29. Deep Bayesian Active Learning for Preference Modeling in Large Language Models 2 upvotes, #29 of 2024-06-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.