Daily Papers of 2024-12-06

  1. VisionZip: Longer is Better but Not Necessary in Vision Language Models 96 upvotes, #1 of 2024-12-06
  2. Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
  3. NVILA: Efficient Frontier Visual Language Models 47 upvotes, #3 of 2024-12-06
  4. Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction 43 upvotes, #4 of 2024-12-06
  5. Evaluating Language Models as Synthetic Data Generators 39 upvotes, #5 of 2024-12-06
  6. Structured 3D Latents for Scalable and Versatile 3D Generation 36 upvotes, #6 of 2024-12-06
  7. Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection 32 upvotes, #7 of 2024-12-06
  8. A Noise is Worth Diffusion Guidance 26 upvotes, #8 of 2024-12-06
  9. Negative Token Merging: Image-based Adversarial Feature Guidance 21 upvotes, #9 of 2024-12-06
  10. MV-Adapter: Multi-view Consistent Image Generation Made Easy 20 upvotes, #10 of 2024-12-06
  11. AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models 20 upvotes, #10 of 2024-12-06
  12. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation 16 upvotes, #12 of 2024-12-06
  13. Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis 15 upvotes, #13 of 2024-12-06
  14. Densing Law of LLMs 14 upvotes, #14 of 2024-12-06
  15. HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing 13 upvotes, #15 of 2024-12-06
  16. Personalized Multimodal Large Language Models: A Survey 12 upvotes, #16 of 2024-12-06
  17. OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows 10 upvotes, #17 of 2024-12-06
  18. Monet: Mixture of Monosemantic Experts for Transformers 10 upvotes, #17 of 2024-12-06
  19. Discriminative Fine-tuning of LVLMs 10 upvotes, #17 of 2024-12-06
  20. Towards Universal Soccer Video Understanding 9 upvotes, #20 of 2024-12-06
  21. Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement 9 upvotes, #20 of 2024-12-06
  22. KV Shifting Attention Enhances Language Modeling 8 upvotes, #22 of 2024-12-06
  23. MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation 8 upvotes, #22 of 2024-12-06
  24. ZipAR: Accelerating Autoregressive Image Generation through Spatial Locality 7 upvotes, #24 of 2024-12-06
  25. 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion 7 upvotes, #24 of 2024-12-06
  26. Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension 6 upvotes, #26 of 2024-12-06
  27. p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay 6 upvotes, #26 of 2024-12-06
  28. MRGen: Diffusion-based Controllable Data Engine for MRI Segmentation towards Unannotated Modalities 5 upvotes, #28 of 2024-12-06
  29. SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction 3 upvotes, #29 of 2024-12-06
  30. Challenges in Trustworthy Human Evaluation of Chatbots 2 upvotes, #30 of 2024-12-06
  31. Establishing Task Scaling Laws via Compute-Efficient Model Ladders 2 upvotes, #30 of 2024-12-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.