Daily Papers of 2024-12-06
- VisionZip: Longer is Better but Not Necessary in Vision Language Models 96 upvotes, #1 of 2024-12-06
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
- NVILA: Efficient Frontier Visual Language Models 47 upvotes, #3 of 2024-12-06
- Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction 43 upvotes, #4 of 2024-12-06
- Evaluating Language Models as Synthetic Data Generators 39 upvotes, #5 of 2024-12-06
- Structured 3D Latents for Scalable and Versatile 3D Generation 36 upvotes, #6 of 2024-12-06
- Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection 32 upvotes, #7 of 2024-12-06
- A Noise is Worth Diffusion Guidance 26 upvotes, #8 of 2024-12-06
- Negative Token Merging: Image-based Adversarial Feature Guidance 21 upvotes, #9 of 2024-12-06
- MV-Adapter: Multi-view Consistent Image Generation Made Easy 20 upvotes, #10 of 2024-12-06
- AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models 20 upvotes, #10 of 2024-12-06
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation 16 upvotes, #12 of 2024-12-06
- Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis 15 upvotes, #13 of 2024-12-06
- Densing Law of LLMs 14 upvotes, #14 of 2024-12-06
- HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing 13 upvotes, #15 of 2024-12-06
- Personalized Multimodal Large Language Models: A Survey 12 upvotes, #16 of 2024-12-06
- OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows 10 upvotes, #17 of 2024-12-06
- Monet: Mixture of Monosemantic Experts for Transformers 10 upvotes, #17 of 2024-12-06
- Discriminative Fine-tuning of LVLMs 10 upvotes, #17 of 2024-12-06
- Towards Universal Soccer Video Understanding 9 upvotes, #20 of 2024-12-06
- Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement 9 upvotes, #20 of 2024-12-06
- KV Shifting Attention Enhances Language Modeling 8 upvotes, #22 of 2024-12-06
- MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation 8 upvotes, #22 of 2024-12-06
- ZipAR: Accelerating Autoregressive Image Generation through Spatial Locality 7 upvotes, #24 of 2024-12-06
- 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion 7 upvotes, #24 of 2024-12-06
- Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension 6 upvotes, #26 of 2024-12-06
- p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay 6 upvotes, #26 of 2024-12-06
- MRGen: Diffusion-based Controllable Data Engine for MRI Segmentation towards Unannotated Modalities 5 upvotes, #28 of 2024-12-06
- SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction 3 upvotes, #29 of 2024-12-06
- Challenges in Trustworthy Human Evaluation of Chatbots 2 upvotes, #30 of 2024-12-06
- Establishing Task Scaling Laws via Compute-Efficient Model Ladders 2 upvotes, #30 of 2024-12-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.