Daily Papers of 2024-11-27
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent 68 upvotes, #1 of 2024-11-27
- Star Attention: Efficient LLM Inference over Long Sequences 42 upvotes, #2 of 2024-11-27
- Pathways on the Image Manifold: Image Editing via Video Generation 29 upvotes, #3 of 2024-11-27
- MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
- Rethinking Token Reduction in MLLMs: Towards a Unified Paradigm for Training-Free Acceleration 18 upvotes, #4 of 2024-11-27
- SketchAgent: Language-Driven Sequential Sketch Generation 14 upvotes, #6 of 2024-11-27
- TEXGen: a Generative Diffusion Model for Mesh Textures 13 upvotes, #7 of 2024-11-27
- VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models 10 upvotes, #8 of 2024-11-27
- Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens 9 upvotes, #9 of 2024-11-27
- SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE 8 upvotes, #10 of 2024-11-27
- Learning 3D Representations from Procedural 3D Programs 8 upvotes, #10 of 2024-11-27
- FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity 7 upvotes, #12 of 2024-11-27
- EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality 5 upvotes, #13 of 2024-11-27
- SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis 5 upvotes, #13 of 2024-11-27
- DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting 5 upvotes, #13 of 2024-11-27
- AnchorCrafter: Animate CyberAnchors Saling Your Products via Human-Object Interacting Video Generation 5 upvotes, #13 of 2024-11-27
- MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts 3 upvotes, #17 of 2024-11-27
- Controllable Human Image Generation with Personalized Multi-Garments 3 upvotes, #17 of 2024-11-27
- Visual Counter Turing Test (VCT^2): Discovering the Challenges for AI-Generated Image Detection and Introducing Visual AI Index (V_AI) 1 upvotes, #19 of 2024-11-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.