Daily Papers of 2024-11-27

  1. ShowUI: One Vision-Language-Action Model for GUI Visual Agent 68 upvotes, #1 of 2024-11-27
  2. Star Attention: Efficient LLM Inference over Long Sequences 42 upvotes, #2 of 2024-11-27
  3. Pathways on the Image Manifold: Image Editing via Video Generation 29 upvotes, #3 of 2024-11-27
  4. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
  5. Rethinking Token Reduction in MLLMs: Towards a Unified Paradigm for Training-Free Acceleration 18 upvotes, #4 of 2024-11-27
  6. SketchAgent: Language-Driven Sequential Sketch Generation 14 upvotes, #6 of 2024-11-27
  7. TEXGen: a Generative Diffusion Model for Mesh Textures 13 upvotes, #7 of 2024-11-27
  8. VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models 10 upvotes, #8 of 2024-11-27
  9. Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens 9 upvotes, #9 of 2024-11-27
  10. SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE 8 upvotes, #10 of 2024-11-27
  11. Learning 3D Representations from Procedural 3D Programs 8 upvotes, #10 of 2024-11-27
  12. FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity 7 upvotes, #12 of 2024-11-27
  13. EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality 5 upvotes, #13 of 2024-11-27
  14. SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis 5 upvotes, #13 of 2024-11-27
  15. DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting 5 upvotes, #13 of 2024-11-27
  16. AnchorCrafter: Animate CyberAnchors Saling Your Products via Human-Object Interacting Video Generation 5 upvotes, #13 of 2024-11-27
  17. MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts 3 upvotes, #17 of 2024-11-27
  18. Controllable Human Image Generation with Personalized Multi-Garments 3 upvotes, #17 of 2024-11-27
  19. Visual Counter Turing Test (VCT^2): Discovering the Challenges for AI-Generated Image Detection and Introducing Visual AI Index (V_AI) 1 upvotes, #19 of 2024-11-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.