Daily Papers of 2024-04-09

  1. Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs 52 upvotes, #1 of 2024-04-09
  2. ByteEdit: Boost, Comply and Accelerate Generative Image Editing 22 upvotes, #2 of 2024-04-09
  3. SpatialTracker: Tracking Any 2D Pixels in 3D Space 21 upvotes, #3 of 2024-04-09
  4. SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing 21 upvotes, #3 of 2024-04-09
  5. BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion 20 upvotes, #5 of 2024-04-09
  6. UniFL: Improve Stable Diffusion via Unified Feedback Learning 20 upvotes, #5 of 2024-04-09
  7. MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators 19 upvotes, #7 of 2024-04-09
  8. MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 15 upvotes, #8 of 2024-04-09
  9. PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations 13 upvotes, #9 of 2024-04-09
  10. YaART: Yet Another ART Rendering Technology 13 upvotes, #9 of 2024-04-09
  11. Aligning Diffusion Models by Optimizing Human Utility 11 upvotes, #11 of 2024-04-09
  12. Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models 9 upvotes, #12 of 2024-04-09
  13. MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation 9 upvotes, #12 of 2024-04-09
  14. DATENeRF: Depth-Aware Text-based Editing of NeRFs 7 upvotes, #14 of 2024-04-09
  15. Koala: Key frame-conditioned long video-LLM 5 upvotes, #15 of 2024-04-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.