Daily Papers of 2024-04-09
- Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs 52 upvotes, #1 of 2024-04-09
- ByteEdit: Boost, Comply and Accelerate Generative Image Editing 22 upvotes, #2 of 2024-04-09
- SpatialTracker: Tracking Any 2D Pixels in 3D Space 21 upvotes, #3 of 2024-04-09
- SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing 21 upvotes, #3 of 2024-04-09
- BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion 20 upvotes, #5 of 2024-04-09
- UniFL: Improve Stable Diffusion via Unified Feedback Learning 20 upvotes, #5 of 2024-04-09
- MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators 19 upvotes, #7 of 2024-04-09
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 15 upvotes, #8 of 2024-04-09
- PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations 13 upvotes, #9 of 2024-04-09
- YaART: Yet Another ART Rendering Technology 13 upvotes, #9 of 2024-04-09
- Aligning Diffusion Models by Optimizing Human Utility 11 upvotes, #11 of 2024-04-09
- Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models 9 upvotes, #12 of 2024-04-09
- MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation 9 upvotes, #12 of 2024-04-09
- DATENeRF: Depth-Aware Text-based Editing of NeRFs 7 upvotes, #14 of 2024-04-09
- Koala: Key frame-conditioned long video-LLM 5 upvotes, #15 of 2024-04-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.