Daily Papers of 2024-12-03

  1. X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models 61 upvotes, #1 of 2024-12-03
  2. o1-Coder: an o1 Replication for Coding 34 upvotes, #2 of 2024-12-03
  3. Open-Sora Plan: Open-Source Large Video Generation Model 30 upvotes, #3 of 2024-12-03
  4. Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis 27 upvotes, #4 of 2024-12-03
  5. VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation 24 upvotes, #5 of 2024-12-03
  6. SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters 21 upvotes, #6 of 2024-12-03
  7. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait 21 upvotes, #6 of 2024-12-03
  8. TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video 19 upvotes, #8 of 2024-12-03
  9. GATE OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation 17 upvotes, #9 of 2024-12-03
  10. Steering Rectified Flow Models in the Vector Field for Controlled Image Generation 15 upvotes, #10 of 2024-12-03
  11. Efficient Track Anything 14 upvotes, #11 of 2024-12-03
  12. The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning 13 upvotes, #12 of 2024-12-03
  13. TinyFusion: Diffusion Transformers Learned Shallow 13 upvotes, #12 of 2024-12-03
  14. VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models 13 upvotes, #12 of 2024-12-03
  15. WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model 10 upvotes, #15 of 2024-12-03
  16. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge 9 upvotes, #16 of 2024-12-03
  17. VLSBench: Unveiling Visual Leakage in Multimodal Safety 8 upvotes, #17 of 2024-12-03
  18. Art-Free Generative Models: Art Creation Without Graphic Art Knowledge 8 upvotes, #17 of 2024-12-03
  19. Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation 8 upvotes, #17 of 2024-12-03
  20. VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information 7 upvotes, #20 of 2024-12-03
  21. PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos 6 upvotes, #21 of 2024-12-03
  22. A Simple and Provable Scaling Law for the Test-Time Compute of Large Language Models 4 upvotes, #22 of 2024-12-03
  23. Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting 4 upvotes, #22 of 2024-12-03
  24. World-consistent Video Diffusion with Explicit 3D Modeling 4 upvotes, #22 of 2024-12-03
  25. Collaborative Instance Navigation: Leveraging Agent Self-Dialogue to Minimize User Input 3 upvotes, #25 of 2024-12-03
  26. AMO Sampler: Enhancing Text Rendering with Overshooting 2 upvotes, #26 of 2024-12-03
  27. Improving speaker verification robustness with synthetic emotional utterances 2 upvotes, #26 of 2024-12-03
  28. HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving 2 upvotes, #26 of 2024-12-03
  29. Towards Cross-Lingual Audio Abuse Detection in Low-Resource Settings with Few-Shot Learning 1 upvotes, #29 of 2024-12-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.