Daily Papers of 2024-12-03
- X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models 61 upvotes, #1 of 2024-12-03
- o1-Coder: an o1 Replication for Coding 34 upvotes, #2 of 2024-12-03
- Open-Sora Plan: Open-Source Large Video Generation Model 30 upvotes, #3 of 2024-12-03
- Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis 27 upvotes, #4 of 2024-12-03
- VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation 24 upvotes, #5 of 2024-12-03
- SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters 21 upvotes, #6 of 2024-12-03
- FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait 21 upvotes, #6 of 2024-12-03
- TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video 19 upvotes, #8 of 2024-12-03
- GATE OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation 17 upvotes, #9 of 2024-12-03
- Steering Rectified Flow Models in the Vector Field for Controlled Image Generation 15 upvotes, #10 of 2024-12-03
- Efficient Track Anything 14 upvotes, #11 of 2024-12-03
- The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning 13 upvotes, #12 of 2024-12-03
- TinyFusion: Diffusion Transformers Learned Shallow 13 upvotes, #12 of 2024-12-03
- VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models 13 upvotes, #12 of 2024-12-03
- WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model 10 upvotes, #15 of 2024-12-03
- INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge 9 upvotes, #16 of 2024-12-03
- VLSBench: Unveiling Visual Leakage in Multimodal Safety 8 upvotes, #17 of 2024-12-03
- Art-Free Generative Models: Art Creation Without Graphic Art Knowledge 8 upvotes, #17 of 2024-12-03
- Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation 8 upvotes, #17 of 2024-12-03
- VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information 7 upvotes, #20 of 2024-12-03
- PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos 6 upvotes, #21 of 2024-12-03
- A Simple and Provable Scaling Law for the Test-Time Compute of Large Language Models 4 upvotes, #22 of 2024-12-03
- Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting 4 upvotes, #22 of 2024-12-03
- World-consistent Video Diffusion with Explicit 3D Modeling 4 upvotes, #22 of 2024-12-03
- Collaborative Instance Navigation: Leveraging Agent Self-Dialogue to Minimize User Input 3 upvotes, #25 of 2024-12-03
- AMO Sampler: Enhancing Text Rendering with Overshooting 2 upvotes, #26 of 2024-12-03
- Improving speaker verification robustness with synthetic emotional utterances 2 upvotes, #26 of 2024-12-03
- HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving 2 upvotes, #26 of 2024-12-03
- Towards Cross-Lingual Audio Abuse Detection in Low-Resource Settings with Few-Shot Learning 1 upvotes, #29 of 2024-12-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.