Daily Papers of 2025-01-08
- REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models 79 upvotes, #1 of 2025-01-08
- Cosmos World Foundation Model Platform for Physical AI 61 upvotes, #2 of 2025-01-08
- LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token 48 upvotes, #3 of 2025-01-08
- MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models 40 upvotes, #4 of 2025-01-08
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos 40 upvotes, #4 of 2025-01-08
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control 22 upvotes, #6 of 2025-01-08
- PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides 18 upvotes, #7 of 2025-01-08
- OpenOmni: Large Language Models Pivot Zero-shot Omnimodal Alignment across Language with Real-time Self-Aware Emotional Speech Synthesis 16 upvotes, #8 of 2025-01-08
- Dolphin: Closed-loop Open-ended Auto-research through Thinking, Practice, and Feedback 14 upvotes, #9 of 2025-01-08
- Magic Mirror: ID-Preserved Video Generation in Video Diffusion Transformers 14 upvotes, #9 of 2025-01-08
- Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model 9 upvotes, #11 of 2025-01-08
- MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splatting 9 upvotes, #11 of 2025-01-08
- Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers 8 upvotes, #13 of 2025-01-08
- MagicFace: High-Fidelity Facial Expression Editing with Action-Unit Control 5 upvotes, #14 of 2025-01-08
- Generalizable Origin Identification for Text-Guided Image-to-Image Diffusion Models 3 upvotes, #15 of 2025-01-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.