Daily Papers of 2025-01-22

  1. Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training 84 upvotes, #1 of 2025-01-22
  2. MMVU: Measuring Expert-Level Multi-Discipline Video Understanding 79 upvotes, #2 of 2025-01-22
  3. Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models 61 upvotes, #3 of 2025-01-22
  4. UI-TARS: Pioneering Automated GUI Interaction with Native Agents 47 upvotes, #4 of 2025-01-22
  5. TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space 46 upvotes, #5 of 2025-01-22
  6. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 39 upvotes, #6 of 2025-01-22
  7. Reasoning Language Models: A Blueprint 30 upvotes, #7 of 2025-01-22
  8. Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation 28 upvotes, #8 of 2025-01-22
  9. Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks 26 upvotes, #9 of 2025-01-22
  10. Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments 22 upvotes, #10 of 2025-01-22
  11. Video Depth Anything: Consistent Depth Estimation for Super-Long Videos 21 upvotes, #11 of 2025-01-22
  12. Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise 20 upvotes, #12 of 2025-01-22
  13. Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement 14 upvotes, #13 of 2025-01-22
  14. EMO2: End-Effector Guided Audio-Driven Avatar Video Generation 12 upvotes, #14 of 2025-01-22
  15. GPS as a Control Signal for Image Generation 12 upvotes, #14 of 2025-01-22
  16. Taming Teacher Forcing for Masked Autoregressive Video Generation 10 upvotes, #16 of 2025-01-22
  17. MSTS: A Multimodal Safety Test Suite for Vision-Language Models 8 upvotes, #17 of 2025-01-22
  18. The Geometry of Tokens in Internal Representations of Large Language Models 8 upvotes, #17 of 2025-01-22
  19. Panoramic Interests: Stylistic-Content Aware Personalized Headline Generation 4 upvotes, #19 of 2025-01-22
  20. Fixing Imbalanced Attention to Mitigate In-Context Hallucination of Large Vision-Language Model 4 upvotes, #19 of 2025-01-22

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.