Daily Papers of 2025-03-19

  1. RWKV-7 "Goose" with Expressive Dynamic State Evolution 131 upvotes, #1 of 2025-03-19
  2. DAPO: An Open-Source LLM Reinforcement Learning System at Scale 104 upvotes, #2 of 2025-03-19
  3. Impossible Videos 52 upvotes, #3 of 2025-03-19
  4. Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM 41 upvotes, #4 of 2025-03-19
  5. DeepPerception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding 29 upvotes, #5 of 2025-03-19
  6. Infinite Mobility: Scalable High-Fidelity Synthesis of Articulated Objects via Procedural Generation 26 upvotes, #6 of 2025-03-19
  7. CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era 24 upvotes, #7 of 2025-03-19
  8. AudioX: Diffusion Transformer for Anything-to-Audio Generation 21 upvotes, #8 of 2025-03-19
  9. Aligning Multimodal LLM with Human Preference: A Survey 21 upvotes, #8 of 2025-03-19
  10. Frac-Connections: Fractional Extension of Hyper-Connections 19 upvotes, #10 of 2025-03-19
  11. Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control 16 upvotes, #11 of 2025-03-19
  12. FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis 15 upvotes, #12 of 2025-03-19
  13. Atlas: Multi-Scale Attention Improves Long Context Image Modeling 11 upvotes, #13 of 2025-03-19
  14. Concat-ID: Towards Universal Identity-Preserving Video Synthesis 10 upvotes, #14 of 2025-03-19
  15. Measuring AI Ability to Complete Long Tasks 10 upvotes, #14 of 2025-03-19
  16. Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection 9 upvotes, #16 of 2025-03-19
  17. MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification 9 upvotes, #16 of 2025-03-19
  18. Temporal Consistency for LLM Reasoning Process Error Identification 9 upvotes, #16 of 2025-03-19
  19. Florenz: Scaling Laws for Systematic Generalization in Vision-Language Models 7 upvotes, #19 of 2025-03-19
  20. MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs 7 upvotes, #19 of 2025-03-19
  21. Towards Self-Improving Systematic Cognition for Next-Generation Foundation MLLMs 6 upvotes, #21 of 2025-03-19
  22. EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees 5 upvotes, #22 of 2025-03-19
  23. PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models 5 upvotes, #22 of 2025-03-19
  24. Pensez: Less Data, Better Reasoning -- Rethinking French LLM 5 upvotes, #22 of 2025-03-19
  25. PyGDA: A Python Library for Graph Domain Adaptation 4 upvotes, #25 of 2025-03-19
  26. Learning to Inference Adaptively for Multimodal Large Language Models 4 upvotes, #25 of 2025-03-19
  27. RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground Simulation 3 upvotes, #27 of 2025-03-19
  28. KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation 3 upvotes, #27 of 2025-03-19
  29. Hyperbolic Safety-Aware Vision-Language Models 3 upvotes, #27 of 2025-03-19
  30. MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain Specific Generative Modeling 3 upvotes, #27 of 2025-03-19
  31. CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous Driving 1 upvotes, #31 of 2025-03-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.