Daily Papers of 2025-03-19
- RWKV-7 "Goose" with Expressive Dynamic State Evolution 131 upvotes, #1 of 2025-03-19
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale 104 upvotes, #2 of 2025-03-19
- Impossible Videos 52 upvotes, #3 of 2025-03-19
- Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM 41 upvotes, #4 of 2025-03-19
- DeepPerception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding 29 upvotes, #5 of 2025-03-19
- Infinite Mobility: Scalable High-Fidelity Synthesis of Articulated Objects via Procedural Generation 26 upvotes, #6 of 2025-03-19
- CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era 24 upvotes, #7 of 2025-03-19
- AudioX: Diffusion Transformer for Anything-to-Audio Generation 21 upvotes, #8 of 2025-03-19
- Aligning Multimodal LLM with Human Preference: A Survey 21 upvotes, #8 of 2025-03-19
- Frac-Connections: Fractional Extension of Hyper-Connections 19 upvotes, #10 of 2025-03-19
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control 16 upvotes, #11 of 2025-03-19
- FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis 15 upvotes, #12 of 2025-03-19
- Atlas: Multi-Scale Attention Improves Long Context Image Modeling 11 upvotes, #13 of 2025-03-19
- Concat-ID: Towards Universal Identity-Preserving Video Synthesis 10 upvotes, #14 of 2025-03-19
- Measuring AI Ability to Complete Long Tasks 10 upvotes, #14 of 2025-03-19
- Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection 9 upvotes, #16 of 2025-03-19
- MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification 9 upvotes, #16 of 2025-03-19
- Temporal Consistency for LLM Reasoning Process Error Identification 9 upvotes, #16 of 2025-03-19
- Florenz: Scaling Laws for Systematic Generalization in Vision-Language Models 7 upvotes, #19 of 2025-03-19
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs 7 upvotes, #19 of 2025-03-19
- Towards Self-Improving Systematic Cognition for Next-Generation Foundation MLLMs 6 upvotes, #21 of 2025-03-19
- EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees 5 upvotes, #22 of 2025-03-19
- PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models 5 upvotes, #22 of 2025-03-19
- Pensez: Less Data, Better Reasoning -- Rethinking French LLM 5 upvotes, #22 of 2025-03-19
- PyGDA: A Python Library for Graph Domain Adaptation 4 upvotes, #25 of 2025-03-19
- Learning to Inference Adaptively for Multimodal Large Language Models 4 upvotes, #25 of 2025-03-19
- RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground Simulation 3 upvotes, #27 of 2025-03-19
- KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation 3 upvotes, #27 of 2025-03-19
- Hyperbolic Safety-Aware Vision-Language Models 3 upvotes, #27 of 2025-03-19
- MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain Specific Generative Modeling 3 upvotes, #27 of 2025-03-19
- CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous Driving 1 upvotes, #31 of 2025-03-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.