Daily Papers of 2025-12-12
- T-pro 2.0: An Efficient Russian Hybrid-Reasoning Model and Playground 83 upvotes, #1 of 2025-12-12
- Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving 44 upvotes, #2 of 2025-12-12
- Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation 42 upvotes, #3 of 2025-12-12
- OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification 32 upvotes, #4 of 2025-12-12
- Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning 31 upvotes, #5 of 2025-12-12
- BEAVER: An Efficient Deterministic LLM Verifier 29 upvotes, #6 of 2025-12-12
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos 27 upvotes, #7 of 2025-12-12
- Thinking with Images via Self-Calling Agent 21 upvotes, #8 of 2025-12-12
- Stronger Normalization-Free Transformers 18 upvotes, #9 of 2025-12-12
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction 15 upvotes, #10 of 2025-12-12
- Evaluating Gemini Robotics Policies in a Veo World Simulator 15 upvotes, #10 of 2025-12-12
- From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models 15 upvotes, #10 of 2025-12-12
- StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space 12 upvotes, #13 of 2025-12-12
- X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale 6 upvotes, #14 of 2025-12-12
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization 6 upvotes, #14 of 2025-12-12
- Confucius Code Agent: An Open-sourced AI Software Engineer at Industrial Scale 5 upvotes, #16 of 2025-12-12
- The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality 5 upvotes, #16 of 2025-12-12
- MoRel: Long-Range Flicker-Free 4D Motion Modeling via Anchor Relay-based Bidirectional Blending with Hierarchical Densification 4 upvotes, #18 of 2025-12-12
- Fed-SE: Federated Self-Evolution for Privacy-Constrained Multi-Environment LLM Agents 3 upvotes, #19 of 2025-12-12
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos 3 upvotes, #19 of 2025-12-12
- MOA: Multi-Objective Alignment for Role-Playing Agents 3 upvotes, #19 of 2025-12-12
- ReViSE: Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning 3 upvotes, #19 of 2025-12-12
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task 3 upvotes, #19 of 2025-12-12
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance 2 upvotes, #24 of 2025-12-12
- DragMesh: Interactive 3D Generation Made Easy 1 upvotes, #25 of 2025-12-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.