Daily Papers of 2025-12-22

  1. Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows 106 upvotes, #1 of 2025-12-22
  2. PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence 72 upvotes, #2 of 2025-12-22
  3. Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding 63 upvotes, #3 of 2025-12-22
  4. When Reasoning Meets Its Laws 54 upvotes, #4 of 2025-12-22
  5. Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience 48 upvotes, #5 of 2025-12-22
  6. 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation 42 upvotes, #6 of 2025-12-22
  7. Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing 36 upvotes, #7 of 2025-12-22
  8. Are We on the Right Way to Assessing LLM-as-a-Judge? 32 upvotes, #8 of 2025-12-22
  9. Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers 22 upvotes, #9 of 2025-12-22
  10. An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges 21 upvotes, #10 of 2025-12-22
  11. GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation 19 upvotes, #11 of 2025-12-22
  12. RadarGen: Automotive Radar Point Cloud Generation from Cameras 17 upvotes, #12 of 2025-12-22
  13. HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering 12 upvotes, #13 of 2025-12-22
  14. Bolmo: Byteifying the Next Generation of Language Models 11 upvotes, #14 of 2025-12-22
  15. 3D-RE-GEN: 3D Reconstruction of Indoor Scenes with a Generative Framework 11 upvotes, #14 of 2025-12-22
  16. Meta-RL Induces Exploration in Language Agents 10 upvotes, #16 of 2025-12-22
  17. Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs 10 upvotes, #16 of 2025-12-22
  18. Animate Any Character in Any World 10 upvotes, #16 of 2025-12-22
  19. SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories 9 upvotes, #19 of 2025-12-22
  20. StageVAR: Stage-Aware Acceleration for Visual Autoregressive Models 7 upvotes, #20 of 2025-12-22
  21. A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos 4 upvotes, #21 of 2025-12-22
  22. MineTheGap: Automatic Mining of Biases in Text-to-Image Models 2 upvotes, #22 of 2025-12-22

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.