Daily Papers of 2025-10-27

  1. DeepAgent: A General Reasoning Agent with Scalable Toolsets 92 upvotes, #1 of 2025-10-27
  2. Reasoning with Sampling: Your Base Model is Smarter Than You Think 44 upvotes, #2 of 2025-10-27
  3. Video-As-Prompt: Unified Semantic Control for Video Generation 44 upvotes, #2 of 2025-10-27
  4. WorldGrow: Generating Infinite 3D World 40 upvotes, #4 of 2025-10-27
  5. A Definition of AGI 33 upvotes, #5 of 2025-10-27
  6. Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation 30 upvotes, #6 of 2025-10-27
  7. From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model 29 upvotes, #7 of 2025-10-27
  8. Sparser Block-Sparse Attention via Token Permutation 23 upvotes, #8 of 2025-10-27
  9. UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning 22 upvotes, #9 of 2025-10-27
  10. Visual Diffusion Models are Geometric Solvers 18 upvotes, #10 of 2025-10-27
  11. Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs 12 upvotes, #11 of 2025-10-27
  12. RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling 11 upvotes, #12 of 2025-10-27
  13. RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging 11 upvotes, #12 of 2025-10-27
  14. Model Merging with Functional Dual Anchors 11 upvotes, #12 of 2025-10-27
  15. ARC-Encoder: learning compressed text representations for large language models 5 upvotes, #15 of 2025-10-27
  16. Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost 4 upvotes, #16 of 2025-10-27
  17. Redefining Retrieval Evaluation in the Era of LLMs 4 upvotes, #16 of 2025-10-27
  18. PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis 4 upvotes, #16 of 2025-10-27
  19. Document Understanding, Measurement, and Manipulation Using Category Theory 4 upvotes, #16 of 2025-10-27
  20. Taming Modality Entanglement in Continual Audio-Visual Segmentation 3 upvotes, #20 of 2025-10-27
  21. Soft Instruction De-escalation Defense 3 upvotes, #20 of 2025-10-27
  22. Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video 3 upvotes, #20 of 2025-10-27
  23. AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite 3 upvotes, #20 of 2025-10-27
  24. Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers 2 upvotes, #24 of 2025-10-27
  25. PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments 2 upvotes, #24 of 2025-10-27
  26. ALICE-LRI: A General Method for Lossless Range Image Generation for Spinning LiDAR Sensors without Calibration Metadata 1 upvotes, #26 of 2025-10-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.