Daily Papers of 2025-10-31

  1. The End of Manual Decoding: Towards Truly End-to-End Language Models 113 upvotes, #1 of 2025-10-31
  2. Emu3.5: Native Multimodal Models are World Learners 102 upvotes, #2 of 2025-10-31
  3. Kimi Linear: An Expressive, Efficient Attention Architecture 97 upvotes, #3 of 2025-10-31
  4. Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games 45 upvotes, #4 of 2025-10-31
  5. Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning 40 upvotes, #5 of 2025-10-31
  6. Exploring Conditions for Diffusion models in Robotic Control 39 upvotes, #6 of 2025-10-31
  7. Surfer 2: The Next Generation of Cross-Platform Computer Use Agents 36 upvotes, #7 of 2025-10-31
  8. AMO-Bench: Large Language Models Still Struggle in High School Math Competitions 33 upvotes, #8 of 2025-10-31
  9. Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark 32 upvotes, #9 of 2025-10-31
  10. The Quest for Generalizable Motion Generation: Data, Model, and Evaluation 27 upvotes, #10 of 2025-10-31
  11. The Era of Agentic Organization: Learning to Organize with Language Models 23 upvotes, #11 of 2025-10-31
  12. OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes 21 upvotes, #12 of 2025-10-31
  13. MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency 16 upvotes, #13 of 2025-10-31
  14. EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis 9 upvotes, #14 of 2025-10-31
  15. Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets 9 upvotes, #14 of 2025-10-31
  16. OmniLayout: Enabling Coarse-to-Fine Learning with LLMs for Universal Document Layout Generation 9 upvotes, #14 of 2025-10-31
  17. MedVLSynther: Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs 6 upvotes, #17 of 2025-10-31
  18. FullPart: Generating each 3D Part at Full Resolution 5 upvotes, #18 of 2025-10-31
  19. Remote Labor Index: Measuring AI Automation of Remote Work 5 upvotes, #18 of 2025-10-31
  20. CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs 4 upvotes, #20 of 2025-10-31
  21. PORTool: Tool-Use LLM Training with Rewarded Tree 4 upvotes, #20 of 2025-10-31
  22. CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark 4 upvotes, #20 of 2025-10-31
  23. L^2M^3OF: A Large Language Multimodal Model for Metal-Organic Frameworks 2 upvotes, #23 of 2025-10-31
  24. Performance Trade-offs of Optimizing Small Language Models for E-Commerce 2 upvotes, #23 of 2025-10-31
  25. CityRiSE: Reasoning Urban Socio-Economic Status in Vision-Language Models via Reinforcement Learning 2 upvotes, #23 of 2025-10-31
  26. POWSM: A Phonetic Open Whisper-Style Speech Foundation Model 2 upvotes, #23 of 2025-10-31
  27. EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation 2 upvotes, #23 of 2025-10-31
  28. Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing 2 upvotes, #23 of 2025-10-31
  29. ChartAB: A Benchmark for Chart Grounding & Dense Alignment 1 upvotes, #29 of 2025-10-31

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.