Daily Papers of 2025-12-16

  1. Memory in the Age of AI Agents 106 upvotes, #1 of 2025-12-16
  2. QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management 98 upvotes, #2 of 2025-12-16
  3. Towards Scalable Pre-training of Visual Tokenizers for Generation 91 upvotes, #3 of 2025-12-16
  4. ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding 87 upvotes, #4 of 2025-12-16
  5. LongVie 2: Multimodal Controllable Ultra-Long Video World Model 69 upvotes, #5 of 2025-12-16
  6. Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows 49 upvotes, #6 of 2025-12-16
  7. NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents 42 upvotes, #7 of 2025-12-16
  8. KlingAvatar 2.0 Technical Report 40 upvotes, #8 of 2025-12-16
  9. Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics 39 upvotes, #9 of 2025-12-16
  10. MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment 25 upvotes, #10 of 2025-12-16
  11. Openpi Comet: Competition Solution For 2025 BEHAVIOR Challenge 17 upvotes, #11 of 2025-12-16
  12. Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos 15 upvotes, #12 of 2025-12-16
  13. WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment 13 upvotes, #13 of 2025-12-16
  14. Towards Interactive Intelligence for Digital Humans 11 upvotes, #14 of 2025-12-16
  15. DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning 10 upvotes, #15 of 2025-12-16
  16. V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions 9 upvotes, #16 of 2025-12-16
  17. Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection 9 upvotes, #16 of 2025-12-16
  18. CAPTAIN: Semantic Feature Injection for Memorization Mitigation in Text-to-Image Diffusion Models 8 upvotes, #18 of 2025-12-16
  19. What matters for Representation Alignment: Global Information or Spatial Structure? 8 upvotes, #18 of 2025-12-16
  20. VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer 8 upvotes, #18 of 2025-12-16
  21. Rethinking Expert Trajectory Utilization in LLM Post-training 7 upvotes, #21 of 2025-12-16
  22. Few-Step Distillation for Text-to-Image Generation: A Practical Guide 7 upvotes, #21 of 2025-12-16
  23. Image Diffusion Preview with Consistency Solver 7 upvotes, #21 of 2025-12-16
  24. Aesthetic Alignment Risks Assimilation: How Image Generation and Reward Models Reinforce Beauty Bias and Ideological "Censorship" 6 upvotes, #24 of 2025-12-16
  25. GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation 6 upvotes, #24 of 2025-12-16
  26. LitePT: Lighter Yet Stronger Point Transformer 6 upvotes, #24 of 2025-12-16
  27. AutoMV: An Automatic Multi-Agent System for Music Video Generation 5 upvotes, #27 of 2025-12-16
  28. RecTok: Reconstruction Distillation along Rectified Flow 4 upvotes, #28 of 2025-12-16
  29. FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos 3 upvotes, #29 of 2025-12-16
  30. Flowception: Temporally Expansive Flow Matching for Video Generation 3 upvotes, #29 of 2025-12-16
  31. State over Tokens: Characterizing the Role of Reasoning Tokens 3 upvotes, #29 of 2025-12-16
  32. FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models 3 upvotes, #29 of 2025-12-16
  33. I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners 3 upvotes, #29 of 2025-12-16
  34. Inferring Compositional 4D Scenes without Ever Seeing One 2 upvotes, #34 of 2025-12-16
  35. START: Spatial and Textual Learning for Chart Understanding 2 upvotes, #34 of 2025-12-16
  36. CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence 2 upvotes, #34 of 2025-12-16
  37. Directional Textual Inversion for Personalized Text-to-Image Generation 2 upvotes, #34 of 2025-12-16
  38. DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders 2 upvotes, #34 of 2025-12-16
  39. Towards Visual Re-Identification of Fish using Fine-Grained Classification for Electronic Monitoring in Fisheries 1 upvotes, #39 of 2025-12-16
  40. Learning Robot Manipulation from Audio World Models 1 upvotes, #39 of 2025-12-16
  41. KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification 1 upvotes, #41 of 2025-12-16

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.