Daily Papers of 2026-01-27

  1. Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs 181 upvotes, #1 of 2026-01-27
  2. daVinci-Dev: Agent-native Mid-training for Software Engineering 123 upvotes, #2 of 2026-01-27
  3. The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation 55 upvotes, #3 of 2026-01-27
  4. Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility 41 upvotes, #4 of 2026-01-27
  5. Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability 39 upvotes, #5 of 2026-01-27
  6. Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers 33 upvotes, #6 of 2026-01-27
  7. iFSQ: Improving FSQ for Image Generation with 1 Line of Code 31 upvotes, #7 of 2026-01-27
  8. DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints 24 upvotes, #8 of 2026-01-27
  9. Self-Refining Video Sampling 24 upvotes, #8 of 2026-01-27
  10. Masked Depth Modeling for Spatial Perception 22 upvotes, #10 of 2026-01-27
  11. VIBEVOICE-ASR Technical Report 19 upvotes, #11 of 2026-01-27
  12. Agentic Very Long Video Understanding 17 upvotes, #12 of 2026-01-27
  13. CGPT: Cluster-Guided Partial Tables with LLM-Generated Supervision for Table Retrieval 14 upvotes, #13 of 2026-01-27
  14. AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation 14 upvotes, #13 of 2026-01-27
  15. TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models 10 upvotes, #15 of 2026-01-27
  16. STAR: Semantic Table Representation with Header-Aware Clustering and Adaptive Weighted Fusion 9 upvotes, #16 of 2026-01-27
  17. A Mechanistic View on Video Generation as World Models: State and Dynamics 9 upvotes, #16 of 2026-01-27
  18. SkyReels-V3 Technique Report 9 upvotes, #16 of 2026-01-27
  19. Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents 9 upvotes, #16 of 2026-01-27
  20. C-RADIOv4 (Tech Report) 8 upvotes, #20 of 2026-01-27
  21. SAGE: Steerable Agentic Data Generation for Deep Search with Execution Feedback 8 upvotes, #20 of 2026-01-27
  22. Diffusion In Diffusion: Reclaiming Global Coherence in Semi-Autoregressive Diffusion 7 upvotes, #22 of 2026-01-27
  23. IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance 7 upvotes, #22 of 2026-01-27
  24. DRPG (Decompose, Retrieve, Plan, Generate): An Agentic Framework for Academic Rebuttal 7 upvotes, #22 of 2026-01-27
  25. Yunjue Agent Tech Report: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks 7 upvotes, #22 of 2026-01-27
  26. One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment 7 upvotes, #22 of 2026-01-27
  27. Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts 5 upvotes, #27 of 2026-01-27
  28. PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues 5 upvotes, #27 of 2026-01-27
  29. End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions 5 upvotes, #27 of 2026-01-27
  30. Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction 4 upvotes, #30 of 2026-01-27
  31. TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors 3 upvotes, #31 of 2026-01-27
  32. Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models 2 upvotes, #32 of 2026-01-27
  33. The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning 2 upvotes, #32 of 2026-01-27
  34. Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control 2 upvotes, #32 of 2026-01-27
  35. Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests 2 upvotes, #32 of 2026-01-27
  36. UI Remix: Supporting UI Design Through Interactive Example Retrieval and Remixing 2 upvotes, #32 of 2026-01-27
  37. MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts 2 upvotes, #32 of 2026-01-27
  38. Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing 1 upvotes, #38 of 2026-01-27
  39. RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-Agents 1 upvotes, #38 of 2026-01-27
  40. HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMs 1 upvotes, #38 of 2026-01-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.