Daily Papers of 2025-02-25

  1. VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing 71 upvotes, #1 of 2025-02-25
  2. Thus Spake Long-Context Large Language Model 66 upvotes, #2 of 2025-02-25
  3. Slamming: Training a Speech Language Model on One GPU in a Day 65 upvotes, #3 of 2025-02-25
  4. DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks 51 upvotes, #4 of 2025-02-25
  5. Audio-FLAN: A Preliminary Release 32 upvotes, #5 of 2025-02-25
  6. GCC: Generative Color Constancy via Diffusing a Color Checker 27 upvotes, #6 of 2025-02-25
  7. Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning 24 upvotes, #7 of 2025-02-25
  8. CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models 23 upvotes, #8 of 2025-02-25
  9. Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment 23 upvotes, #8 of 2025-02-25
  10. RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers 19 upvotes, #10 of 2025-02-25
  11. Stable-SPAM: How to Train in 4-Bit More Stably than 16-Bit Adam 16 upvotes, #11 of 2025-02-25
  12. Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models 15 upvotes, #12 of 2025-02-25
  13. Beyond Release: Access Considerations for Generative AI Systems 11 upvotes, #13 of 2025-02-25
  14. Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation 11 upvotes, #13 of 2025-02-25
  15. Mobile-Agent-V: Learning Mobile Device Operation Through Video-Guided Multi-Agent Collaboration 11 upvotes, #13 of 2025-02-25
  16. X-Dancer: Expressive Music to Human Dance Video Generation 11 upvotes, #13 of 2025-02-25
  17. Forecasting Open-Weight AI Model Growth on Hugging Face 10 upvotes, #17 of 2025-02-25
  18. Grounded Persuasive Language Generation for Automated Marketing 10 upvotes, #17 of 2025-02-25
  19. TAG: A Decentralized Framework for Multi-Agent Hierarchical Reinforcement Learning 8 upvotes, #19 of 2025-02-25
  20. Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties 7 upvotes, #20 of 2025-02-25
  21. Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models 6 upvotes, #21 of 2025-02-25
  22. InductionBench: LLMs Fail in the Simplest Complexity Class 6 upvotes, #21 of 2025-02-25
  23. Can Community Notes Replace Professional Fact-Checkers? 5 upvotes, #23 of 2025-02-25
  24. Pandora3D: A Comprehensive Framework for High-Quality 3D Shape and Texture Generation 5 upvotes, #23 of 2025-02-25
  25. MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use 4 upvotes, #25 of 2025-02-25
  26. Early-Exit and Instant Confidence Translation Quality Estimation 3 upvotes, #26 of 2025-02-25
  27. Mind the Gap! Static and Interactive Evaluations of Large Audio Models 3 upvotes, #26 of 2025-02-25
  28. MONSTER: Monash Scalable Time Series Evaluation Repository 2 upvotes, #28 of 2025-02-25
  29. Self-Taught Agentic Long Context Understanding 2 upvotes, #28 of 2025-02-25
  30. The snake in the Brownian sphere 1 upvotes, #30 of 2025-02-25
  31. M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment 1 upvotes, #30 of 2025-02-25
  32. Diagnosing COVID-19 Severity from Chest X-Ray Images Using ViT and CNN Architectures 1 upvotes, #30 of 2025-02-25
  33. MegaLoc: One Retrieval to Place Them All 1 upvotes, #30 of 2025-02-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.