Daily Papers of 2025-01-07

  1. STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution 47 upvotes, #1 of 2025-01-07
  2. Test-time Computing: from System-1 Thinking to System-2 Thinking 36 upvotes, #2 of 2025-01-07
  3. BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 34 upvotes, #3 of 2025-01-07
  4. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction 32 upvotes, #4 of 2025-01-07
  5. Personalized Graph-Based Retrieval for Large Language Models 26 upvotes, #5 of 2025-01-07
  6. Scaling Laws for Floating Point Quantization Training 24 upvotes, #6 of 2025-01-07
  7. TransPixar: Advancing Text-to-Video Generation with Transparency 21 upvotes, #7 of 2025-01-07
  8. METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring 19 upvotes, #8 of 2025-01-07
  9. Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 19 upvotes, #8 of 2025-01-07
  10. Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models 16 upvotes, #10 of 2025-01-07
  11. GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking 16 upvotes, #10 of 2025-01-07
  12. DepthMaster: Taming Diffusion Models for Monocular Depth Estimation 15 upvotes, #12 of 2025-01-07
  13. PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models 13 upvotes, #13 of 2025-01-07
  14. ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use 9 upvotes, #14 of 2025-01-07
  15. AutoPresent: Designing Structured Visuals from Scratch 8 upvotes, #15 of 2025-01-07
  16. Ingredients: Blending Custom Photos with Video Diffusion Transformers 8 upvotes, #15 of 2025-01-07
  17. Samba-asr state-of-the-art speech recognition leveraging structured state-space models 7 upvotes, #17 of 2025-01-07
  18. Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation 6 upvotes, #18 of 2025-01-07
  19. ProTracker: Probabilistic Integration for Robust and Accurate Point Tracking 4 upvotes, #19 of 2025-01-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.