Daily Papers of 2025-02-10

  1. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach 105 upvotes, #1 of 2025-02-10
  2. Goku: Flow Based Video Generative Foundation Models 82 upvotes, #2 of 2025-02-10
  3. VideoRoPE: What Makes for Good Video Rotary Position Embedding? 60 upvotes, #3 of 2025-02-10
  4. Fast Video Generation with Sliding Tile Attention 46 upvotes, #4 of 2025-02-10
  5. QuEST: Stable Training of LLMs with 1-Bit Weights and Activations 40 upvotes, #5 of 2025-02-10
  6. AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting 28 upvotes, #6 of 2025-02-10
  7. FlashVideo:Flowing Fidelity to Detail for Efficient High-Resolution Video Generation 22 upvotes, #7 of 2025-02-10
  8. Agency Is Frame-Dependent 21 upvotes, #8 of 2025-02-10
  9. DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails 20 upvotes, #9 of 2025-02-10
  10. Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models 18 upvotes, #10 of 2025-02-10
  11. Generating Symbolic World Models via Test-time Scaling of Large Language Models 16 upvotes, #11 of 2025-02-10
  12. CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance 10 upvotes, #12 of 2025-02-10
  13. On-device Sora: Enabling Diffusion-Based Text-to-Video Generation for Mobile Devices 10 upvotes, #12 of 2025-02-10
  14. CMoE: Fast Carving of Mixture-of-Experts for Efficient LLM Inference 10 upvotes, #12 of 2025-02-10
  15. Linear Correlation in LM's Compositional Generalization and Hallucination 10 upvotes, #12 of 2025-02-10
  16. No Task Left Behind: Isotropic Model Merging with Common and Task-Specific Subspaces 10 upvotes, #12 of 2025-02-10
  17. QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation 10 upvotes, #12 of 2025-02-10
  18. Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More 9 upvotes, #18 of 2025-02-10
  19. ARR: Question Answering with Large Language Models via Analyzing, Retrieving, and Reasoning 7 upvotes, #19 of 2025-02-10
  20. Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs 7 upvotes, #19 of 2025-02-10
  21. YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment 5 upvotes, #21 of 2025-02-10
  22. Value-Based Deep RL Scales Predictably 5 upvotes, #21 of 2025-02-10
  23. Continuous 3D Perception Model with Persistent State 3 upvotes, #23 of 2025-02-10
  24. Adaptive Semantic Prompt Caching with VectorQ 3 upvotes, #23 of 2025-02-10
  25. MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf 3 upvotes, #23 of 2025-02-10
  26. SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs 2 upvotes, #26 of 2025-02-10
  27. Intelligent Sensing-to-Action for Robust Autonomy at the Edge: Opportunities and Challenges 0 upvotes, #27 of 2025-02-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.