Daily Papers of 2025-02-25
- VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing 71 upvotes, #1 of 2025-02-25
- Thus Spake Long-Context Large Language Model 66 upvotes, #2 of 2025-02-25
- Slamming: Training a Speech Language Model on One GPU in a Day 65 upvotes, #3 of 2025-02-25
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks 51 upvotes, #4 of 2025-02-25
- Audio-FLAN: A Preliminary Release 32 upvotes, #5 of 2025-02-25
- GCC: Generative Color Constancy via Diffusing a Color Checker 27 upvotes, #6 of 2025-02-25
- Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning 24 upvotes, #7 of 2025-02-25
- CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models 23 upvotes, #8 of 2025-02-25
- Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment 23 upvotes, #8 of 2025-02-25
- RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers 19 upvotes, #10 of 2025-02-25
- Stable-SPAM: How to Train in 4-Bit More Stably than 16-Bit Adam 16 upvotes, #11 of 2025-02-25
- Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models 15 upvotes, #12 of 2025-02-25
- Beyond Release: Access Considerations for Generative AI Systems 11 upvotes, #13 of 2025-02-25
- Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation 11 upvotes, #13 of 2025-02-25
- Mobile-Agent-V: Learning Mobile Device Operation Through Video-Guided Multi-Agent Collaboration 11 upvotes, #13 of 2025-02-25
- X-Dancer: Expressive Music to Human Dance Video Generation 11 upvotes, #13 of 2025-02-25
- Forecasting Open-Weight AI Model Growth on Hugging Face 10 upvotes, #17 of 2025-02-25
- Grounded Persuasive Language Generation for Automated Marketing 10 upvotes, #17 of 2025-02-25
- TAG: A Decentralized Framework for Multi-Agent Hierarchical Reinforcement Learning 8 upvotes, #19 of 2025-02-25
- Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties 7 upvotes, #20 of 2025-02-25
- Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models 6 upvotes, #21 of 2025-02-25
- InductionBench: LLMs Fail in the Simplest Complexity Class 6 upvotes, #21 of 2025-02-25
- Can Community Notes Replace Professional Fact-Checkers? 5 upvotes, #23 of 2025-02-25
- Pandora3D: A Comprehensive Framework for High-Quality 3D Shape and Texture Generation 5 upvotes, #23 of 2025-02-25
- MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use 4 upvotes, #25 of 2025-02-25
- Early-Exit and Instant Confidence Translation Quality Estimation 3 upvotes, #26 of 2025-02-25
- Mind the Gap! Static and Interactive Evaluations of Large Audio Models 3 upvotes, #26 of 2025-02-25
- MONSTER: Monash Scalable Time Series Evaluation Repository 2 upvotes, #28 of 2025-02-25
- Self-Taught Agentic Long Context Understanding 2 upvotes, #28 of 2025-02-25
- The snake in the Brownian sphere 1 upvotes, #30 of 2025-02-25
- M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment 1 upvotes, #30 of 2025-02-25
- Diagnosing COVID-19 Severity from Chest X-Ray Images Using ViT and CNN Architectures 1 upvotes, #30 of 2025-02-25
- MegaLoc: One Retrieval to Place Them All 1 upvotes, #30 of 2025-02-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.