Daily Papers of 2025-03-12
- Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia 92 upvotes, #1 of 2025-03-12
- LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL 78 upvotes, #2 of 2025-03-12
- YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
- Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning 36 upvotes, #4 of 2025-03-12
- MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice 32 upvotes, #5 of 2025-03-12
- Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model 31 upvotes, #6 of 2025-03-12
- Gemini Embedding: Generalizable Embeddings from Gemini 31 upvotes, #6 of 2025-03-12
- Video Action Differencing 30 upvotes, #8 of 2025-03-12
- UniF^2ace: Fine-grained Face Understanding and Generation with Unified Multimodal Models 28 upvotes, #9 of 2025-03-12
- SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories 24 upvotes, #10 of 2025-03-12
- Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling 23 upvotes, #11 of 2025-03-12
- Implicit Reasoning in Transformers is Reasoning through Shortcuts 19 upvotes, #12 of 2025-03-12
- LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization 18 upvotes, #13 of 2025-03-12
- Exploiting Instruction-Following Retrievers for Malicious Information Retrieval 16 upvotes, #14 of 2025-03-12
- OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models 16 upvotes, #14 of 2025-03-12
- CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing 11 upvotes, #16 of 2025-03-12
- "Principal Components" Enable A New Language of Images 11 upvotes, #16 of 2025-03-12
- Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru 10 upvotes, #18 of 2025-03-12
- VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering 9 upvotes, #19 of 2025-03-12
- Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Enhancement Protocol 8 upvotes, #20 of 2025-03-12
- ^RFLAV: Rolling Flow matching for infinite Audio Video generation 8 upvotes, #20 of 2025-03-12
- Mixture of Experts Made Intrinsically Interpretable 7 upvotes, #22 of 2025-03-12
- AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models 7 upvotes, #22 of 2025-03-12
- Inductive Moment Matching 6 upvotes, #24 of 2025-03-12
- AI-native Memory 2.0: Second Me 6 upvotes, #24 of 2025-03-12
- Referring to Any Person 6 upvotes, #24 of 2025-03-12
- BiasEdit: Debiasing Stereotyped Language Models via Model Editing 6 upvotes, #24 of 2025-03-12
- LocAgent: Graph-Guided LLM Agents for Code Localization 6 upvotes, #24 of 2025-03-12
- Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation 5 upvotes, #29 of 2025-03-12
- RayFlow: Instance-Aware Diffusion Acceleration via Adaptive Flow Trajectories 5 upvotes, #29 of 2025-03-12
- Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents 5 upvotes, #29 of 2025-03-12
- Evaluating Intelligence via Trial and Error 4 upvotes, #32 of 2025-03-12
- Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence 4 upvotes, #32 of 2025-03-12
- ObjectMover: Generative Object Movement with Video Prior 4 upvotes, #32 of 2025-03-12
- QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension 4 upvotes, #32 of 2025-03-12
- Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts 3 upvotes, #36 of 2025-03-12
- Ideas in Inference-time Scaling can Benefit Generative Pre-training Algorithms 2 upvotes, #37 of 2025-03-12
- NullFace: Training-Free Localized Face Anonymization 2 upvotes, #37 of 2025-03-12
- OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction 1 upvotes, #39 of 2025-03-12
- PlainQAFact: Automatic Factuality Evaluation Metric for Biomedical Plain Language Summaries Generation 1 upvotes, #39 of 2025-03-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.