Daily Papers of 2025-03-11
- Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders 212 upvotes, #1 of 2025-03-11
- SEAP: Training-free Sparse Expert Activation Pruning Unlock the Brainpower of Large Language Models 63 upvotes, #2 of 2025-03-11
- MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scale Reinforcement Learning 53 upvotes, #3 of 2025-03-11
- Automated Movie Generation via Multi-Agent CoT Planning 40 upvotes, #4 of 2025-03-11
- Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning 37 upvotes, #5 of 2025-03-11
- VACE: All-in-One Video Creation and Editing 37 upvotes, #5 of 2025-03-11
- FedRand: Enhancing Privacy in Federated Learning with Randomized LoRA Subparameter Updates 28 upvotes, #7 of 2025-03-11
- DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs 27 upvotes, #8 of 2025-03-11
- EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer 24 upvotes, #9 of 2025-03-11
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models 22 upvotes, #10 of 2025-03-11
- Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning 20 upvotes, #11 of 2025-03-11
- AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning 19 upvotes, #12 of 2025-03-11
- FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation 18 upvotes, #13 of 2025-03-11
- WritingBench: A Comprehensive Benchmark for Generative Writing 16 upvotes, #14 of 2025-03-11
- SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing 15 upvotes, #15 of 2025-03-11
- Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment 15 upvotes, #15 of 2025-03-11
- Agent models: Internalizing Chain-of-Action Generation into Reasoning models 14 upvotes, #17 of 2025-03-11
- MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning 14 upvotes, #17 of 2025-03-11
- DreamRelation: Relation-Centric Video Customization 13 upvotes, #19 of 2025-03-11
- LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning 12 upvotes, #20 of 2025-03-11
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement 9 upvotes, #21 of 2025-03-11
- Effective and Efficient Masked Image Generation Models 9 upvotes, #21 of 2025-03-11
- Words or Vision: Do Vision-Language Models Have Blind Faith in Text? 8 upvotes, #23 of 2025-03-11
- PE3R: Perception-Efficient 3D Reconstruction 8 upvotes, #23 of 2025-03-11
- This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs 7 upvotes, #25 of 2025-03-11
- State-offset Tuning: State-based Parameter-Efficient Fine-Tuning for State Space Models 5 upvotes, #26 of 2025-03-11
- BlackGoose Rimer: Harnessing RWKV-7 as a Simple yet Superior Replacement for Transformers in Large-Scale Time Series Modeling 5 upvotes, #26 of 2025-03-11
- YOLOE: Real-Time Seeing Anything 5 upvotes, #26 of 2025-03-11
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations 4 upvotes, #29 of 2025-03-11
- DiffCLIP: Differential Attention Meets CLIP 4 upvotes, #29 of 2025-03-11
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation 4 upvotes, #29 of 2025-03-11
- Efficient Distillation of Classifier-Free Guidance using Adapters 4 upvotes, #29 of 2025-03-11
- Detection Avoidance Techniques for Large Language Models 4 upvotes, #29 of 2025-03-11
- Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries 3 upvotes, #34 of 2025-03-11
- What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization 3 upvotes, #34 of 2025-03-11
- ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks 3 upvotes, #34 of 2025-03-11
- Should VLMs be Pre-trained with Image Data? 3 upvotes, #34 of 2025-03-11
- Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces 2 upvotes, #38 of 2025-03-11
- Novel Object 6D Pose Estimation with a Single Reference View 2 upvotes, #38 of 2025-03-11
- Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model 2 upvotes, #38 of 2025-03-11
- Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs 2 upvotes, #38 of 2025-03-11
- A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning 2 upvotes, #38 of 2025-03-11
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models 2 upvotes, #38 of 2025-03-11
- REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding 2 upvotes, #38 of 2025-03-11
- HumanMM: Global Human Motion Recovery from Multi-shot Videos 2 upvotes, #38 of 2025-03-11
- Feynman-Kac Correctors in Diffusion: Annealing, Guidance, and Product of Experts 1 upvotes, #46 of 2025-03-11
- NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp Detection 1 upvotes, #46 of 2025-03-11
- PhiloBERTA: A Transformer-Based Cross-Lingual Analysis of Greek and Latin Lexicons 1 upvotes, #46 of 2025-03-11
- Symbolic Mixture-of-Experts: Adaptive Skill-based Routing for Heterogeneous Reasoning 1 upvotes, #46 of 2025-03-11
- RePO: ReLU-based Preference Optimization 1 upvotes, #46 of 2025-03-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.