Daily Papers of 2025-02-18
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention 134 upvotes, #1 of 2025-02-18
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? 41 upvotes, #2 of 2025-02-18
- Learning Getting-Up Policies for Real-World Humanoid Robots 36 upvotes, #3 of 2025-02-18
- ReLearn: Unlearning via Learning for Large Language Models 28 upvotes, #4 of 2025-02-18
- I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models 27 upvotes, #5 of 2025-02-18
- How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training 21 upvotes, #6 of 2025-02-18
- IHEval: Evaluating Language Models on Following the Instruction Hierarchy 18 upvotes, #7 of 2025-02-18
- CRANE: Reasoning with constrained LLM generation 18 upvotes, #7 of 2025-02-18
- Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation 16 upvotes, #9 of 2025-02-18
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation 16 upvotes, #9 of 2025-02-18
- System Message Generation for User Preferences using Open-Source Models 15 upvotes, #11 of 2025-02-18
- Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening 15 upvotes, #11 of 2025-02-18
- Intuitive physics understanding emerges from self-supervised pretraining on natural videos 13 upvotes, #13 of 2025-02-18
- Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs 11 upvotes, #14 of 2025-02-18
- Talk Structurally, Act Hierarchically: A Collaborative Framework for LLM Multi-Agent Systems 10 upvotes, #15 of 2025-02-18
- SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors 10 upvotes, #15 of 2025-02-18
- The Mirage of Model Editing: Revisiting Evaluation in the Wild 10 upvotes, #15 of 2025-02-18
- Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents 9 upvotes, #18 of 2025-02-18
- video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model 8 upvotes, #19 of 2025-02-18
- One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs 7 upvotes, #20 of 2025-02-18
- SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL 7 upvotes, #20 of 2025-02-18
- MagicArticulate: Make Your 3D Models Articulation-Ready 7 upvotes, #20 of 2025-02-18
- Dyve: Thinking Fast and Slow for Dynamic Process Verification 6 upvotes, #23 of 2025-02-18
- Cuckoo: An IE Free Rider Hatched by Massive Nutrition in LLM's Nest 6 upvotes, #23 of 2025-02-18
- Building A Proof-Oriented Programmer That Is 64% Better Than GPT-4o Under Data Scarsity 6 upvotes, #23 of 2025-02-18
- EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling 5 upvotes, #26 of 2025-02-18
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning 5 upvotes, #26 of 2025-02-18
- PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning 5 upvotes, #26 of 2025-02-18
- Can a Single Model Master Both Multi-turn Conversations and Tool Use? CALM: A Unified Conversational Agentic Language Model 4 upvotes, #29 of 2025-02-18
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking 4 upvotes, #29 of 2025-02-18
- ILIAS: Instance-Level Image retrieval At Scale 4 upvotes, #29 of 2025-02-18
- Towards Data-Efficient Pretraining for Atomic Property Prediction 3 upvotes, #32 of 2025-02-18
- Large Language Models and Mathematical Reasoning Failures 3 upvotes, #32 of 2025-02-18
- Diffusion Models without Classifier-free Guidance 3 upvotes, #32 of 2025-02-18
- Better Embeddings with Coupled Adam 1 upvotes, #35 of 2025-02-18
- Data Valuation using Neural Networks for Efficient Instruction Fine-Tuning 1 upvotes, #35 of 2025-02-18
- ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability 1 upvotes, #37 of 2025-02-18
- Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance 0 upvotes, #37 of 2025-02-18
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.