Daily Papers of 2025-05-26
- TabSTAR: A Foundation Tabular Model With Semantically Target-Aware Representations 109 upvotes, #1 of 2025-05-26
- QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning 83 upvotes, #2 of 2025-05-26
- Distilling LLM Agent into Small Models with Retrieval and Code Tools 75 upvotes, #3 of 2025-05-26
- Quartet: Native FP4 Training Can Be Optimal for Large Language Models 73 upvotes, #4 of 2025-05-26
- Reasoning Model is Stubborn: Diagnosing Instruction Overriding in Reasoning Models 63 upvotes, #5 of 2025-05-26
- One RL to See Them All: Visual Triple Unified Reinforcement Learning 59 upvotes, #6 of 2025-05-26
- PhyX: Does Your Model Have the "Wits" for Physical Reasoning? 47 upvotes, #7 of 2025-05-26
- QwenLong-CPRS: Towards infty-LLMs with Dynamic Context Optimization 40 upvotes, #8 of 2025-05-26
- Scaling Image and Video Generation via Test-Time Evolutionary Search 39 upvotes, #9 of 2025-05-26
- Model Already Knows the Best Noise: Bayesian Active Noise Selection via Attention in Video Diffusion Model 30 upvotes, #10 of 2025-05-26
- MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback 30 upvotes, #10 of 2025-05-26
- VeriThinker: Learning to Verify Makes Reasoning Model Efficient 24 upvotes, #12 of 2025-05-26
- Diffusion Classifiers Understand Compositionality, but Conditions Apply 19 upvotes, #13 of 2025-05-26
- Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention 18 upvotes, #14 of 2025-05-26
- AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models 17 upvotes, #15 of 2025-05-26
- s3: You Don't Need That Much Data to Train a Search Agent via RL 16 upvotes, #16 of 2025-05-26
- Position of Uncertainty: A Cross-Linguistic Study of Positional Bias in Large Language Models 16 upvotes, #16 of 2025-05-26
- Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection 15 upvotes, #18 of 2025-05-26
- Time-R1: Towards Comprehensive Temporal Reasoning in LLMs 14 upvotes, #19 of 2025-05-26
- Thought-Augmented Policy Optimization: Bridging External Guidance and Internal Capabilities 14 upvotes, #19 of 2025-05-26
- FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow 14 upvotes, #19 of 2025-05-26
- Clear Nights Ahead: Towards Multi-Weather Nighttime Image Restoration 11 upvotes, #22 of 2025-05-26
- Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning 10 upvotes, #23 of 2025-05-26
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs 10 upvotes, #23 of 2025-05-26
- Synthetic Data RL: Task Definition Is All You Need 10 upvotes, #23 of 2025-05-26
- Speechless: Speech Instruction Training Without Speech for Low Resource Languages 10 upvotes, #23 of 2025-05-26
- ScanBot: Towards Intelligent Surface Scanning in Embodied Robotic Systems 9 upvotes, #27 of 2025-05-26
- Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models 9 upvotes, #27 of 2025-05-26
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study 8 upvotes, #29 of 2025-05-26
- RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning 7 upvotes, #30 of 2025-05-26
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning 6 upvotes, #31 of 2025-05-26
- Interactive Post-Training for Vision-Language-Action Models 6 upvotes, #31 of 2025-05-26
- DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation 6 upvotes, #31 of 2025-05-26
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection 5 upvotes, #34 of 2025-05-26
- Large Language Models Implicitly Learn to See and Hear Just By Reading 5 upvotes, #34 of 2025-05-26
- On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning 5 upvotes, #34 of 2025-05-26
- Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep Networks 4 upvotes, #37 of 2025-05-26
- Value-Guided Search for Efficient Chain-of-Thought Reasoning 4 upvotes, #37 of 2025-05-26
- Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering 3 upvotes, #39 of 2025-05-26
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models 3 upvotes, #39 of 2025-05-26
- TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios 2 upvotes, #41 of 2025-05-26
- NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning 2 upvotes, #41 of 2025-05-26
- Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA 2 upvotes, #41 of 2025-05-26
- FREESON: Retriever-Free Retrieval-Augmented Reasoning via Corpus-Traversing MCTS 2 upvotes, #41 of 2025-05-26
- FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation 1 upvotes, #45 of 2025-05-26
- NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities 1 upvotes, #45 of 2025-05-26
- Universal Biological Sequence Reranking for Improved De Novo Peptide Sequencing 0 upvotes, #47 of 2025-05-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.