Daily Papers of 2025-05-26

  1. TabSTAR: A Foundation Tabular Model With Semantically Target-Aware Representations 109 upvotes, #1 of 2025-05-26
  2. QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning 83 upvotes, #2 of 2025-05-26
  3. Distilling LLM Agent into Small Models with Retrieval and Code Tools 75 upvotes, #3 of 2025-05-26
  4. Quartet: Native FP4 Training Can Be Optimal for Large Language Models 73 upvotes, #4 of 2025-05-26
  5. Reasoning Model is Stubborn: Diagnosing Instruction Overriding in Reasoning Models 63 upvotes, #5 of 2025-05-26
  6. One RL to See Them All: Visual Triple Unified Reinforcement Learning 59 upvotes, #6 of 2025-05-26
  7. PhyX: Does Your Model Have the "Wits" for Physical Reasoning? 47 upvotes, #7 of 2025-05-26
  8. QwenLong-CPRS: Towards infty-LLMs with Dynamic Context Optimization 40 upvotes, #8 of 2025-05-26
  9. Scaling Image and Video Generation via Test-Time Evolutionary Search 39 upvotes, #9 of 2025-05-26
  10. Model Already Knows the Best Noise: Bayesian Active Noise Selection via Attention in Video Diffusion Model 30 upvotes, #10 of 2025-05-26
  11. MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback 30 upvotes, #10 of 2025-05-26
  12. VeriThinker: Learning to Verify Makes Reasoning Model Efficient 24 upvotes, #12 of 2025-05-26
  13. Diffusion Classifiers Understand Compositionality, but Conditions Apply 19 upvotes, #13 of 2025-05-26
  14. Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention 18 upvotes, #14 of 2025-05-26
  15. AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models 17 upvotes, #15 of 2025-05-26
  16. s3: You Don't Need That Much Data to Train a Search Agent via RL 16 upvotes, #16 of 2025-05-26
  17. Position of Uncertainty: A Cross-Linguistic Study of Positional Bias in Large Language Models 16 upvotes, #16 of 2025-05-26
  18. Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection 15 upvotes, #18 of 2025-05-26
  19. Time-R1: Towards Comprehensive Temporal Reasoning in LLMs 14 upvotes, #19 of 2025-05-26
  20. Thought-Augmented Policy Optimization: Bridging External Guidance and Internal Capabilities 14 upvotes, #19 of 2025-05-26
  21. FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow 14 upvotes, #19 of 2025-05-26
  22. Clear Nights Ahead: Towards Multi-Weather Nighttime Image Restoration 11 upvotes, #22 of 2025-05-26
  23. Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning 10 upvotes, #23 of 2025-05-26
  24. RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs 10 upvotes, #23 of 2025-05-26
  25. Synthetic Data RL: Task Definition Is All You Need 10 upvotes, #23 of 2025-05-26
  26. Speechless: Speech Instruction Training Without Speech for Low Resource Languages 10 upvotes, #23 of 2025-05-26
  27. ScanBot: Towards Intelligent Surface Scanning in Embodied Robotic Systems 9 upvotes, #27 of 2025-05-26
  28. Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models 9 upvotes, #27 of 2025-05-26
  29. Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study 8 upvotes, #29 of 2025-05-26
  30. RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning 7 upvotes, #30 of 2025-05-26
  31. Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning 6 upvotes, #31 of 2025-05-26
  32. Interactive Post-Training for Vision-Language-Action Models 6 upvotes, #31 of 2025-05-26
  33. DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation 6 upvotes, #31 of 2025-05-26
  34. ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection 5 upvotes, #34 of 2025-05-26
  35. Large Language Models Implicitly Learn to See and Hear Just By Reading 5 upvotes, #34 of 2025-05-26
  36. On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning 5 upvotes, #34 of 2025-05-26
  37. Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep Networks 4 upvotes, #37 of 2025-05-26
  38. Value-Guided Search for Efficient Chain-of-Thought Reasoning 4 upvotes, #37 of 2025-05-26
  39. Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering 3 upvotes, #39 of 2025-05-26
  40. Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models 3 upvotes, #39 of 2025-05-26
  41. TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios 2 upvotes, #41 of 2025-05-26
  42. NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning 2 upvotes, #41 of 2025-05-26
  43. Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA 2 upvotes, #41 of 2025-05-26
  44. FREESON: Retriever-Free Retrieval-Augmented Reasoning via Corpus-Traversing MCTS 2 upvotes, #41 of 2025-05-26
  45. FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation 1 upvotes, #45 of 2025-05-26
  46. NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities 1 upvotes, #45 of 2025-05-26
  47. Universal Biological Sequence Reranking for Improved De Novo Peptide Sequencing 0 upvotes, #47 of 2025-05-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.