Daily Papers of 2026-01-19
- Your Group-Relative Advantage Is Biased 144 upvotes, #1 of 2026-01-19
- RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation 57 upvotes, #2 of 2026-01-19
- The Poisoned Apple Effect: Strategic Manipulation of Mediated Markets via Technology Expansion of AI Agents 46 upvotes, #3 of 2026-01-19
- Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text 38 upvotes, #4 of 2026-01-19
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts 34 upvotes, #5 of 2026-01-19
- When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs 26 upvotes, #6 of 2026-01-19
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models 24 upvotes, #7 of 2026-01-19
- ShapeR: Robust Conditional 3D Shape Generation from Casual Captures 20 upvotes, #8 of 2026-01-19
- Future Optical Flow Prediction Improves Robot Control & Video Generation 19 upvotes, #9 of 2026-01-19
- FrankenMotion: Part-level Human Motion Generation and Composition 18 upvotes, #10 of 2026-01-19
- Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM 17 upvotes, #11 of 2026-01-19
- BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search 17 upvotes, #11 of 2026-01-19
- ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection 15 upvotes, #13 of 2026-01-19
- Reasoning Models Generate Societies of Thought 11 upvotes, #14 of 2026-01-19
- PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models 11 upvotes, #14 of 2026-01-19
- Language of Thought Shapes Output Diversity in Large Language Models 9 upvotes, #16 of 2026-01-19
- PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records 8 upvotes, #17 of 2026-01-19
- Building Production-Ready Probes For Gemini 7 upvotes, #18 of 2026-01-19
- More Images, More Problems? A Controlled Analysis of VLM Failure Modes 6 upvotes, #19 of 2026-01-19
- AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems 4 upvotes, #20 of 2026-01-19
- PhyRPR: Training-Free Physics-Constrained Video Generation 3 upvotes, #21 of 2026-01-19
- What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge 3 upvotes, #21 of 2026-01-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.