Daily Papers of 2025-10-14
- QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs 164 upvotes, #1 of 2025-10-14
- Diffusion Transformers with Representation Autoencoders 155 upvotes, #2 of 2025-10-14
- Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States 50 upvotes, #3 of 2025-10-14
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs 45 upvotes, #4 of 2025-10-14
- Spotlight on Token Perception for Multimodal Reinforcement Learning 35 upvotes, #5 of 2025-10-14
- RLFR: Extending Reinforcement Learning for LLMs with Flow Environment 35 upvotes, #5 of 2025-10-14
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models 33 upvotes, #7 of 2025-10-14
- Demystifying Reinforcement Learning in Agentic Reasoning 30 upvotes, #8 of 2025-10-14
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training 29 upvotes, #9 of 2025-10-14
- AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration 28 upvotes, #10 of 2025-10-14
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions 27 upvotes, #11 of 2025-10-14
- Building a Foundational Guardrail for General Agentic Systems via Synthetic Data 26 upvotes, #12 of 2025-10-14
- Don't Just Fine-tune the Agent, Tune the Environment 26 upvotes, #12 of 2025-10-14
- DocReward: A Document Reward Model for Structuring and Stylizing 26 upvotes, #12 of 2025-10-14
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems 26 upvotes, #12 of 2025-10-14
- GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving 25 upvotes, #16 of 2025-10-14
- Making Mathematical Reasoning Adaptive 22 upvotes, #17 of 2025-10-14
- FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs 19 upvotes, #18 of 2025-10-14
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning 18 upvotes, #19 of 2025-10-14
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning 17 upvotes, #20 of 2025-10-14
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes 16 upvotes, #21 of 2025-10-14
- On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models 14 upvotes, #22 of 2025-10-14
- SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models 14 upvotes, #22 of 2025-10-14
- SwarmSys: Decentralized Swarm-Inspired Agents for Scalable and Adaptive Reasoning 13 upvotes, #24 of 2025-10-14
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images 13 upvotes, #24 of 2025-10-14
- High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting 11 upvotes, #26 of 2025-10-14
- FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding 11 upvotes, #26 of 2025-10-14
- Stable Video Infinity: Infinite-Length Video Generation with Error Recycling 10 upvotes, #28 of 2025-10-14
- ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding 10 upvotes, #28 of 2025-10-14
- Self-Improving LLM Agents at Test-Time 9 upvotes, #30 of 2025-10-14
- Skill-Targeted Adaptive Training 9 upvotes, #30 of 2025-10-14
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections 8 upvotes, #32 of 2025-10-14
- HUME: Measuring the Human-Model Performance Gap in Text Embedding Task 8 upvotes, #32 of 2025-10-14
- PEAR: Phase Entropy Aware Reward for Efficient Reasoning 7 upvotes, #34 of 2025-10-14
- The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs 6 upvotes, #35 of 2025-10-14
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference 6 upvotes, #35 of 2025-10-14
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing 5 upvotes, #37 of 2025-10-14
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation 5 upvotes, #37 of 2025-10-14
- InfiniHuman: Infinite 3D Human Creation with Precise Control 5 upvotes, #37 of 2025-10-14
- World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge 4 upvotes, #40 of 2025-10-14
- oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning 4 upvotes, #40 of 2025-10-14
- Multimodal Policy Internalization for Conversational Agents 4 upvotes, #40 of 2025-10-14
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining 4 upvotes, #40 of 2025-10-14
- A Tale of LLMs and Induced Small Proxies: Scalable Agents for Knowledge Mining 3 upvotes, #44 of 2025-10-14
- Graph Diffusion Transformers are In-Context Molecular Designers 3 upvotes, #44 of 2025-10-14
- LLaMAX2: Your Translation-Enhanced Model also Performs Well in Reasoning 3 upvotes, #44 of 2025-10-14
- VLM-Guided Adaptive Negative Prompting for Creative Generation 3 upvotes, #44 of 2025-10-14
- AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model 3 upvotes, #44 of 2025-10-14
- Through the Perspective of LiDAR: A Feature-Enriched and Uncertainty-Aware Annotation Pipeline for Terrestrial Point Cloud Segmentation 2 upvotes, #49 of 2025-10-14
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs 2 upvotes, #49 of 2025-10-14
- The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution 2 upvotes, #49 of 2025-10-14
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models 2 upvotes, #49 of 2025-10-14
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment 2 upvotes, #49 of 2025-10-14
- Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis Behavior 1 upvotes, #54 of 2025-10-14
- Are Large Reasoning Models Interruptible? 1 upvotes, #54 of 2025-10-14
- MultiCOIN: Multi-Modal COntrollable Video INbetweening 0 upvotes, #56 of 2025-10-14
- The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers 0 upvotes, #56 of 2025-10-14
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.