Tianyu Pang
Tianyu Pang on Hugging Face Daily Papers: 26 papers, 4 in the top 3 of their day, 814 upvotes.
- Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models 41 upvotes, #9 of 2026-06-10
- Rethinking the Divergence Regularization in LLM RL 33 upvotes, #11 of 2026-06-10
- Variational Reasoning for Language Models 66 upvotes, #5 of 2025-09-29
- Language Models Can Learn from Verbal Feedback Without Scalar Rewards 64 upvotes, #6 of 2025-09-29
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use 64 upvotes, #5 of 2025-09-03
- Fostering Video Reasoning via Next-Event Prediction 27 upvotes, #11 of 2025-05-29
- Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment 8 upvotes, #37 of 2025-05-28
- Lifelong Safety Alignment for Language Models 23 upvotes, #15 of 2025-05-27
- QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design 38 upvotes, #6 of 2025-05-23
- BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms 4 upvotes, #32 of 2025-05-22
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization 34 upvotes, #3 of 2025-05-21
- FlowReasoner: Reinforcing Query-Level Meta-Agents 46 upvotes, #3 of 2025-04-22
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation 18 upvotes, #11 of 2025-04-18
- Efficient Process Reward Model Training via Active Learning 13 upvotes, #14 of 2025-04-16
- Understanding R1-Zero-Like Training: A Critical Perspective 36 upvotes, #7 of 2025-04-03
- Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs 11 upvotes, #14 of 2025-02-18
- Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models 18 upvotes, #4 of 2024-12-30
- When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training 13 upvotes, #5 of 2024-11-21
- Improving Long-Text Alignment for Text-to-Image Diffusion Models 13 upvotes, #8 of 2024-10-17
- Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates 6 upvotes, #20 of 2024-10-11
- RegMix: Data Mixture as Regression for Language Model Pre-training 24 upvotes, #7 of 2024-07-02
- Bootstrapping Language Models with DPO Implicit Rewards 34 upvotes, #3 of 2024-06-19
- Weak-to-Strong Jailbreaking on Large Language Models 16 upvotes, #8 of 2024-01-31
- LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition 34 upvotes, #1 of 2023-07-26
- Efficient Diffusion Policies for Offline Reinforcement Learning 2 upvotes, #5 of 2023-06-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.