Tianyu Pang

Tianyu Pang on Hugging Face Daily Papers: 26 papers, 4 in the top 3 of their day, 814 upvotes.

  1. Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models 41 upvotes, #9 of 2026-06-10
  2. Rethinking the Divergence Regularization in LLM RL 33 upvotes, #11 of 2026-06-10
  3. Variational Reasoning for Language Models 66 upvotes, #5 of 2025-09-29
  4. Language Models Can Learn from Verbal Feedback Without Scalar Rewards 64 upvotes, #6 of 2025-09-29
  5. VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use 64 upvotes, #5 of 2025-09-03
  6. Fostering Video Reasoning via Next-Event Prediction 27 upvotes, #11 of 2025-05-29
  7. Reinforcing General Reasoning without Verifiers 26 upvotes, #17 of 2025-05-28
  8. Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment 8 upvotes, #37 of 2025-05-28
  9. Lifelong Safety Alignment for Language Models 23 upvotes, #15 of 2025-05-27
  10. QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design 38 upvotes, #6 of 2025-05-23
  11. BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms 4 upvotes, #32 of 2025-05-22
  12. Optimizing Anytime Reasoning via Budget Relative Policy Optimization 34 upvotes, #3 of 2025-05-21
  13. FlowReasoner: Reinforcing Query-Level Meta-Agents 46 upvotes, #3 of 2025-04-22
  14. NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation 18 upvotes, #11 of 2025-04-18
  15. Efficient Process Reward Model Training via Active Learning 13 upvotes, #14 of 2025-04-16
  16. Understanding R1-Zero-Like Training: A Critical Perspective 36 upvotes, #7 of 2025-04-03
  17. Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs 11 upvotes, #14 of 2025-02-18
  18. Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models 18 upvotes, #4 of 2024-12-30
  19. When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training 13 upvotes, #5 of 2024-11-21
  20. Improving Long-Text Alignment for Text-to-Image Diffusion Models 13 upvotes, #8 of 2024-10-17
  21. Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates 6 upvotes, #20 of 2024-10-11
  22. RegMix: Data Mixture as Regression for Language Model Pre-training 24 upvotes, #7 of 2024-07-02
  23. Bootstrapping Language Models with DPO Implicit Rewards 34 upvotes, #3 of 2024-06-19
  24. Weak-to-Strong Jailbreaking on Large Language Models 16 upvotes, #8 of 2024-01-31
  25. LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition 34 upvotes, #1 of 2023-07-26
  26. Efficient Diffusion Policies for Offline Reinforcement Learning 2 upvotes, #5 of 2023-06-01

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.