Tianyi Zhou

Tianyi Zhou on Hugging Face Daily Papers: 63 papers, 15 in the top 3 of their day, 1,964 upvotes.

  1. A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review 71 upvotes, #12 of 2026-10-02
  2. How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review 48 upvotes, #8 of 2026-08-14
  3. Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 13 upvotes, #13 of 2026-08-10
  4. Weak-to-Strong On-Policy Distillation 56 upvotes, #4 of 2026-08-03
  5. Visual Contrastive Self-Distillation 51 upvotes, #4 of 2026-07-24
  6. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction 9 upvotes, #26 of 2026-06-30
  7. Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models 14 upvotes, #6 of 2026-06-22
  8. Guava: An Effective and Universal Harness for Embodied Manipulation 28 upvotes, #4 of 2026-06-18
  9. Self-Evolving Visual Questioner 15 upvotes, #14 of 2026-06-17
  10. Skip a Layer or Loop It? Learning Program-of-Layers in LLMs 24 upvotes, #12 of 2026-06-15
  11. When is Your LLM Steerable? 8 upvotes, #27 of 2026-06-15
  12. AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering? 1 upvotes, #58 of 2026-06-02
  13. Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks 20 upvotes, #5 of 2026-04-24
  14. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents 28 upvotes, #7 of 2026-04-21
  15. When AI Navigates the Fog of War 28 upvotes, #8 of 2026-03-19
  16. Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook 26 upvotes, #4 of 2026-02-18
  17. What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis 14 upvotes, #9 of 2026-02-16
  18. TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models 10 upvotes, #15 of 2026-01-27
  19. Schoenfeld's Anatomy of Mathematical Reasoning by Language Models 14 upvotes, #4 of 2025-12-26
  20. Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction 23 upvotes, #8 of 2025-12-23
  21. V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions 9 upvotes, #16 of 2025-12-16
  22. Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs 24 upvotes, #7 of 2025-11-11
  23. ChartAB: A Benchmark for Chart Grounding & Dense Alignment 1 upvotes, #29 of 2025-10-31
  24. BLIP3o-NEXT: Next Frontier of Native Image Generation 21 upvotes, #12 of 2025-10-20
  25. Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play 123 upvotes, #3 of 2025-10-01
  26. Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory 13 upvotes, #14 of 2025-09-26
  27. VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding 7 upvotes, #21 of 2025-08-12
  28. Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs 28 upvotes, #8 of 2025-07-11
  29. FaSTA^*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing 38 upvotes, #2 of 2025-06-27
  30. Where to find Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test 26 upvotes, #5 of 2025-06-27
  31. Optimizing Length Compression in Large Reasoning Models 10 upvotes, #16 of 2025-06-18
  32. Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency 46 upvotes, #4 of 2025-06-17
  33. BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 80 upvotes, #1 of 2025-05-15
  34. Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs 6 upvotes, #10 of 2025-05-02
  35. WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents 18 upvotes, #11 of 2025-04-23
  36. Exploring Expert Failures Improves LLM Agent Tuning 11 upvotes, #16 of 2025-04-18
  37. ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness 45 upvotes, #3 of 2025-04-17
  38. How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients 39 upvotes, #4 of 2025-04-16
  39. Towards Visual Text Grounding of Multimodal Large Language Model 13 upvotes, #11 of 2025-04-11
  40. C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing 58 upvotes, #3 of 2025-04-11
  41. Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? 35 upvotes, #3 of 2025-04-10
  42. Efficient Reinforcement Finetuning via Adaptive Curriculum Learning 9 upvotes, #14 of 2025-04-09
  43. CoSTAast: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing 70 upvotes, #2 of 2025-03-14
  44. R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model 49 upvotes, #4 of 2025-03-10
  45. ATLaS: Agent Tuning via Learning Critical Steps 7 upvotes, #16 of 2025-03-05
  46. R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts 43 upvotes, #3 of 2025-02-28
  47. On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective 44 upvotes, #2 of 2025-02-20
  48. GUI Agents: A Survey 22 upvotes, #5 of 2024-12-19
  49. Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
  50. DynaSaur: Large Language Agents Beyond Predefined Actions 13 upvotes, #13 of 2024-11-05
  51. What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective 57 upvotes, #2 of 2024-11-01
  52. Diffusion Curriculum: Synthetic-to-Real Generative Curriculum Learning via Image-Guided Diffusion 13 upvotes, #9 of 2024-10-21
  53. BenTo: Benchmark Task Reduction with In-Context Transferability 20 upvotes, #11 of 2024-10-18
  54. Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free 45 upvotes, #1 of 2024-10-16
  55. WALL-E: World Alignment by Rule Learning Improves World Model-based LLM Agents 48 upvotes, #1 of 2024-10-11
  56. Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA 4 upvotes, #43 of 2024-10-10
  57. AUTOHALLUSION: Automatic Generation of Hallucination Benchmarks for Vision-Language Models 11 upvotes, #11 of 2024-06-28
  58. ODIN: Disentangled Reward Mitigates Hacking in RLHF 14 upvotes, #9 of 2024-02-13
  59. TrustLLM: Trustworthiness in Large Language Models 69 upvotes, #1 of 2024-01-12
  60. HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models 27 upvotes, #2 of 2023-10-24
  61. AlpaGasus: Training A Better Alpaca with Fewer Data 24 upvotes, #4 of 2023-07-18
  62. Diffusion Models Beat GANs on Image Classification 20 upvotes, #5 of 2023-07-18
  63. InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models 5 upvotes, #5 of 2023-06-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.