Tianyi Zhou
Tianyi Zhou on Hugging Face Daily Papers: 63 papers, 15 in the top 3 of their day, 1,964 upvotes.
- A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review 71 upvotes, #12 of 2026-10-02
- How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review 48 upvotes, #8 of 2026-08-14
- Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning 13 upvotes, #13 of 2026-08-10
- Weak-to-Strong On-Policy Distillation 56 upvotes, #4 of 2026-08-03
- Visual Contrastive Self-Distillation 51 upvotes, #4 of 2026-07-24
- Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction 9 upvotes, #26 of 2026-06-30
- Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models 14 upvotes, #6 of 2026-06-22
- Guava: An Effective and Universal Harness for Embodied Manipulation 28 upvotes, #4 of 2026-06-18
- Self-Evolving Visual Questioner 15 upvotes, #14 of 2026-06-17
- Skip a Layer or Loop It? Learning Program-of-Layers in LLMs 24 upvotes, #12 of 2026-06-15
- When is Your LLM Steerable? 8 upvotes, #27 of 2026-06-15
- AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering? 1 upvotes, #58 of 2026-06-02
- Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks 20 upvotes, #5 of 2026-04-24
- ClawEnvKit: Automatic Environment Generation for Claw-Like Agents 28 upvotes, #7 of 2026-04-21
- When AI Navigates the Fog of War 28 upvotes, #8 of 2026-03-19
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook 26 upvotes, #4 of 2026-02-18
- What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis 14 upvotes, #9 of 2026-02-16
- TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models 10 upvotes, #15 of 2026-01-27
- Schoenfeld's Anatomy of Mathematical Reasoning by Language Models 14 upvotes, #4 of 2025-12-26
- Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction 23 upvotes, #8 of 2025-12-23
- V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions 9 upvotes, #16 of 2025-12-16
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs 24 upvotes, #7 of 2025-11-11
- ChartAB: A Benchmark for Chart Grounding & Dense Alignment 1 upvotes, #29 of 2025-10-31
- BLIP3o-NEXT: Next Frontier of Native Image Generation 21 upvotes, #12 of 2025-10-20
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play 123 upvotes, #3 of 2025-10-01
- Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory 13 upvotes, #14 of 2025-09-26
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding 7 upvotes, #21 of 2025-08-12
- Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs 28 upvotes, #8 of 2025-07-11
- FaSTA^*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing 38 upvotes, #2 of 2025-06-27
- Where to find Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test 26 upvotes, #5 of 2025-06-27
- Optimizing Length Compression in Large Reasoning Models 10 upvotes, #16 of 2025-06-18
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency 46 upvotes, #4 of 2025-06-17
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 80 upvotes, #1 of 2025-05-15
- Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs 6 upvotes, #10 of 2025-05-02
- WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents 18 upvotes, #11 of 2025-04-23
- Exploring Expert Failures Improves LLM Agent Tuning 11 upvotes, #16 of 2025-04-18
- ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness 45 upvotes, #3 of 2025-04-17
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients 39 upvotes, #4 of 2025-04-16
- Towards Visual Text Grounding of Multimodal Large Language Model 13 upvotes, #11 of 2025-04-11
- C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing 58 upvotes, #3 of 2025-04-11
- Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? 35 upvotes, #3 of 2025-04-10
- Efficient Reinforcement Finetuning via Adaptive Curriculum Learning 9 upvotes, #14 of 2025-04-09
- CoSTAast: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing 70 upvotes, #2 of 2025-03-14
- R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model 49 upvotes, #4 of 2025-03-10
- ATLaS: Agent Tuning via Learning Critical Steps 7 upvotes, #16 of 2025-03-05
- R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts 43 upvotes, #3 of 2025-02-28
- On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective 44 upvotes, #2 of 2025-02-20
- GUI Agents: A Survey 22 upvotes, #5 of 2024-12-19
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
- DynaSaur: Large Language Agents Beyond Predefined Actions 13 upvotes, #13 of 2024-11-05
- What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective 57 upvotes, #2 of 2024-11-01
- Diffusion Curriculum: Synthetic-to-Real Generative Curriculum Learning via Image-Guided Diffusion 13 upvotes, #9 of 2024-10-21
- BenTo: Benchmark Task Reduction with In-Context Transferability 20 upvotes, #11 of 2024-10-18
- Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free 45 upvotes, #1 of 2024-10-16
- WALL-E: World Alignment by Rule Learning Improves World Model-based LLM Agents 48 upvotes, #1 of 2024-10-11
- Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA 4 upvotes, #43 of 2024-10-10
- AUTOHALLUSION: Automatic Generation of Hallucination Benchmarks for Vision-Language Models 11 upvotes, #11 of 2024-06-28
- ODIN: Disentangled Reward Mitigates Hacking in RLHF 14 upvotes, #9 of 2024-02-13
- TrustLLM: Trustworthiness in Large Language Models 69 upvotes, #1 of 2024-01-12
- HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models 27 upvotes, #2 of 2023-10-24
- AlpaGasus: Training A Better Alpaca with Fewer Data 24 upvotes, #4 of 2023-07-18
- Diffusion Models Beat GANs on Image Classification 20 upvotes, #5 of 2023-07-18
- InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models 5 upvotes, #5 of 2023-06-06
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.