Ge Zhang

Ge Zhang on Hugging Face Daily Papers: 119 papers, 43 in the top 3 of their day, 6,434 upvotes.

  1. Aspire: Can Models Self-Evolve from Vague Goals? 228 upvotes, #3 of 2026-09-03
  2. S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? 39 upvotes, #8 of 2026-09-03
  3. HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? 264 upvotes, #2 of 2026-09-03
  4. SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers 99 upvotes, #3 of 2026-09-02
  5. OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs 30 upvotes, #6 of 2026-08-24
  6. StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows 9 upvotes, #21 of 2026-08-19
  7. EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments 17 upvotes, #20 of 2026-07-07
  8. Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields 21 upvotes, #15 of 2026-06-10
  9. Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO 7 upvotes, #23 of 2026-05-08
  10. In-Place Test-Time Training 28 upvotes, #14 of 2026-04-08
  11. Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation 11 upvotes, #10 of 2026-04-06
  12. Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining 8 upvotes, #21 of 2026-03-13
  13. \$OneMillion-Bench: How Far are Language Agents from Human Experts? 26 upvotes, #9 of 2026-03-10
  14. Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization 22 upvotes, #8 of 2026-02-27
  15. BABE: Biology Arena BEnchmark 10 upvotes, #27 of 2026-02-06
  16. Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities 19 upvotes, #15 of 2026-02-06
  17. Context Forcing: Consistent Autoregressive Video Generation with Long Context 35 upvotes, #6 of 2026-02-06
  18. ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocation 42 upvotes, #7 of 2026-01-30
  19. The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning 48 upvotes, #4 of 2026-01-12
  20. Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space 54 upvotes, #2 of 2026-01-02
  21. AutoMV: An Automatic Multi-Agent System for Music Video Generation 5 upvotes, #27 of 2025-12-16
  22. NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents 42 upvotes, #7 of 2025-12-16
  23. From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence 248 upvotes, #1 of 2025-12-02
  24. DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains 4 upvotes, #14 of 2025-11-17
  25. Virtual Width Networks 34 upvotes, #3 of 2025-11-17
  26. MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs 17 upvotes, #10 of 2025-11-11
  27. RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization 7 upvotes, #18 of 2025-11-11
  28. MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity 7 upvotes, #9 of 2025-11-06
  29. Scaling Latent Reasoning via Looped Language Models 201 upvotes, #1 of 2025-10-30
  30. A^2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning 22 upvotes, #11 of 2025-10-20
  31. Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures 10 upvotes, #23 of 2025-10-17
  32. OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs 45 upvotes, #4 of 2025-10-14
  33. ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems 26 upvotes, #12 of 2025-10-14
  34. Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation 43 upvotes, #5 of 2025-10-02
  35. Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution 15 upvotes, #11 of 2025-10-02
  36. VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
  37. Towards Personalized Deep Research: Benchmarks and Evaluations 27 upvotes, #13 of 2025-09-30
  38. FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning 29 upvotes, #7 of 2025-09-19
  39. Reverse-Engineered Reasoning for Open-Ended Generation 142 upvotes, #1 of 2025-09-09
  40. Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? 54 upvotes, #4 of 2025-09-05
  41. UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning 112 upvotes, #2 of 2025-09-03
  42. TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling 76 upvotes, #2 of 2025-08-27
  43. FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction 61 upvotes, #3 of 2025-08-21
  44. Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL 114 upvotes, #1 of 2025-08-20
  45. MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents 16 upvotes, #5 of 2025-08-20
  46. WideSearch: Benchmarking Agentic Broad Info-Seeking 102 upvotes, #2 of 2025-08-12
  47. VeriGUI: Verifiable Long-Chain GUI Dataset 137 upvotes, #2 of 2025-08-07
  48. Efficient Agents: Building Effective Agents While Reducing Cost 79 upvotes, #3 of 2025-08-07
  49. Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference 120 upvotes, #1 of 2025-08-06
  50. Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving 103 upvotes, #1 of 2025-08-01
  51. A Systematic Analysis of Hybrid Linear Attention 22 upvotes, #9 of 2025-07-10
  52. First Return, Entropy-Eliciting Explore 23 upvotes, #8 of 2025-07-10
  53. CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization 38 upvotes, #6 of 2025-07-09
  54. A Survey on Latent Reasoning 78 upvotes, #2 of 2025-07-09
  55. Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving 66 upvotes, #3 of 2025-07-08
  56. OAgents: An Empirical Study of Building Effective Agents 34 upvotes, #4 of 2025-06-24
  57. Scaling Test-time Compute for LLM Agents 57 upvotes, #2 of 2025-06-18
  58. TaskCraft: Automated Generation of Agentic Tasks 30 upvotes, #10 of 2025-06-17
  59. VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation 14 upvotes, #15 of 2025-05-21
  60. General-Reasoner: Advancing LLM Reasoning Across All Domains 20 upvotes, #10 of 2025-05-21
  61. AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection 25 upvotes, #9 of 2025-05-13
  62. FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models 27 upvotes, #7 of 2025-05-06
  63. IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs 22 upvotes, #8 of 2025-04-23
  64. ReTool: Reinforcement Learning for Strategic Tool Use in LLMs 58 upvotes, #2 of 2025-04-17
  65. COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values 41 upvotes, #5 of 2025-04-09
  66. Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models 11 upvotes, #19 of 2025-03-25
  67. A Comprehensive Survey on Long Context Language Modeling 47 upvotes, #3 of 2025-03-24
  68. FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis 15 upvotes, #12 of 2025-03-19
  69. Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers 17 upvotes, #8 of 2025-03-17
  70. YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
  71. Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? 26 upvotes, #7 of 2025-02-27
  72. CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models 23 upvotes, #8 of 2025-02-25
  73. Audio-FLAN: A Preliminary Release 32 upvotes, #5 of 2025-02-25
  74. SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines 92 upvotes, #3 of 2025-02-21
  75. Generating Symbolic World Models via Test-time Scaling of Large Language Models 16 upvotes, #11 of 2025-02-10
  76. MAGA: MAssive Genre-Audience Reformulation to Pretraining Corpus Expansion 17 upvotes, #11 of 2025-02-07
  77. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey 49 upvotes, #3 of 2024-12-30
  78. PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos 6 upvotes, #21 of 2024-12-03
  79. OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision 42 upvotes, #2 of 2024-11-12
  80. OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models 99 upvotes, #1 of 2024-11-08
  81. M2rc-Eval: Massively Multilingual Repository-level Code Completion Evaluation 6 upvotes, #16 of 2024-11-04
  82. AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions 36 upvotes, #2 of 2024-10-30
  83. A Comparative Study on Reasoning Patterns of OpenAI's o1 Model 15 upvotes, #15 of 2024-10-18
  84. Can MLLMs Understand the Deep Implication Behind Chinese Images? 7 upvotes, #23 of 2024-10-18
  85. MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models 17 upvotes, #7 of 2024-10-16
  86. Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 26 upvotes, #6 of 2024-10-15
  87. ING-VP: MLLMs cannot Play Easy Vision-based Games Yet 8 upvotes, #26 of 2024-10-10
  88. General Preference Modeling with Preference Representations for Aligning Language Models 5 upvotes, #14 of 2024-10-03
  89. MIO: A Foundation Model on Multimodal Tokens 46 upvotes, #2 of 2024-09-30
  90. HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models 39 upvotes, #1 of 2024-09-25
  91. OmniBench: Towards The Future of Universal Omni-Language Models 24 upvotes, #4 of 2024-09-25
  92. Towards a Unified View of Preference Learning for Large Language Models: A Survey 67 upvotes, #1 of 2024-09-10
  93. FuzzCoder: Byte-level Fuzzing Test via Large Language Model 44 upvotes, #3 of 2024-09-06
  94. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-09-05
  95. Foundation Models for Music: A Survey 35 upvotes, #3 of 2024-08-27
  96. TableBench: A Comprehensive and Complex Benchmark for Table Question Answering 48 upvotes, #1 of 2024-08-21
  97. I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm 30 upvotes, #2 of 2024-08-16
  98. DDK: Distilling Domain Knowledge for Efficient Large Language Models 18 upvotes, #4 of 2024-07-25
  99. LongIns: A Challenging Long-context Instruction-based Exam for LLMs 18 upvotes, #6 of 2024-06-26
  100. MantisScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 13 upvotes, #8 of 2024-06-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.