Songyang Zhang

Songyang Zhang on Hugging Face Daily Papers: 24 papers, 13 in the top 3 of their day, 1,157 upvotes.

  1. Intern-S1: A Scientific Multimodal Foundation Model 242 upvotes, #1 of 2025-08-22
  2. Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis 3 upvotes, #15 of 2025-08-21
  3. CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward 32 upvotes, #4 of 2025-08-06
  4. Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination 77 upvotes, #1 of 2025-07-15
  5. CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards 16 upvotes, #8 of 2025-07-15
  6. Rethinking Verification for LLM Code Generation: From Generation to Testing 28 upvotes, #6 of 2025-07-10
  7. Coding Triangle: How Does Large Language Model Understand Code? 19 upvotes, #12 of 2025-07-09
  8. Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective 36 upvotes, #9 of 2025-05-27
  9. Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 55 upvotes, #3 of 2025-02-11
  10. Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement 14 upvotes, #13 of 2025-01-22
  11. Are Your LLMs Capable of Stable Reasoning? 87 upvotes, #1 of 2024-12-18
  12. CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution 55 upvotes, #3 of 2024-10-22
  13. ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs 13 upvotes, #8 of 2024-10-17
  14. HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models 39 upvotes, #1 of 2024-09-25
  15. UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios 21 upvotes, #2 of 2024-09-02
  16. NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window? 39 upvotes, #1 of 2024-07-17
  17. GTA: A Benchmark for General Tool Agents 8 upvotes, #15 of 2024-07-12
  18. InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
  19. Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
  20. Adapting LLaMA Decoder to Vision Transformer 12 upvotes, #6 of 2024-04-11
  21. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
  22. InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
  23. InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning 19 upvotes, #2 of 2024-02-12
  24. InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.