Songyang Zhang
Songyang Zhang on Hugging Face Daily Papers: 24 papers, 13 in the top 3 of their day, 1,157 upvotes.
- Intern-S1: A Scientific Multimodal Foundation Model 242 upvotes, #1 of 2025-08-22
- Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis 3 upvotes, #15 of 2025-08-21
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward 32 upvotes, #4 of 2025-08-06
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination 77 upvotes, #1 of 2025-07-15
- CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards 16 upvotes, #8 of 2025-07-15
- Rethinking Verification for LLM Code Generation: From Generation to Testing 28 upvotes, #6 of 2025-07-10
- Coding Triangle: How Does Large Language Model Understand Code? 19 upvotes, #12 of 2025-07-09
- Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective 36 upvotes, #9 of 2025-05-27
- Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 55 upvotes, #3 of 2025-02-11
- Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement 14 upvotes, #13 of 2025-01-22
- Are Your LLMs Capable of Stable Reasoning? 87 upvotes, #1 of 2024-12-18
- CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution 55 upvotes, #3 of 2024-10-22
- ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs 13 upvotes, #8 of 2024-10-17
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models 39 upvotes, #1 of 2024-09-25
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios 21 upvotes, #2 of 2024-09-02
- NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window? 39 upvotes, #1 of 2024-07-17
- GTA: A Benchmark for General Tool Agents 8 upvotes, #15 of 2024-07-12
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
- Adapting LLaMA Decoder to Vision Transformer 12 upvotes, #6 of 2024-04-11
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
- InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
- InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning 19 upvotes, #2 of 2024-02-12
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.