Lijun Wu
Lijun Wu on Hugging Face Daily Papers: 25 papers, 5 in the top 3 of their day, 1,568 upvotes.
- MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale 117 upvotes, #3 of 2026-04-07
- Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale 125 upvotes, #1 of 2026-03-27
- MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods 57 upvotes, #5 of 2026-01-30
- Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility 41 upvotes, #4 of 2026-01-27
- ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch 10 upvotes, #13 of 2026-01-26
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value 38 upvotes, #5 of 2025-12-17
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning 19 upvotes, #10 of 2025-10-08
- Sequential Diffusion Language Models 36 upvotes, #10 of 2025-09-30
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing 100 upvotes, #4 of 2025-09-29
- ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning 9 upvotes, #17 of 2025-09-26
- A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers 135 upvotes, #2 of 2025-09-01
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 170 upvotes, #1 of 2025-08-26
- Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning 33 upvotes, #4 of 2025-07-24
- REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once 28 upvotes, #5 of 2025-07-15
- GTR-CoT: Graph Traversal as Visual Chain of Thought for Molecular Structure Recognition 13 upvotes, #15 of 2025-06-10
- Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
- CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges 16 upvotes, #4 of 2025-04-29
- A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis 27 upvotes, #7 of 2025-04-18
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
- LEMMA: Learning from Errors for MatheMatical Advancement in LLMs 13 upvotes, #14 of 2025-03-25
- MathFusion: Enhancing Mathematic Problem-solving of LLM through Instruction Fusion 22 upvotes, #17 of 2025-03-21
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer 20 upvotes, #7 of 2025-03-20
- NatureLM: Deciphering the Language of Nature for Scientific Discovery 17 upvotes, #11 of 2025-02-12
- Multi-Agent Collaborative Data Selection for Efficient LLM Pretraining 17 upvotes, #6 of 2024-10-14
- MolXPT: Wrapping Molecules with Text for Generative Pre-training 1 upvotes, #16 of 2023-05-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.