Lijun Wu

Lijun Wu on Hugging Face Daily Papers: 25 papers, 5 in the top 3 of their day, 1,568 upvotes.

  1. MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale 117 upvotes, #3 of 2026-04-07
  2. Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale 125 upvotes, #1 of 2026-03-27
  3. MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods 57 upvotes, #5 of 2026-01-30
  4. Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility 41 upvotes, #4 of 2026-01-27
  5. ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch 10 upvotes, #13 of 2026-01-26
  6. OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value 38 upvotes, #5 of 2025-12-17
  7. Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning 19 upvotes, #10 of 2025-10-08
  8. Sequential Diffusion Language Models 36 upvotes, #10 of 2025-09-30
  9. MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing 100 upvotes, #4 of 2025-09-29
  10. ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning 9 upvotes, #17 of 2025-09-26
  11. A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers 135 upvotes, #2 of 2025-09-01
  12. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 170 upvotes, #1 of 2025-08-26
  13. Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning 33 upvotes, #4 of 2025-07-24
  14. REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once 28 upvotes, #5 of 2025-07-15
  15. GTR-CoT: Graph Traversal as Visual Chain of Thought for Molecular Structure Recognition 13 upvotes, #15 of 2025-06-10
  16. Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
  17. CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges 16 upvotes, #4 of 2025-04-29
  18. A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis 27 upvotes, #7 of 2025-04-18
  19. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
  20. LEMMA: Learning from Errors for MatheMatical Advancement in LLMs 13 upvotes, #14 of 2025-03-25
  21. MathFusion: Enhancing Mathematic Problem-solving of LLM through Instruction Fusion 22 upvotes, #17 of 2025-03-21
  22. MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer 20 upvotes, #7 of 2025-03-20
  23. NatureLM: Deciphering the Language of Nature for Scientific Discovery 17 upvotes, #11 of 2025-02-12
  24. Multi-Agent Collaborative Data Selection for Efficient LLM Pretraining 17 upvotes, #6 of 2024-10-14
  25. MolXPT: Wrapping Molecules with Text for Generative Pre-training 1 upvotes, #16 of 2023-05-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.