Dongfu Jiang

Dongfu Jiang on Hugging Face Daily Papers: 23 papers, 6 in the top 3 of their day, 1,623 upvotes.

  1. Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion 11 upvotes, #18 of 2026-06-17
  2. Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 15 upvotes, #16 of 2026-06-16
  3. Cosmos 3: Omnimodal World Models for Physical AI 115 upvotes, #1 of 2026-06-04
  4. AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? 29 upvotes, #7 of 2026-06-04
  5. RewardHarness: Self-Evolving Agentic Post-Training 9 upvotes, #20 of 2026-05-14
  6. Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction 105 upvotes, #2 of 2026-05-08
  7. Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 36 upvotes, #4 of 2026-04-15
  8. ClawBench: Can AI Agents Complete Everyday Online Tasks? 255 upvotes, #3 of 2026-04-10
  9. Watch Before You Answer: Learning from Visually Grounded Post-Training 35 upvotes, #11 of 2026-04-08
  10. OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis 91 upvotes, #3 of 2026-03-24
  11. Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 60 upvotes, #3 of 2026-03-20
  12. Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning 20 upvotes, #20 of 2025-09-30
  13. VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
  14. VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use 64 upvotes, #5 of 2025-09-03
  15. StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
  16. QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design 38 upvotes, #6 of 2025-05-23
  17. General-Reasoner: Advancing LLM Reasoning Across All Domains 20 upvotes, #10 of 2025-05-21
  18. ACECODER: Acing Coder RL via Automated Test-Case Synthesis 23 upvotes, #3 of 2025-02-05
  19. MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks 34 upvotes, #5 of 2024-10-15
  20. MantisScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 13 upvotes, #8 of 2024-06-24
  21. WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences 11 upvotes, #16 of 2024-06-18
  22. GenAI Arena: An Open Evaluation Platform for Generative Models 18 upvotes, #4 of 2024-06-10
  23. LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion 6 upvotes, #4 of 2023-06-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.