Jianfeng Gao

Jianfeng Gao on Hugging Face Daily Papers: 34 papers, 14 in the top 3 of their day, 1,132 upvotes.

  1. Towards a Medical AI Scientist 86 upvotes, #2 of 2026-03-31
  2. FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
  3. Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math 37 upvotes, #3 of 2025-05-01
  4. Reinforcement Learning for Reasoning in Large Language Models with One Training Example 88 upvotes, #1 of 2025-04-30
  5. MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention 8 upvotes, #8 of 2025-04-29
  6. Magma: A Foundation Model for Multimodal AI Agents 46 upvotes, #5 of 2025-02-19
  7. SCBench: A KV Cache-Centric Analysis of Long-Context Methods 8 upvotes, #13 of 2024-12-16
  8. TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies 2 upvotes, #17 of 2024-12-16
  9. OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation 10 upvotes, #14 of 2024-12-13
  10. Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
  11. TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models 14 upvotes, #11 of 2024-10-15
  12. Improving Autonomous AI Agents with Reflective Tree Search and Self-Learning 8 upvotes, #17 of 2024-10-04
  13. GRIN: GRadient-INformed MoE 14 upvotes, #7 of 2024-09-19
  14. Matryoshka Multimodal Models 29 upvotes, #3 of 2024-05-28
  15. List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs 14 upvotes, #8 of 2024-04-26
  16. Pix2Gif: Motion-Guided Diffusion for GIF Generation 12 upvotes, #9 of 2024-03-08
  17. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models 87 upvotes, #3 of 2024-02-28
  18. An Interactive Agent Foundation Model 29 upvotes, #3 of 2024-02-09
  19. TrustLLM: Trustworthiness in Large Language Models 69 upvotes, #1 of 2024-01-12
  20. Localized Symbolic Knowledge Distillation for Visual Commonsense Models 3 upvotes, #8 of 2023-12-11
  21. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
  22. Visual In-Context Prompting 18 upvotes, #7 of 2023-11-23
  23. GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
  24. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents 50 upvotes, #2 of 2023-11-10
  25. Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs 12 upvotes, #8 of 2023-11-07
  26. LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing 42 upvotes, #2 of 2023-11-02
  27. Teaching Language Models to Self-Improve through Interactive Demonstrations 12 upvotes, #8 of 2023-10-23
  28. Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V 29 upvotes, #4 of 2023-10-18
  29. Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
  30. MindAgent: Emergent Gaming Interaction 12 upvotes, #9 of 2023-09-19
  31. An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models 19 upvotes, #6 of 2023-09-19
  32. Semantic-SAM: Segment and Recognize Anything at Any Granularity 23 upvotes, #3 of 2023-07-11
  33. Demystifying GPT Self-Repair for Code Generation 21 upvotes, #3 of 2023-06-19
  34. Augmenting Language Models with Long-Term Memory 19 upvotes, #4 of 2023-06-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.