Jianfeng Gao
Jianfeng Gao on Hugging Face Daily Papers: 34 papers, 14 in the top 3 of their day, 1,132 upvotes.
- Towards a Medical AI Scientist 86 upvotes, #2 of 2026-03-31
- FlowRL: Matching Reward Distributions for LLM Reasoning 100 upvotes, #2 of 2025-09-19
- Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math 37 upvotes, #3 of 2025-05-01
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example 88 upvotes, #1 of 2025-04-30
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention 8 upvotes, #8 of 2025-04-29
- Magma: A Foundation Model for Multimodal AI Agents 46 upvotes, #5 of 2025-02-19
- SCBench: A KV Cache-Centric Analysis of Long-Context Methods 8 upvotes, #13 of 2024-12-16
- TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies 2 upvotes, #17 of 2024-12-16
- OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation 10 upvotes, #14 of 2024-12-13
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models 14 upvotes, #11 of 2024-10-15
- Improving Autonomous AI Agents with Reflective Tree Search and Self-Learning 8 upvotes, #17 of 2024-10-04
- GRIN: GRadient-INformed MoE 14 upvotes, #7 of 2024-09-19
- Matryoshka Multimodal Models 29 upvotes, #3 of 2024-05-28
- List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs 14 upvotes, #8 of 2024-04-26
- Pix2Gif: Motion-Guided Diffusion for GIF Generation 12 upvotes, #9 of 2024-03-08
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models 87 upvotes, #3 of 2024-02-28
- An Interactive Agent Foundation Model 29 upvotes, #3 of 2024-02-09
- TrustLLM: Trustworthiness in Large Language Models 69 upvotes, #1 of 2024-01-12
- Localized Symbolic Knowledge Distillation for Visual Commonsense Models 3 upvotes, #8 of 2023-12-11
- LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
- Visual In-Context Prompting 18 upvotes, #7 of 2023-11-23
- GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
- LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents 50 upvotes, #2 of 2023-11-10
- Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs 12 upvotes, #8 of 2023-11-07
- LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing 42 upvotes, #2 of 2023-11-02
- Teaching Language Models to Self-Improve through Interactive Demonstrations 12 upvotes, #8 of 2023-10-23
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V 29 upvotes, #4 of 2023-10-18
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
- MindAgent: Emergent Gaming Interaction 12 upvotes, #9 of 2023-09-19
- An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models 19 upvotes, #6 of 2023-09-19
- Semantic-SAM: Segment and Recognize Anything at Any Granularity 23 upvotes, #3 of 2023-07-11
- Demystifying GPT Self-Repair for Code Generation 21 upvotes, #3 of 2023-06-19
- Augmenting Language Models with Long-Term Memory 19 upvotes, #4 of 2023-06-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.