Jifeng Dai

Jifeng Dai on Hugging Face Daily Papers: 19 papers, 8 in the top 3 of their day, 835 upvotes.

  1. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models 14 upvotes, #6 of 2025-07-21
  2. VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models 71 upvotes, #1 of 2025-04-24
  3. Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy 47 upvotes, #2 of 2025-03-27
  4. GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing 45 upvotes, #4 of 2025-03-14
  5. VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
  6. Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding 7 upvotes, #8 of 2025-01-16
  7. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 35 upvotes, #3 of 2024-12-16
  8. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 103 upvotes, #1 of 2024-12-09
  9. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization 61 upvotes, #1 of 2024-11-22
  10. PUMA: Empowering Unified MLLM with Multi-granular Visual Generation 51 upvotes, #4 of 2024-10-22
  11. MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models 56 upvotes, #1 of 2024-08-07
  12. Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams 11 upvotes, #10 of 2024-07-08
  13. InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
  14. Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling 39 upvotes, #4 of 2024-01-30
  15. ControlLLM: Augment Language Models with Tools by Searching on Graphs 18 upvotes, #3 of 2023-10-30
  16. The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World 13 upvotes, #8 of 2023-08-04
  17. Ghost in the Minecraft: Generally Capable Agents for Open-World Enviroments via Large Language Models with Text-based Knowledge and Memory 4 upvotes, #8 of 2023-05-30
  18. VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks 4 upvotes, #6 of 2023-05-19
  19. InternChat: Solving Vision-Centric Tasks by Interacting with Chatbots Beyond Language 5 upvotes, #5 of 2023-05-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.