王维汉

王维汉 on Hugging Face Daily Papers: 8 papers, 3 in the top 3 of their day, 487 upvotes.

  1. GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning 176 upvotes, #1 of 2025-07-02
  2. MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models 40 upvotes, #4 of 2025-01-08
  3. CogVLM2: Visual Language Models for Image and Video Understanding 55 upvotes, #2 of 2024-08-30
  4. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer 31 upvotes, #5 of 2024-08-13
  5. CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion 15 upvotes, #5 of 2024-03-11
  6. CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations 9 upvotes, #9 of 2024-02-07
  7. CogAgent: A Visual Language Model for GUI Agents 32 upvotes, #3 of 2023-12-15
  8. CogVLM: Visual Expert for Pretrained Language Models 27 upvotes, #4 of 2023-11-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.