Jiaqi Wang

Jiaqi Wang on Hugging Face Daily Papers: 69 papers, 26 in the top 3 of their day, 3,102 upvotes.

  1. JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence 200 upvotes, #1 of 2026-06-16
  2. AdaCodec: A Predictive Visual Code for Video MLLMs 5 upvotes, #27 of 2026-06-05
  3. Not only where, But when: Temporal Scheduling for RLVR 6 upvotes, #46 of 2026-06-02
  4. LoMo: Local Modality Substitution for Deeper Vision-Language Fusion 23 upvotes, #17 of 2026-05-29
  5. Channel-wise Vector Quantization 15 upvotes, #23 of 2026-05-26
  6. DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders 5 upvotes, #37 of 2026-05-22
  7. WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation 45 upvotes, #9 of 2026-05-15
  8. Co-Evolving Policy Distillation 67 upvotes, #3 of 2026-05-01
  9. Near-Future Policy Optimization 68 upvotes, #2 of 2026-04-23
  10. EasyVideoR1: Easier RL for Video Understanding 40 upvotes, #6 of 2026-04-21
  11. Self-Distilled RLVR 157 upvotes, #3 of 2026-04-06
  12. Visual-ERM: Reward Modeling for Visual Equivalence 21 upvotes, #8 of 2026-03-16
  13. DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing 78 upvotes, #3 of 2026-02-13
  14. Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition 28 upvotes, #15 of 2026-02-10
  15. Unified Personalized Reward Model for Vision Generation 19 upvotes, #14 of 2026-02-04
  16. UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing 75 upvotes, #6 of 2026-02-03
  17. SS4D: Native 4D Generative Model via Structured Spacetime Latents 12 upvotes, #17 of 2025-12-17
  18. EtCon: Edit-then-Consolidate for Reliable Knowledge Editing 7 upvotes, #13 of 2025-12-11
  19. ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning 45 upvotes, #5 of 2025-12-05
  20. Think Visually, Reason Textually: Vision-Language Synergy in ARC 8 upvotes, #20 of 2025-11-26
  21. UniREditBench: A Unified Reasoning-based Image Editing Benchmark 36 upvotes, #5 of 2025-11-04
  22. Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning 27 upvotes, #6 of 2025-11-03
  23. STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence 18 upvotes, #18 of 2025-10-29
  24. UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation 66 upvotes, #4 of 2025-10-22
  25. RLFR: Extending Reinforcement Learning for LLMs with Flow Environment 35 upvotes, #5 of 2025-10-14
  26. G^2RPO: Granular GRPO for Precise Reward in Flow Models 5 upvotes, #26 of 2025-10-09
  27. SPARK: Synergistic Policy And Reward Co-Evolving Framework 16 upvotes, #21 of 2025-09-29
  28. CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning 31 upvotes, #10 of 2025-09-29
  29. MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing 100 upvotes, #4 of 2025-09-29
  30. SIM-CoT: Supervised Implicit Chain-of-Thought 35 upvotes, #2 of 2025-09-25
  31. Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning 85 upvotes, #2 of 2025-08-29
  32. CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning 35 upvotes, #3 of 2025-08-28
  33. SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience 46 upvotes, #6 of 2025-08-07
  34. Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models 61 upvotes, #2 of 2025-08-04
  35. SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction 37 upvotes, #7 of 2025-07-22
  36. ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing 25 upvotes, #8 of 2025-06-25
  37. GeometryZero: Improving Geometry Solving for LLM with Group Contrastive Policy Optimization 3 upvotes, #37 of 2025-06-10
  38. Visual Agentic Reinforcement Fine-Tuning 31 upvotes, #6 of 2025-05-21
  39. Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning 87 upvotes, #2 of 2025-05-07
  40. HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance 11 upvotes, #10 of 2025-04-09
  41. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation 35 upvotes, #4 of 2025-02-20
  42. Light-A-Video: Training-free Video Relighting via Progressive Light Fusion 37 upvotes, #6 of 2025-02-13
  43. VideoRoPE: What Makes for Good Video Rotary Position Embedding? 60 upvotes, #3 of 2025-02-10
  44. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 39 upvotes, #6 of 2025-01-22
  45. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? 36 upvotes, #5 of 2025-01-13
  46. BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 34 upvotes, #3 of 2025-01-07
  47. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction 32 upvotes, #4 of 2025-01-07
  48. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
  49. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models 20 upvotes, #7 of 2024-12-11
  50. X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models 61 upvotes, #1 of 2024-12-03
  51. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models 34 upvotes, #1 of 2024-10-24
  52. SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree 61 upvotes, #2 of 2024-10-22
  53. Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate 36 upvotes, #7 of 2024-10-10
  54. BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way 10 upvotes, #24 of 2024-10-10
  55. VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models 11 upvotes, #6 of 2024-07-17
  56. InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
  57. Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 33 upvotes, #5 of 2024-06-21
  58. MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs 52 upvotes, #1 of 2024-06-18
  59. MotionClone: Training-Free Motion Cloning for Controllable Video Generation 38 upvotes, #3 of 2024-06-13
  60. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions 61 upvotes, #1 of 2024-06-07
  61. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
  62. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
  63. InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
  64. InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
  65. Gemini vs GPT-4V: A Preliminary Comparison and Combination of Vision-Language Models Through Qualitative Cases 18 upvotes, #5 of 2023-12-27
  66. HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image 22 upvotes, #3 of 2023-12-08
  67. Alpha-CLIP: A CLIP Model Focusing on Wherever You Want 33 upvotes, #3 of 2023-12-07
  68. OneLLM: One Framework to Align All Modalities with Language 23 upvotes, #6 of 2023-12-06
  69. GPT4Point: A Unified Framework for Point-Language Understanding and Generation 9 upvotes, #18 of 2023-12-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.