Hongsheng LI

Hongsheng LI on Hugging Face Daily Papers: 46 papers, 14 in the top 3 of their day, 1,313 upvotes.

  1. UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents 38 upvotes, #12 of 2025-05-28
  2. EnerVerse-AC: Envisioning Embodied Environments with Action Condition 21 upvotes, #7 of 2025-05-16
  3. WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch 16 upvotes, #11 of 2025-05-13
  4. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT 39 upvotes, #3 of 2025-05-02
  5. LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects 21 upvotes, #3 of 2025-04-29
  6. From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning 15 upvotes, #14 of 2025-04-23
  7. UI-R1: Enhancing Action Prediction of GUI Agents by Reinforcement Learning 54 upvotes, #3 of 2025-03-28
  8. LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis 25 upvotes, #6 of 2025-03-28
  9. Lumina-Image 2.0: A Unified and Efficient Image Generative Framework 18 upvotes, #11 of 2025-03-28
  10. Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning 33 upvotes, #4 of 2025-03-17
  11. GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing 45 upvotes, #4 of 2025-03-14
  12. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency 27 upvotes, #9 of 2025-02-14
  13. Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step 31 upvotes, #4 of 2025-01-24
  14. IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models 13 upvotes, #8 of 2025-01-24
  15. EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation 45 upvotes, #1 of 2025-01-06
  16. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 35 upvotes, #3 of 2024-12-16
  17. EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM 21 upvotes, #7 of 2024-12-13
  18. StreamChat: Chatting with Streaming Video 17 upvotes, #9 of 2024-12-12
  19. BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices 39 upvotes, #2 of 2024-11-19
  20. PUMA: Empowering Unified MLLM with Multi-granular Visual Generation 51 upvotes, #4 of 2024-10-22
  21. Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow 16 upvotes, #9 of 2024-10-11
  22. MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code 44 upvotes, #2 of 2024-10-11
  23. MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines 33 upvotes, #3 of 2024-09-20
  24. LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation 19 upvotes, #4 of 2024-08-29
  25. GenCA: A Text-conditioned Generative Model for Realistic and Drivable Codec Avatars 17 upvotes, #7 of 2024-08-28
  26. Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining 29 upvotes, #2 of 2024-08-06
  27. AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents 29 upvotes, #3 of 2024-07-26
  28. MAVIS: Mathematical Visual Instruction Tuning 26 upvotes, #5 of 2024-07-12
  29. Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models 16 upvotes, #10 of 2024-06-18
  30. Phased Consistency Model 38 upvotes, #1 of 2024-05-29
  31. Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control 9 upvotes, #14 of 2024-05-28
  32. Urban Architect: Steerable 3D Urban Scene Generation with Layout Prior 9 upvotes, #7 of 2024-04-11
  33. MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? 45 upvotes, #1 of 2024-03-22
  34. Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation 10 upvotes, #13 of 2024-03-21
  35. FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis 6 upvotes, #10 of 2024-03-20
  36. GiT: Towards Generalist Vision Transformer through Universal Language Interface 23 upvotes, #4 of 2024-03-15
  37. SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models 17 upvotes, #8 of 2024-02-09
  38. Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling 39 upvotes, #4 of 2024-01-30
  39. SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models 14 upvotes, #7 of 2023-11-14
  40. ImageBind-LLM: Multi-modality Instruction Tuning 17 upvotes, #7 of 2023-09-08
  41. Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following 13 upvotes, #7 of 2023-09-04
  42. Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification 24 upvotes, #2 of 2023-08-16
  43. Tiny LVLM-eHub: Early Multimodal Experiments with Bard 11 upvotes, #10 of 2023-08-08
  44. Meta-Transformer: A Unified Framework for Multimodal Learning 45 upvotes, #2 of 2023-07-21
  45. JourneyDB: A Benchmark for Generative Image Understanding 20 upvotes, #5 of 2023-07-04
  46. Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising 4 upvotes, #8 of 2023-05-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.