Hongsheng LI
Hongsheng LI on Hugging Face Daily Papers: 46 papers, 14 in the top 3 of their day, 1,313 upvotes.
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents 38 upvotes, #12 of 2025-05-28
- EnerVerse-AC: Envisioning Embodied Environments with Action Condition 21 upvotes, #7 of 2025-05-16
- WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch 16 upvotes, #11 of 2025-05-13
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT 39 upvotes, #3 of 2025-05-02
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects 21 upvotes, #3 of 2025-04-29
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning 15 upvotes, #14 of 2025-04-23
- UI-R1: Enhancing Action Prediction of GUI Agents by Reinforcement Learning 54 upvotes, #3 of 2025-03-28
- LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis 25 upvotes, #6 of 2025-03-28
- Lumina-Image 2.0: A Unified and Efficient Image Generative Framework 18 upvotes, #11 of 2025-03-28
- Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning 33 upvotes, #4 of 2025-03-17
- GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing 45 upvotes, #4 of 2025-03-14
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency 27 upvotes, #9 of 2025-02-14
- Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step 31 upvotes, #4 of 2025-01-24
- IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models 13 upvotes, #8 of 2025-01-24
- EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation 45 upvotes, #1 of 2025-01-06
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding 35 upvotes, #3 of 2024-12-16
- EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM 21 upvotes, #7 of 2024-12-13
- StreamChat: Chatting with Streaming Video 17 upvotes, #9 of 2024-12-12
- BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices 39 upvotes, #2 of 2024-11-19
- PUMA: Empowering Unified MLLM with Multi-granular Visual Generation 51 upvotes, #4 of 2024-10-22
- Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow 16 upvotes, #9 of 2024-10-11
- MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code 44 upvotes, #2 of 2024-10-11
- MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines 33 upvotes, #3 of 2024-09-20
- LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation 19 upvotes, #4 of 2024-08-29
- GenCA: A Text-conditioned Generative Model for Realistic and Drivable Codec Avatars 17 upvotes, #7 of 2024-08-28
- Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining 29 upvotes, #2 of 2024-08-06
- AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents 29 upvotes, #3 of 2024-07-26
- MAVIS: Mathematical Visual Instruction Tuning 26 upvotes, #5 of 2024-07-12
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models 16 upvotes, #10 of 2024-06-18
- Phased Consistency Model 38 upvotes, #1 of 2024-05-29
- Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control 9 upvotes, #14 of 2024-05-28
- Urban Architect: Steerable 3D Urban Scene Generation with Layout Prior 9 upvotes, #7 of 2024-04-11
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? 45 upvotes, #1 of 2024-03-22
- Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation 10 upvotes, #13 of 2024-03-21
- FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis 6 upvotes, #10 of 2024-03-20
- GiT: Towards Generalist Vision Transformer through Universal Language Interface 23 upvotes, #4 of 2024-03-15
- SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models 17 upvotes, #8 of 2024-02-09
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling 39 upvotes, #4 of 2024-01-30
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models 14 upvotes, #7 of 2023-11-14
- ImageBind-LLM: Multi-modality Instruction Tuning 17 upvotes, #7 of 2023-09-08
- Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following 13 upvotes, #7 of 2023-09-04
- Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification 24 upvotes, #2 of 2023-08-16
- Tiny LVLM-eHub: Early Multimodal Experiments with Bard 11 upvotes, #10 of 2023-08-08
- Meta-Transformer: A Unified Framework for Multimodal Learning 45 upvotes, #2 of 2023-07-21
- JourneyDB: A Benchmark for Generative Image Understanding 20 upvotes, #5 of 2023-07-04
- Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising 4 upvotes, #8 of 2023-05-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.