Renrui
Renrui on Hugging Face Daily Papers: 11 papers, 2 in the top 3 of their day, 266 upvotes.
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation 15 upvotes, #12 of 2025-11-21
- MAVIS: Mathematical Visual Instruction Tuning 26 upvotes, #5 of 2024-07-12
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 34 upvotes, #3 of 2024-07-11
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? 45 upvotes, #1 of 2024-03-22
- SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models 17 upvotes, #8 of 2024-02-09
- A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise 14 upvotes, #5 of 2023-12-20
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models 14 upvotes, #7 of 2023-11-14
- MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning 30 upvotes, #5 of 2023-10-06
- ImageBind-LLM: Multi-modality Instruction Tuning 17 upvotes, #7 of 2023-09-08
- Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following 13 upvotes, #7 of 2023-09-04
- JourneyDB: A Benchmark for Generative Image Understanding 20 upvotes, #5 of 2023-07-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.