Daily Papers of 2025-04-11

  1. Kimi-VL Technical Report 113 upvotes, #1 of 2025-04-11
  2. DeepSeek-R1 Thoughtology: Let's <think> about LLM Reasoning 79 upvotes, #2 of 2025-04-11
  3. C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing 58 upvotes, #3 of 2025-04-11
  4. VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning 43 upvotes, #4 of 2025-04-11
  5. VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning 40 upvotes, #5 of 2025-04-11
  6. MM-IFEngine: Towards Multimodal Instruction Following 31 upvotes, #6 of 2025-04-11
  7. HoloPart: Generative 3D Part Amodal Segmentation 25 upvotes, #7 of 2025-04-11
  8. Scaling Laws for Native Multimodal Models Scaling Laws for Native Multimodal Models 24 upvotes, #8 of 2025-04-11
  9. MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations 15 upvotes, #9 of 2025-04-11
  10. SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement 14 upvotes, #10 of 2025-04-11
  11. Towards Visual Text Grounding of Multimodal Large Language Model 13 upvotes, #11 of 2025-04-11
  12. Compass Control: Multi Object Orientation Control for Text-to-Image Generation 7 upvotes, #12 of 2025-04-11
  13. TAPNext: Tracking Any Point (TAP) as Next Token Prediction 4 upvotes, #13 of 2025-04-11
  14. MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detection 3 upvotes, #14 of 2025-04-11
  15. Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction 3 upvotes, #14 of 2025-04-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.