YannQi

YannQi on Hugging Face Daily Papers: 10 papers, 1 in the top 3 of their day, 253 upvotes.

  1. How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining 65 upvotes, #11 of 2026-09-29
  2. CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering 2 upvotes, #35 of 2026-03-03
  3. MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning 2 upvotes, #44 of 2026-02-04
  4. HunyuanOCR Technical Report 19 upvotes, #12 of 2025-11-26
  5. Taming Modality Entanglement in Continual Audio-Visual Segmentation 3 upvotes, #20 of 2025-10-27
  6. Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering 3 upvotes, #25 of 2025-10-21
  7. R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning 105 upvotes, #3 of 2025-09-01
  8. Continuous Speculative Decoding for Autoregressive Image Generation 14 upvotes, #4 of 2024-11-20
  9. Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis 14 upvotes, #5 of 2024-09-11
  10. AVESFormer: Efficient Transformer Design for Real-Time Audio-Visual Segmentation 2 upvotes, #12 of 2024-08-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.