XING SUN

XING SUN on Hugging Face Daily Papers: 35 papers, 7 in the top 3 of their day, 1,419 upvotes.

  1. Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge 19 upvotes, #14 of 2026-08-31
  2. ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL 27 upvotes, #13 of 2026-08-31
  3. Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios 16 upvotes, #14 of 2026-08-27
  4. Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction 26 upvotes, #8 of 2026-07-24
  5. Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process 37 upvotes, #8 of 2026-07-07
  6. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI 56 upvotes, #4 of 2026-06-15
  7. Toward Native Multimodal Modeling: A Roadmap 42 upvotes, #6 of 2026-05-26
  8. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 232 upvotes, #1 of 2026-04-08
  9. HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention 40 upvotes, #6 of 2026-03-31
  10. Deep Tabular Research via Continual Experience-Driven Execution 14 upvotes, #12 of 2026-03-23
  11. MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning 12 upvotes, #21 of 2026-03-20
  12. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion 48 upvotes, #4 of 2026-03-11
  13. ReMiT: RL-Guided Mid-Training for Iterative LLM Evolution 6 upvotes, #26 of 2026-02-09
  14. RISE-Video: Can Video Generators Decode Implicit World Rules? 26 upvotes, #9 of 2026-02-06
  15. Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision 41 upvotes, #4 of 2026-01-28
  16. Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization 108 upvotes, #2 of 2026-01-05
  17. Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models 119 upvotes, #2 of 2026-01-01
  18. SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents 38 upvotes, #7 of 2025-12-30
  19. Streaming Video Instruction Tuning 9 upvotes, #13 of 2025-12-25
  20. SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space 26 upvotes, #11 of 2025-11-26
  21. LTD-Bench: Evaluating Large Language Models by Letting Them Draw 8 upvotes, #10 of 2025-11-05
  22. VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting 41 upvotes, #5 of 2025-10-28
  23. Training-Free Group Relative Policy Optimization 40 upvotes, #11 of 2025-10-10
  24. Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning 10 upvotes, #26 of 2025-09-29
  25. TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill \& Decode Inference 7 upvotes, #11 of 2025-08-25
  26. Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models 15 upvotes, #15 of 2025-06-03
  27. VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model 8 upvotes, #12 of 2025-05-07
  28. VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction 36 upvotes, #2 of 2025-01-06
  29. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
  30. Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models 19 upvotes, #4 of 2024-08-29
  31. VITA: Towards Open-Source Interactive Omni Multimodal LLM 40 upvotes, #1 of 2024-08-12
  32. Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models 14 upvotes, #7 of 2024-08-06
  33. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 13 upvotes, #3 of 2024-06-03
  34. A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise 14 upvotes, #5 of 2023-12-20
  35. Woodpecker: Hallucination Correction for Multimodal Large Language Models 18 upvotes, #3 of 2023-10-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.