XING SUN
XING SUN on Hugging Face Daily Papers: 35 papers, 7 in the top 3 of their day, 1,419 upvotes.
- Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge 19 upvotes, #14 of 2026-08-31
- ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL 27 upvotes, #13 of 2026-08-31
- Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios 16 upvotes, #14 of 2026-08-27
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction 26 upvotes, #8 of 2026-07-24
- Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process 37 upvotes, #8 of 2026-07-07
- From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI 56 upvotes, #4 of 2026-06-15
- Toward Native Multimodal Modeling: A Roadmap 42 upvotes, #6 of 2026-05-26
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding 232 upvotes, #1 of 2026-04-08
- HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention 40 upvotes, #6 of 2026-03-31
- Deep Tabular Research via Continual Experience-Driven Execution 14 upvotes, #12 of 2026-03-23
- MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning 12 upvotes, #21 of 2026-03-20
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion 48 upvotes, #4 of 2026-03-11
- ReMiT: RL-Guided Mid-Training for Iterative LLM Evolution 6 upvotes, #26 of 2026-02-09
- RISE-Video: Can Video Generators Decode Implicit World Rules? 26 upvotes, #9 of 2026-02-06
- Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision 41 upvotes, #4 of 2026-01-28
- Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization 108 upvotes, #2 of 2026-01-05
- Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models 119 upvotes, #2 of 2026-01-01
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents 38 upvotes, #7 of 2025-12-30
- Streaming Video Instruction Tuning 9 upvotes, #13 of 2025-12-25
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space 26 upvotes, #11 of 2025-11-26
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw 8 upvotes, #10 of 2025-11-05
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting 41 upvotes, #5 of 2025-10-28
- Training-Free Group Relative Policy Optimization 40 upvotes, #11 of 2025-10-10
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning 10 upvotes, #26 of 2025-09-29
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill \& Decode Inference 7 upvotes, #11 of 2025-08-25
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models 15 upvotes, #15 of 2025-06-03
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model 8 upvotes, #12 of 2025-05-07
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction 36 upvotes, #2 of 2025-01-06
- MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs 18 upvotes, #4 of 2024-11-27
- Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models 19 upvotes, #4 of 2024-08-29
- VITA: Towards Open-Source Interactive Omni Multimodal LLM 40 upvotes, #1 of 2024-08-12
- Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models 14 upvotes, #7 of 2024-08-06
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 13 upvotes, #3 of 2024-06-03
- A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise 14 upvotes, #5 of 2023-12-20
- Woodpecker: Hallucination Correction for Multimodal Large Language Models 18 upvotes, #3 of 2023-10-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.