Tencent

Tencent on Hugging Face Daily Papers: 119 papers, 19 in the top 3 of their day, 5 paper of the day.

  1. Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL 124 upvotes, #4 of 2026-10-02
  2. AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research 19 upvotes, #38 of 2026-09-29
  3. Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models 30 upvotes, #30 of 2026-09-29
  4. Draft-KV: Learning Useful Latent Communication Between Language Models 7 upvotes, #55 of 2026-09-29
  5. SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL 12 upvotes, #16 of 2026-09-28
  6. RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling 40 upvotes, #6 of 2026-09-24
  7. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay 130 upvotes, #4 of 2026-09-22
  8. OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation 37 upvotes, #8 of 2026-09-21
  9. WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing 47 upvotes, #12 of 2026-09-18
  10. Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge 19 upvotes, #14 of 2026-08-31
  11. ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL 27 upvotes, #13 of 2026-08-31
  12. GameWAM: A World Action Model for Video Games 45 upvotes, #8 of 2026-08-28
  13. WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report 68 upvotes, #3 of 2026-08-26
  14. Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs 35 upvotes, #5 of 2026-08-24
  15. FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving 19 upvotes, #10 of 2026-08-21
  16. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback 31 upvotes, #6 of 2026-08-21
  17. Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection 6 upvotes, #24 of 2026-08-19
  18. CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing 15 upvotes, #14 of 2026-08-19
  19. Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs 13 upvotes, #19 of 2026-08-18
  20. VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? 59 upvotes, #5 of 2026-08-18
  21. YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family 20 upvotes, #6 of 2026-08-10
  22. EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning 43 upvotes, #6 of 2026-08-07
  23. When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents 16 upvotes, #15 of 2026-08-06
  24. WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models 13 upvotes, #20 of 2026-08-06
  25. SkillJack: Persistent Skill Backdoors in Self-Evolving Agents 23 upvotes, #14 of 2026-08-05
  26. A New Role for Relevance: Guiding Corpus Interaction in Agentic Search 93 upvotes, #3 of 2026-07-29
  27. ReferTrack: Referring Then Tracking for Embodied Visual Tracking 53 upvotes, #3 of 2026-07-24
  28. Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction 26 upvotes, #8 of 2026-07-24
  29. Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking 31 upvotes, #3 of 2026-07-23
  30. EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World 92 upvotes, #3 of 2026-07-21
  31. When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers 6 upvotes, #28 of 2026-07-08
  32. HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better 7 upvotes, #25 of 2026-07-08
  33. Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process 37 upvotes, #8 of 2026-07-07
  34. Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming 14 upvotes, #7 of 2026-07-06
  35. SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History 43 upvotes, #6 of 2026-07-01
  36. ReFreeKV: Towards Threshold-Free KV Cache Compression 48 upvotes, #5 of 2026-06-30
  37. Exploring the Design Space of Reward Backpropagation for Flow Matching 10 upvotes, #24 of 2026-06-23
  38. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI 56 upvotes, #4 of 2026-06-15
  39. TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning 18 upvotes, #15 of 2026-06-11
  40. WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models 22 upvotes, #13 of 2026-06-09
  41. FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention 61 upvotes, #6 of 2026-06-09
  42. World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning 29 upvotes, #6 of 2026-06-03
  43. HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs 15 upvotes, #24 of 2026-05-28
  44. EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation 79 upvotes, #2 of 2026-05-27
  45. Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement 16 upvotes, #18 of 2026-05-27
  46. Toward Native Multimodal Modeling: A Roadmap 42 upvotes, #6 of 2026-05-26
  47. UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification 21 upvotes, #13 of 2026-05-11
  48. MiA-Signature: Approximating Global Activation for Long-Context Understanding 54 upvotes, #4 of 2026-05-08
  49. A^2TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping 14 upvotes, #15 of 2026-05-08
  50. Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language 21 upvotes, #9 of 2026-04-22
  51. PlayCoder: Making LLM-Generated GUI Code Playable 26 upvotes, #8 of 2026-04-22
  52. MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping 96 upvotes, #7 of 2026-04-10
  53. UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience 45 upvotes, #3 of 2026-03-26
  54. OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning 7 upvotes, #18 of 2026-03-26
  55. Deep Tabular Research via Continual Experience-Driven Execution 14 upvotes, #12 of 2026-03-23
  56. MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning 12 upvotes, #21 of 2026-03-20
  57. ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation 34 upvotes, #5 of 2026-03-13
  58. Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders 105 upvotes, #1 of 2026-03-09
  59. WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories 17 upvotes, #13 of 2026-03-03
  60. Query-focused and Memory-aware Reranker for Long Context Processing 55 upvotes, #2 of 2026-02-25
  61. The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context 13 upvotes, #18 of 2026-02-12
  62. Free(): Learning to Forget in Malloc-Only Reasoning Models 5 upvotes, #31 of 2026-02-12
  63. Covo-Audio Technical Report 8 upvotes, #30 of 2026-02-11
  64. Locas: Your Models are Principled Initializers of Locally-Supported Parametric Memories 4 upvotes, #40 of 2026-02-11
  65. OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention 12 upvotes, #16 of 2026-02-09
  66. ReMiT: RL-Guided Mid-Training for Iterative LLM Evolution 6 upvotes, #26 of 2026-02-09
  67. CL-bench: A Benchmark for Context Learning 22 upvotes, #16 of 2026-02-05
  68. No Global Plan in Chain-of-Thought: Uncover the Latent Planning Horizon of LLMs 68 upvotes, #3 of 2026-02-04
  69. Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation 25 upvotes, #12 of 2026-02-04
  70. Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning 16 upvotes, #27 of 2026-02-03
  71. Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning 9 upvotes, #10 of 2026-01-29
  72. Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision 41 upvotes, #4 of 2026-01-28
  73. Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification 20 upvotes, #7 of 2026-01-26
  74. Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning 16 upvotes, #8 of 2026-01-22
  75. TCAndon-Router: Adaptive Reasoning Router for Multi-Agent Collaboration 4 upvotes, #22 of 2026-01-12
  76. AT^2PO: Agentic Turn-based Policy Optimization via Tree Search 26 upvotes, #7 of 2026-01-09
  77. Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization 108 upvotes, #2 of 2026-01-05
  78. Improving Multi-step RAG with Hypergraph-based Memory for Long-Context Complex Relational Modeling 96 upvotes, #1 of 2026-01-02
  79. Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models 119 upvotes, #2 of 2026-01-01
  80. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection 13 upvotes, #17 of 2025-12-30
  81. SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents 38 upvotes, #7 of 2025-12-30
  82. Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding 108 upvotes, #1 of 2025-12-29
  83. Make-It-Poseable: Feed-forward Latent Posing Model for 3D Humanoid Character Animation 4 upvotes, #28 of 2025-12-19
  84. N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models 19 upvotes, #14 of 2025-12-19
  85. Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning 17 upvotes, #10 of 2025-12-18
  86. Distribution Matching Variational AutoEncoder 27 upvotes, #8 of 2025-12-09
  87. SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization 8 upvotes, #23 of 2025-12-05
  88. AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition 3 upvotes, #18 of 2025-12-04
  89. Guided Self-Evolving LLMs with Minimal Human Supervision 48 upvotes, #5 of 2025-12-03
  90. Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning 24 upvotes, #12 of 2025-12-02
  91. SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space 26 upvotes, #11 of 2025-11-26
  92. EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control 5 upvotes, #18 of 2025-11-21
  93. MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism 11 upvotes, #10 of 2025-11-17
  94. DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation 49 upvotes, #4 of 2025-11-11
  95. Too Good to be Bad: On the Failure of LLMs to Role-Play Villains 50 upvotes, #1 of 2025-11-10
  96. LTD-Bench: Evaluating Large Language Models by Letting Them Draw 8 upvotes, #10 of 2025-11-05
  97. Continuous Autoregressive Language Models 61 upvotes, #4 of 2025-11-03
  98. The End of Manual Decoding: Towards Truly End-to-End Language Models 113 upvotes, #1 of 2025-10-31
  99. ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks 17 upvotes, #12 of 2025-10-30
  100. Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values 18 upvotes, #8 of 2025-10-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.