Tencent
Tencent on Hugging Face Daily Papers: 119 papers, 19 in the top 3 of their day, 5 paper of the day.
- Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL 124 upvotes, #4 of 2026-10-02
- AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research 19 upvotes, #38 of 2026-09-29
- Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models 30 upvotes, #30 of 2026-09-29
- Draft-KV: Learning Useful Latent Communication Between Language Models 7 upvotes, #55 of 2026-09-29
- SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL 12 upvotes, #16 of 2026-09-28
- RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling 40 upvotes, #6 of 2026-09-24
- GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay 130 upvotes, #4 of 2026-09-22
- OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation 37 upvotes, #8 of 2026-09-21
- WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing 47 upvotes, #12 of 2026-09-18
- Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge 19 upvotes, #14 of 2026-08-31
- ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL 27 upvotes, #13 of 2026-08-31
- GameWAM: A World Action Model for Video Games 45 upvotes, #8 of 2026-08-28
- WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report 68 upvotes, #3 of 2026-08-26
- Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs 35 upvotes, #5 of 2026-08-24
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving 19 upvotes, #10 of 2026-08-21
- SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback 31 upvotes, #6 of 2026-08-21
- Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection 6 upvotes, #24 of 2026-08-19
- CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing 15 upvotes, #14 of 2026-08-19
- Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs 13 upvotes, #19 of 2026-08-18
- VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? 59 upvotes, #5 of 2026-08-18
- YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family 20 upvotes, #6 of 2026-08-10
- EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning 43 upvotes, #6 of 2026-08-07
- When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents 16 upvotes, #15 of 2026-08-06
- WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models 13 upvotes, #20 of 2026-08-06
- SkillJack: Persistent Skill Backdoors in Self-Evolving Agents 23 upvotes, #14 of 2026-08-05
- A New Role for Relevance: Guiding Corpus Interaction in Agentic Search 93 upvotes, #3 of 2026-07-29
- ReferTrack: Referring Then Tracking for Embodied Visual Tracking 53 upvotes, #3 of 2026-07-24
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction 26 upvotes, #8 of 2026-07-24
- Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking 31 upvotes, #3 of 2026-07-23
- EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World 92 upvotes, #3 of 2026-07-21
- When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers 6 upvotes, #28 of 2026-07-08
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better 7 upvotes, #25 of 2026-07-08
- Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process 37 upvotes, #8 of 2026-07-07
- Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming 14 upvotes, #7 of 2026-07-06
- SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History 43 upvotes, #6 of 2026-07-01
- ReFreeKV: Towards Threshold-Free KV Cache Compression 48 upvotes, #5 of 2026-06-30
- Exploring the Design Space of Reward Backpropagation for Flow Matching 10 upvotes, #24 of 2026-06-23
- From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI 56 upvotes, #4 of 2026-06-15
- TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning 18 upvotes, #15 of 2026-06-11
- WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models 22 upvotes, #13 of 2026-06-09
- FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention 61 upvotes, #6 of 2026-06-09
- World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning 29 upvotes, #6 of 2026-06-03
- HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs 15 upvotes, #24 of 2026-05-28
- EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation 79 upvotes, #2 of 2026-05-27
- Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement 16 upvotes, #18 of 2026-05-27
- Toward Native Multimodal Modeling: A Roadmap 42 upvotes, #6 of 2026-05-26
- UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification 21 upvotes, #13 of 2026-05-11
- MiA-Signature: Approximating Global Activation for Long-Context Understanding 54 upvotes, #4 of 2026-05-08
- A^2TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping 14 upvotes, #15 of 2026-05-08
- Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language 21 upvotes, #9 of 2026-04-22
- PlayCoder: Making LLM-Generated GUI Code Playable 26 upvotes, #8 of 2026-04-22
- MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping 96 upvotes, #7 of 2026-04-10
- UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience 45 upvotes, #3 of 2026-03-26
- OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning 7 upvotes, #18 of 2026-03-26
- Deep Tabular Research via Continual Experience-Driven Execution 14 upvotes, #12 of 2026-03-23
- MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning 12 upvotes, #21 of 2026-03-20
- ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation 34 upvotes, #5 of 2026-03-13
- Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders 105 upvotes, #1 of 2026-03-09
- WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories 17 upvotes, #13 of 2026-03-03
- Query-focused and Memory-aware Reranker for Long Context Processing 55 upvotes, #2 of 2026-02-25
- The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context 13 upvotes, #18 of 2026-02-12
- Free(): Learning to Forget in Malloc-Only Reasoning Models 5 upvotes, #31 of 2026-02-12
- Covo-Audio Technical Report 8 upvotes, #30 of 2026-02-11
- Locas: Your Models are Principled Initializers of Locally-Supported Parametric Memories 4 upvotes, #40 of 2026-02-11
- OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention 12 upvotes, #16 of 2026-02-09
- ReMiT: RL-Guided Mid-Training for Iterative LLM Evolution 6 upvotes, #26 of 2026-02-09
- CL-bench: A Benchmark for Context Learning 22 upvotes, #16 of 2026-02-05
- No Global Plan in Chain-of-Thought: Uncover the Latent Planning Horizon of LLMs 68 upvotes, #3 of 2026-02-04
- Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation 25 upvotes, #12 of 2026-02-04
- Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning 16 upvotes, #27 of 2026-02-03
- Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning 9 upvotes, #10 of 2026-01-29
- Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision 41 upvotes, #4 of 2026-01-28
- Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification 20 upvotes, #7 of 2026-01-26
- Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning 16 upvotes, #8 of 2026-01-22
- TCAndon-Router: Adaptive Reasoning Router for Multi-Agent Collaboration 4 upvotes, #22 of 2026-01-12
- AT^2PO: Agentic Turn-based Policy Optimization via Tree Search 26 upvotes, #7 of 2026-01-09
- Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization 108 upvotes, #2 of 2026-01-05
- Improving Multi-step RAG with Hypergraph-based Memory for Long-Context Complex Relational Modeling 96 upvotes, #1 of 2026-01-02
- Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models 119 upvotes, #2 of 2026-01-01
- YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection 13 upvotes, #17 of 2025-12-30
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents 38 upvotes, #7 of 2025-12-30
- Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding 108 upvotes, #1 of 2025-12-29
- Make-It-Poseable: Feed-forward Latent Posing Model for 3D Humanoid Character Animation 4 upvotes, #28 of 2025-12-19
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models 19 upvotes, #14 of 2025-12-19
- Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning 17 upvotes, #10 of 2025-12-18
- Distribution Matching Variational AutoEncoder 27 upvotes, #8 of 2025-12-09
- SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization 8 upvotes, #23 of 2025-12-05
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition 3 upvotes, #18 of 2025-12-04
- Guided Self-Evolving LLMs with Minimal Human Supervision 48 upvotes, #5 of 2025-12-03
- Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning 24 upvotes, #12 of 2025-12-02
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space 26 upvotes, #11 of 2025-11-26
- EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control 5 upvotes, #18 of 2025-11-21
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism 11 upvotes, #10 of 2025-11-17
- DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation 49 upvotes, #4 of 2025-11-11
- Too Good to be Bad: On the Failure of LLMs to Role-Play Villains 50 upvotes, #1 of 2025-11-10
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw 8 upvotes, #10 of 2025-11-05
- Continuous Autoregressive Language Models 61 upvotes, #4 of 2025-11-03
- The End of Manual Decoding: Towards Truly End-to-End Language Models 113 upvotes, #1 of 2025-10-31
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks 17 upvotes, #12 of 2025-10-30
- Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values 18 upvotes, #8 of 2025-10-24
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.