Tsinghua University
Tsinghua University on Hugging Face Daily Papers: 59 papers, 6 in the top 3 of their day, 4 paper of the day.
- RoboCoach: World Models as Active Coaches for Compositional Robot Skills 13 upvotes, #39 of 2026-10-01
- One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices 4 upvotes, #84 of 2026-09-30
- CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents 14 upvotes, #19 of 2026-09-17
- The Router Within: Eliciting Native Skill Routing from a Frozen LLM 7 upvotes, #20 of 2026-09-16
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation 695 upvotes, #1 of 2026-09-15
- Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering 29 upvotes, #11 of 2026-09-03
- Agentic Transaction: Towards ACID-Compliant Agent Systems 27 upvotes, #13 of 2026-08-18
- Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression 13 upvotes, #13 of 2026-08-10
- Uncertainty-Aware World Model for Aerial Image-Goal Navigation 12 upvotes, #15 of 2026-08-10
- AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 94 upvotes, #1 of 2026-08-07
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation 5 upvotes, #31 of 2026-08-05
- Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories 8 upvotes, #25 of 2026-08-05
- ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts 7 upvotes, #27 of 2026-08-05
- Vidu S1: A Real-Time Interactive Video Generation Model 138 upvotes, #1 of 2026-07-10
- AgenticDataBench: A Comprehensive Benchmark for Data Agents 35 upvotes, #6 of 2026-07-03
- CausalMix: Data Mixture as Causal Inference for Language Model Training 19 upvotes, #10 of 2026-07-02
- TurboServe: Serving Streaming Video Generation Efficiently and Economically 34 upvotes, #2 of 2026-07-02
- LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing 82 upvotes, #4 of 2026-06-30
- SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 42 upvotes, #8 of 2026-06-09
- Parametric Social Identity Injection and Diversification in Public Opinion Simulation 1 upvotes, #39 of 2026-06-08
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation 16 upvotes, #18 of 2026-06-04
- Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents 3 upvotes, #34 of 2026-06-04
- Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning 16 upvotes, #18 of 2026-05-27
- SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking 8 upvotes, #34 of 2026-05-26
- Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking 6 upvotes, #34 of 2026-05-22
- Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection 2 upvotes, #38 of 2026-05-21
- Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization 2 upvotes, #42 of 2026-05-14
- KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels 7 upvotes, #23 of 2026-05-08
- Mobile GUI Agents under Real-world Threats: Are We There Yet? 3 upvotes, #28 of 2026-04-16
- TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction 2 upvotes, #40 of 2026-04-14
- 6Bit-Diffusion: Inference-Time Mixed-Precision Quantization for Video Diffusion Models 10 upvotes, #13 of 2026-03-26
- Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning 7 upvotes, #18 of 2026-03-26
- BEAVER: A Training-Free Hierarchical Prompt Compression Method via Structure-Aware Page Selection 11 upvotes, #15 of 2026-03-23
- CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges 6 upvotes, #18 of 2026-03-16
- How Far Can Unsupervised RLVR Scale LLM Training? 52 upvotes, #4 of 2026-03-10
- IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation 1 upvotes, #22 of 2026-03-09
- Imagination Helps Visual Reasoning, But Not Yet in Latent Space 39 upvotes, #5 of 2026-02-27
- JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments 2 upvotes, #20 of 2026-02-26
- SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning 43 upvotes, #3 of 2026-02-20
- BiManiBench: A Hierarchical Benchmark for Evaluating Bimanual Coordination of Multimodal Large Language Models 3 upvotes, #19 of 2026-02-19
- Geometry-Aware Rotary Position Embedding for Consistent Video World Model 9 upvotes, #10 of 2026-02-18
- Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings 9 upvotes, #14 of 2026-02-17
- Context Learning for Multi-Agent Discussion 4 upvotes, #39 of 2026-02-05
- ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding 3 upvotes, #32 of 2026-02-02
- Continual GUI Agents 4 upvotes, #28 of 2026-02-02
- E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models 8 upvotes, #10 of 2026-01-08
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task 3 upvotes, #19 of 2025-12-12
- MIND-V: Hierarchical Video Generation for Long-Horizon Robotic Manipulation with RL-based Physical Alignment 12 upvotes, #10 of 2025-12-10
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation 8 upvotes, #18 of 2025-11-25
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning 11 upvotes, #9 of 2025-11-13
- WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation 12 upvotes, #8 of 2025-11-13
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging 11 upvotes, #12 of 2025-10-27
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views 20 upvotes, #11 of 2025-10-22
- CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving 4 upvotes, #26 of 2025-10-16
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models 12 upvotes, #24 of 2025-10-01
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention 13 upvotes, #23 of 2025-10-01
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR 47 upvotes, #5 of 2025-09-30
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention 109 upvotes, #1 of 2025-09-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.