Tsinghua University

Tsinghua University on Hugging Face Daily Papers: 59 papers, 6 in the top 3 of their day, 4 paper of the day.

  1. RoboCoach: World Models as Active Coaches for Compositional Robot Skills 13 upvotes, #39 of 2026-10-01
  2. One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices 4 upvotes, #84 of 2026-09-30
  3. CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents 14 upvotes, #19 of 2026-09-17
  4. The Router Within: Eliciting Native Skill Routing from a Frozen LLM 7 upvotes, #20 of 2026-09-16
  5. Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation 695 upvotes, #1 of 2026-09-15
  6. Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering 29 upvotes, #11 of 2026-09-03
  7. Agentic Transaction: Towards ACID-Compliant Agent Systems 27 upvotes, #13 of 2026-08-18
  8. Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression 13 upvotes, #13 of 2026-08-10
  9. Uncertainty-Aware World Model for Aerial Image-Goal Navigation 12 upvotes, #15 of 2026-08-10
  10. AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 94 upvotes, #1 of 2026-08-07
  11. Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation 5 upvotes, #31 of 2026-08-05
  12. Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories 8 upvotes, #25 of 2026-08-05
  13. ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts 7 upvotes, #27 of 2026-08-05
  14. Vidu S1: A Real-Time Interactive Video Generation Model 138 upvotes, #1 of 2026-07-10
  15. AgenticDataBench: A Comprehensive Benchmark for Data Agents 35 upvotes, #6 of 2026-07-03
  16. CausalMix: Data Mixture as Causal Inference for Language Model Training 19 upvotes, #10 of 2026-07-02
  17. TurboServe: Serving Streaming Video Generation Efficiently and Economically 34 upvotes, #2 of 2026-07-02
  18. LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing 82 upvotes, #4 of 2026-06-30
  19. SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 42 upvotes, #8 of 2026-06-09
  20. Parametric Social Identity Injection and Diversification in Public Opinion Simulation 1 upvotes, #39 of 2026-06-08
  21. Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation 16 upvotes, #18 of 2026-06-04
  22. Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents 3 upvotes, #34 of 2026-06-04
  23. Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning 16 upvotes, #18 of 2026-05-27
  24. SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking 8 upvotes, #34 of 2026-05-26
  25. Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking 6 upvotes, #34 of 2026-05-22
  26. Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection 2 upvotes, #38 of 2026-05-21
  27. Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization 2 upvotes, #42 of 2026-05-14
  28. KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels 7 upvotes, #23 of 2026-05-08
  29. Mobile GUI Agents under Real-world Threats: Are We There Yet? 3 upvotes, #28 of 2026-04-16
  30. TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction 2 upvotes, #40 of 2026-04-14
  31. 6Bit-Diffusion: Inference-Time Mixed-Precision Quantization for Video Diffusion Models 10 upvotes, #13 of 2026-03-26
  32. Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning 7 upvotes, #18 of 2026-03-26
  33. BEAVER: A Training-Free Hierarchical Prompt Compression Method via Structure-Aware Page Selection 11 upvotes, #15 of 2026-03-23
  34. CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges 6 upvotes, #18 of 2026-03-16
  35. How Far Can Unsupervised RLVR Scale LLM Training? 52 upvotes, #4 of 2026-03-10
  36. IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation 1 upvotes, #22 of 2026-03-09
  37. Imagination Helps Visual Reasoning, But Not Yet in Latent Space 39 upvotes, #5 of 2026-02-27
  38. JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments 2 upvotes, #20 of 2026-02-26
  39. SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning 43 upvotes, #3 of 2026-02-20
  40. BiManiBench: A Hierarchical Benchmark for Evaluating Bimanual Coordination of Multimodal Large Language Models 3 upvotes, #19 of 2026-02-19
  41. Geometry-Aware Rotary Position Embedding for Consistent Video World Model 9 upvotes, #10 of 2026-02-18
  42. Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings 9 upvotes, #14 of 2026-02-17
  43. Context Learning for Multi-Agent Discussion 4 upvotes, #39 of 2026-02-05
  44. ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding 3 upvotes, #32 of 2026-02-02
  45. Continual GUI Agents 4 upvotes, #28 of 2026-02-02
  46. E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models 8 upvotes, #10 of 2026-01-08
  47. JustRL: Scaling a 1.5B LLM with a Simple RL Recipe 22 upvotes, #13 of 2025-12-19
  48. Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task 3 upvotes, #19 of 2025-12-12
  49. MIND-V: Hierarchical Video Generation for Long-Horizon Robotic Manipulation with RL-based Physical Alignment 12 upvotes, #10 of 2025-12-10
  50. Controllable Layer Decomposition for Reversible Multi-Layer Image Generation 8 upvotes, #18 of 2025-11-25
  51. MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning 11 upvotes, #9 of 2025-11-13
  52. WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation 12 upvotes, #8 of 2025-11-13
  53. RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging 11 upvotes, #12 of 2025-10-27
  54. Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views 20 upvotes, #11 of 2025-10-22
  55. CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving 4 upvotes, #26 of 2025-10-16
  56. Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models 12 upvotes, #24 of 2025-10-01
  57. Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention 13 upvotes, #23 of 2025-10-01
  58. Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR 47 upvotes, #5 of 2025-09-30
  59. SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention 109 upvotes, #1 of 2025-09-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.