Peking University

Peking University on Hugging Face Daily Papers: 91 papers, 13 in the top 3 of their day, 6 paper of the day.

  1. DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation 6 upvotes, #73 of 2026-10-02
  2. Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection 7 upvotes, #61 of 2026-09-30
  3. Post-Training Leaves Behavioral Shadows on Unrelated Decisions 271 upvotes, #1 of 2026-09-29
  4. Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning 2 upvotes, #93 of 2026-09-29
  5. RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation 55 upvotes, #3 of 2026-09-28
  6. DeltaWAM: Delta World Action Models for Bimanual Manipulation 19 upvotes, #13 of 2026-09-25
  7. Harness-Zero: Harness Distillation via Agent-as-Harness 37 upvotes, #11 of 2026-09-22
  8. OmniEdu: Open Foundation Models for Learning and Teaching 233 upvotes, #1 of 2026-09-22
  9. ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation 1 upvotes, #24 of 2026-09-14
  10. MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control 8 upvotes, #13 of 2026-09-14
  11. DataFlex-RL: An Evaluation Platform for RLVR Data Policies 162 upvotes, #2 of 2026-09-14
  12. Length-Adaptive Decoding for Masked Diffusion Machine Translation 4 upvotes, #21 of 2026-08-26
  13. The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search 4 upvotes, #28 of 2026-08-25
  14. Verifier-Induced Support Reshaping in On-Policy Optimization 5 upvotes, #23 of 2026-08-17
  15. ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents 10 upvotes, #13 of 2026-08-13
  16. ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? 14 upvotes, #20 of 2026-08-05
  17. UniWorld-Design: From Pixel Generation to Layer-Native Design 21 upvotes, #15 of 2026-08-05
  18. MiniWorld: Democratizing the Training of Video World Models from Scratch 19 upvotes, #16 of 2026-08-05
  19. Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts 8 upvotes, #33 of 2026-08-04
  20. DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents 17 upvotes, #23 of 2026-08-04
  21. Flux-OPD: On-Policy Distillation with Evolving Contexts 44 upvotes, #10 of 2026-07-31
  22. Data Pyramid for Embodied Manipulation 36 upvotes, #7 of 2026-07-28
  23. DataPrep-Bench: Benchmarking LLMs as Training Data Preparators 55 upvotes, #1 of 2026-07-27
  24. K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs 63 upvotes, #2 of 2026-07-24
  25. DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines 137 upvotes, #2 of 2026-07-22
  26. SciForma: Structure-Faithful Generation of Scientific Diagrams 22 upvotes, #10 of 2026-07-22
  27. VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery 10 upvotes, #12 of 2026-07-13
  28. Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation 6 upvotes, #19 of 2026-06-29
  29. GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning 3 upvotes, #10 of 2026-06-22
  30. DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects 72 upvotes, #2 of 2026-06-19
  31. MotionVLA: Vision-Language-Action Model for Humanoid Motion 4 upvotes, #25 of 2026-06-17
  32. Watch, Remember, Reason: Human-View Video Understanding with MLLMs 21 upvotes, #12 of 2026-06-08
  33. The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs 7 upvotes, #22 of 2026-06-05
  34. LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing 24 upvotes, #8 of 2026-06-05
  35. AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding 10 upvotes, #16 of 2026-06-05
  36. OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning 24 upvotes, #16 of 2026-05-28
  37. Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining 144 upvotes, #1 of 2026-05-21
  38. RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting 9 upvotes, #23 of 2026-05-20
  39. StableVLA: Towards Robust Vision-Language-Action Models without Extra Data 15 upvotes, #18 of 2026-05-19
  40. GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding 4 upvotes, #36 of 2026-05-18
  41. VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction 26 upvotes, #15 of 2026-05-15
  42. PresentAgent-2: Towards Generalist Multimodal Presentation Agents 8 upvotes, #22 of 2026-05-14
  43. Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction 1 upvotes, #56 of 2026-05-13
  44. From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills 21 upvotes, #5 of 2026-05-04
  45. Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models 50 upvotes, #4 of 2026-04-30
  46. UniMesh: Unifying 3D Mesh Understanding and Generation 10 upvotes, #16 of 2026-04-22
  47. HSG: Hyperbolic Scene Graph 1 upvotes, #49 of 2026-04-21
  48. Context-Value-Action Architecture for Value-Driven Large Language Model Agents 8 upvotes, #23 of 2026-04-08
  49. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models 200 upvotes, #2 of 2026-04-07
  50. DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models 345 upvotes, #1 of 2026-04-03
  51. HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention 40 upvotes, #6 of 2026-03-31
  52. Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models 19 upvotes, #9 of 2026-03-30
  53. PEARL: Personalized Streaming Video Understanding Model 40 upvotes, #6 of 2026-03-25
  54. FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow 32 upvotes, #7 of 2026-03-23
  55. Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models 21 upvotes, #12 of 2026-03-19
  56. MWM: Mobile World Models for Action-Conditioned Consistent Prediction 0 upvotes, #48 of 2026-03-10
  57. StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation 0 upvotes, #23 of 2026-02-20
  58. MMA: Multimodal Memory Agent 8 upvotes, #14 of 2026-02-19
  59. MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation 3 upvotes, #21 of 2026-02-17
  60. Code2Worlds: Empowering Coding LLMs for 4D World Generation 4 upvotes, #18 of 2026-02-16
  61. Light4D: Training-Free Extreme Viewpoint 4D Video Relighting 2 upvotes, #23 of 2026-02-16
  62. GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning 1 upvotes, #27 of 2026-02-16
  63. TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions 27 upvotes, #8 of 2026-02-12
  64. PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss 41 upvotes, #10 of 2026-02-03
  65. 3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence 1 upvotes, #39 of 2026-01-13
  66. AnyDepth: Depth Estimation Made Easy 9 upvotes, #14 of 2026-01-12
  67. DocDancer: Towards Agentic Document-Grounded Information Seeking 4 upvotes, #19 of 2026-01-09
  68. MDAgent2: Large Language Model for Code Generation and Knowledge Q&A in Molecular Dynamics 7 upvotes, #11 of 2026-01-08
  69. Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations 12 upvotes, #11 of 2025-12-25
  70. DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI 193 upvotes, #1 of 2025-12-23
  71. Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation 1 upvotes, #30 of 2025-12-18
  72. VABench: A Comprehensive Benchmark for Audio-Video Generation 7 upvotes, #19 of 2025-12-18
  73. Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling 40 upvotes, #4 of 2025-12-17
  74. DragMesh: Interactive 3D Generation Made Easy 1 upvotes, #25 of 2025-12-12
  75. From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs 23 upvotes, #6 of 2025-12-10
  76. EgoLCD: Egocentric Video Generation with Long Context Diffusion 5 upvotes, #29 of 2025-12-05
  77. Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation 13 upvotes, #18 of 2025-12-03
  78. MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots 5 upvotes, #12 of 2025-11-27
  79. Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward 31 upvotes, #7 of 2025-11-26
  80. DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation 62 upvotes, #3 of 2025-11-25
  81. EvoVLA: Self-Evolving Vision-Language-Action Model 4 upvotes, #24 of 2025-11-25
  82. Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks 10 upvotes, #19 of 2025-10-30
  83. Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence 11 upvotes, #14 of 2025-10-24
  84. Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback 18 upvotes, #13 of 2025-10-21
  85. Universal Image Restoration Pre-training via Masked Degradation Classification 10 upvotes, #19 of 2025-10-16
  86. WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation 5 upvotes, #26 of 2025-10-09
  87. StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes 1 upvotes, #33 of 2025-09-23
  88. Nav-R1: Reasoning and Navigation in Embodied Scenes 6 upvotes, #10 of 2025-09-16
  89. 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding 15 upvotes, #4 of 2025-08-04
  90. TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios 2 upvotes, #41 of 2025-05-26
  91. TransMLA: Multi-head Latent Attention Is All You Need 41 upvotes, #4 of 2025-02-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.