Peking University
Peking University on Hugging Face Daily Papers: 91 papers, 13 in the top 3 of their day, 6 paper of the day.
- DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation 6 upvotes, #73 of 2026-10-02
- Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection 7 upvotes, #61 of 2026-09-30
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions 271 upvotes, #1 of 2026-09-29
- Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning 2 upvotes, #93 of 2026-09-29
- RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation 55 upvotes, #3 of 2026-09-28
- DeltaWAM: Delta World Action Models for Bimanual Manipulation 19 upvotes, #13 of 2026-09-25
- Harness-Zero: Harness Distillation via Agent-as-Harness 37 upvotes, #11 of 2026-09-22
- OmniEdu: Open Foundation Models for Learning and Teaching 233 upvotes, #1 of 2026-09-22
- ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation 1 upvotes, #24 of 2026-09-14
- MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control 8 upvotes, #13 of 2026-09-14
- DataFlex-RL: An Evaluation Platform for RLVR Data Policies 162 upvotes, #2 of 2026-09-14
- Length-Adaptive Decoding for Masked Diffusion Machine Translation 4 upvotes, #21 of 2026-08-26
- The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search 4 upvotes, #28 of 2026-08-25
- Verifier-Induced Support Reshaping in On-Policy Optimization 5 upvotes, #23 of 2026-08-17
- ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents 10 upvotes, #13 of 2026-08-13
- ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? 14 upvotes, #20 of 2026-08-05
- UniWorld-Design: From Pixel Generation to Layer-Native Design 21 upvotes, #15 of 2026-08-05
- MiniWorld: Democratizing the Training of Video World Models from Scratch 19 upvotes, #16 of 2026-08-05
- Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts 8 upvotes, #33 of 2026-08-04
- DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents 17 upvotes, #23 of 2026-08-04
- Flux-OPD: On-Policy Distillation with Evolving Contexts 44 upvotes, #10 of 2026-07-31
- Data Pyramid for Embodied Manipulation 36 upvotes, #7 of 2026-07-28
- DataPrep-Bench: Benchmarking LLMs as Training Data Preparators 55 upvotes, #1 of 2026-07-27
- K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs 63 upvotes, #2 of 2026-07-24
- DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines 137 upvotes, #2 of 2026-07-22
- SciForma: Structure-Faithful Generation of Scientific Diagrams 22 upvotes, #10 of 2026-07-22
- VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery 10 upvotes, #12 of 2026-07-13
- Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation 6 upvotes, #19 of 2026-06-29
- GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning 3 upvotes, #10 of 2026-06-22
- DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects 72 upvotes, #2 of 2026-06-19
- MotionVLA: Vision-Language-Action Model for Humanoid Motion 4 upvotes, #25 of 2026-06-17
- Watch, Remember, Reason: Human-View Video Understanding with MLLMs 21 upvotes, #12 of 2026-06-08
- The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs 7 upvotes, #22 of 2026-06-05
- LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing 24 upvotes, #8 of 2026-06-05
- AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding 10 upvotes, #16 of 2026-06-05
- OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning 24 upvotes, #16 of 2026-05-28
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining 144 upvotes, #1 of 2026-05-21
- RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting 9 upvotes, #23 of 2026-05-20
- StableVLA: Towards Robust Vision-Language-Action Models without Extra Data 15 upvotes, #18 of 2026-05-19
- GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding 4 upvotes, #36 of 2026-05-18
- VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction 26 upvotes, #15 of 2026-05-15
- PresentAgent-2: Towards Generalist Multimodal Presentation Agents 8 upvotes, #22 of 2026-05-14
- Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction 1 upvotes, #56 of 2026-05-13
- From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills 21 upvotes, #5 of 2026-05-04
- Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models 50 upvotes, #4 of 2026-04-30
- UniMesh: Unifying 3D Mesh Understanding and Generation 10 upvotes, #16 of 2026-04-22
- HSG: Hyperbolic Scene Graph 1 upvotes, #49 of 2026-04-21
- Context-Value-Action Architecture for Value-Driven Large Language Model Agents 8 upvotes, #23 of 2026-04-08
- OpenWorldLib: A Unified Codebase and Definition of Advanced World Models 200 upvotes, #2 of 2026-04-07
- DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models 345 upvotes, #1 of 2026-04-03
- HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention 40 upvotes, #6 of 2026-03-31
- Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models 19 upvotes, #9 of 2026-03-30
- PEARL: Personalized Streaming Video Understanding Model 40 upvotes, #6 of 2026-03-25
- FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow 32 upvotes, #7 of 2026-03-23
- Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models 21 upvotes, #12 of 2026-03-19
- MWM: Mobile World Models for Action-Conditioned Consistent Prediction 0 upvotes, #48 of 2026-03-10
- StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation 0 upvotes, #23 of 2026-02-20
- MMA: Multimodal Memory Agent 8 upvotes, #14 of 2026-02-19
- MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation 3 upvotes, #21 of 2026-02-17
- Code2Worlds: Empowering Coding LLMs for 4D World Generation 4 upvotes, #18 of 2026-02-16
- Light4D: Training-Free Extreme Viewpoint 4D Video Relighting 2 upvotes, #23 of 2026-02-16
- GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning 1 upvotes, #27 of 2026-02-16
- TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions 27 upvotes, #8 of 2026-02-12
- PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss 41 upvotes, #10 of 2026-02-03
- 3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence 1 upvotes, #39 of 2026-01-13
- AnyDepth: Depth Estimation Made Easy 9 upvotes, #14 of 2026-01-12
- DocDancer: Towards Agentic Document-Grounded Information Seeking 4 upvotes, #19 of 2026-01-09
- MDAgent2: Large Language Model for Code Generation and Knowledge Q&A in Molecular Dynamics 7 upvotes, #11 of 2026-01-08
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations 12 upvotes, #11 of 2025-12-25
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI 193 upvotes, #1 of 2025-12-23
- Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation 1 upvotes, #30 of 2025-12-18
- VABench: A Comprehensive Benchmark for Audio-Video Generation 7 upvotes, #19 of 2025-12-18
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling 40 upvotes, #4 of 2025-12-17
- DragMesh: Interactive 3D Generation Made Easy 1 upvotes, #25 of 2025-12-12
- From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs 23 upvotes, #6 of 2025-12-10
- EgoLCD: Egocentric Video Generation with Long Context Diffusion 5 upvotes, #29 of 2025-12-05
- Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation 13 upvotes, #18 of 2025-12-03
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots 5 upvotes, #12 of 2025-11-27
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward 31 upvotes, #7 of 2025-11-26
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation 62 upvotes, #3 of 2025-11-25
- EvoVLA: Self-Evolving Vision-Language-Action Model 4 upvotes, #24 of 2025-11-25
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks 10 upvotes, #19 of 2025-10-30
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence 11 upvotes, #14 of 2025-10-24
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback 18 upvotes, #13 of 2025-10-21
- Universal Image Restoration Pre-training via Masked Degradation Classification 10 upvotes, #19 of 2025-10-16
- WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation 5 upvotes, #26 of 2025-10-09
- StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes 1 upvotes, #33 of 2025-09-23
- Nav-R1: Reasoning and Navigation in Embodied Scenes 6 upvotes, #10 of 2025-09-16
- 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding 15 upvotes, #4 of 2025-08-04
- TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios 2 upvotes, #41 of 2025-05-26
- TransMLA: Multi-head Latent Attention Is All You Need 41 upvotes, #4 of 2025-02-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.