Zhejiang University
Zhejiang University on Hugging Face Daily Papers: 40 papers, 6 in the top 3 of their day, 3 paper of the day.
- CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding 10 upvotes, #46 of 2026-10-01
- Scaling Properties of Same-Family On-Policy Distillation 321 upvotes, #4 of 2026-09-30
- PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation 166 upvotes, #6 of 2026-09-30
- SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue 100 upvotes, #1 of 2026-09-24
- RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? 28 upvotes, #22 of 2026-09-09
- ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes 25 upvotes, #13 of 2026-09-03
- LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation 29 upvotes, #12 of 2026-08-31
- Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion 33 upvotes, #7 of 2026-08-25
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence 43 upvotes, #10 of 2026-08-14
- Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination 4 upvotes, #27 of 2026-08-10
- INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models 16 upvotes, #23 of 2026-07-31
- Pass the Baton: Trajectory-Relayed On-Policy Distillation 33 upvotes, #6 of 2026-07-29
- Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning 18 upvotes, #12 of 2026-07-02
- From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion 32 upvotes, #11 of 2026-06-12
- VIA-SD: Verification via Intra-Model Routing for Speculative Decoding 35 upvotes, #10 of 2026-06-12
- Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning 24 upvotes, #12 of 2026-06-09
- Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing 1 upvotes, #43 of 2026-06-05
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching 16 upvotes, #18 of 2026-06-04
- Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? 30 upvotes, #12 of 2026-06-02
- Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios 31 upvotes, #14 of 2026-06-01
- Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer 37 upvotes, #11 of 2026-06-01
- TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction 51 upvotes, #5 of 2026-05-26
- PanoWorld: Towards Spatial Supersensing in 360^circ Panorama World 21 upvotes, #18 of 2026-05-15
- Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models 10 upvotes, #27 of 2026-05-15
- Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective 14 upvotes, #16 of 2026-04-15
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents 141 upvotes, #1 of 2026-04-14
- OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering 25 upvotes, #16 of 2026-04-10
- Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO 42 upvotes, #12 of 2026-02-10
- MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments 13 upvotes, #15 of 2026-02-09
- InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning 12 upvotes, #16 of 2026-02-09
- Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes 2 upvotes, #22 of 2026-01-09
- InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields 95 upvotes, #2 of 2026-01-07
- Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems 7 upvotes, #14 of 2026-01-01
- Evaluating Parameter Efficient Methods for RLVR 24 upvotes, #2 of 2025-12-31
- Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality 46 upvotes, #3 of 2025-12-10
- ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation 20 upvotes, #15 of 2025-12-03
- EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering 26 upvotes, #16 of 2025-09-30
- GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts 27 upvotes, #13 of 2025-09-30
- Taming LLMs by Scaling Learning Rates with Gradient Grouping 36 upvotes, #4 of 2025-06-03
- MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization 78 upvotes, #1 of 2025-04-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.