National University of Singapore
National University of Singapore on Hugging Face Daily Papers: 70 papers, 10 in the top 3 of their day, 2 paper of the day.
- FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation 19 upvotes, #3 of 2026-10-05
- Agent Priors-guided Policy Learning 77 upvotes, #9 of 2026-10-02
- MaLiang-Harness: A Programmable Path to Image and Video Generation 410 upvotes, #2 of 2026-09-30
- StoryEngine: A State-Grounded Agentic Framework for Video Storytelling 10 upvotes, #54 of 2026-09-30
- Omni-IO Skills: Harnessing Your Agent Omni-Native 292 upvotes, #5 of 2026-09-30
- Structured Residual Connectivity Matters for Diffusion Transformers 17 upvotes, #44 of 2026-09-29
- Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles 7 upvotes, #27 of 2026-09-22
- Grounded Action Model: 3D Grounding as a Foundation for Robotics 89 upvotes, #6 of 2026-09-22
- What Does Privileged Information Add to On-Policy Self-Distillation? 36 upvotes, #23 of 2026-09-18
- AI for Games in the Foundation Model Era 145 upvotes, #4 of 2026-09-16
- Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models 70 upvotes, #6 of 2026-09-14
- When Models Edit Too Much: On the Fidelity of Minimal Code Edits 10 upvotes, #16 of 2026-09-07
- LMSM: LLM Security Framework Inspired by Linux Security Modules 5 upvotes, #26 of 2026-08-31
- Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models 185 upvotes, #1 of 2026-08-28
- Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements 105 upvotes, #3 of 2026-08-19
- V-RAE: Rethinking Video Latent Spaces for Generation 30 upvotes, #8 of 2026-08-19
- Latent On-Policy Self-Distillation 24 upvotes, #10 of 2026-08-17
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 87 upvotes, #5 of 2026-08-14
- ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning 4 upvotes, #32 of 2026-08-05
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space 10 upvotes, #30 of 2026-08-04
- RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models 9 upvotes, #17 of 2026-08-03
- Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning 27 upvotes, #10 of 2026-07-08
- Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator 3 upvotes, #35 of 2026-07-08
- NoPA: Non-Parametric Online 3D Scene Graph Generation 9 upvotes, #21 of 2026-07-02
- World Action Models: A Survey 56 upvotes, #6 of 2026-06-23
- OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation 27 upvotes, #9 of 2026-06-17
- Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents 73 upvotes, #3 of 2026-06-15
- An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models 6 upvotes, #30 of 2026-06-15
- One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA 16 upvotes, #17 of 2026-06-10
- Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data 1 upvotes, #44 of 2026-06-09
- Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents 1 upvotes, #44 of 2026-06-09
- Audio Interaction Model 108 upvotes, #2 of 2026-06-04
- Q-ARVD: Quantizing Autoregressive Video Diffusion Models 21 upvotes, #18 of 2026-05-22
- Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation 131 upvotes, #2 of 2026-05-21
- Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs 28 upvotes, #13 of 2026-05-21
- Audio-Visual Intelligence in Large Foundation Models 32 upvotes, #10 of 2026-05-08
- Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment 9 upvotes, #12 of 2026-04-28
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents 108 upvotes, #6 of 2026-04-10
- DMax: Aggressive Parallel Decoding for dLLMs 50 upvotes, #9 of 2026-04-10
- Paper Espresso: From Paper Overload to Research Insight 12 upvotes, #25 of 2026-04-07
- ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration 13 upvotes, #23 of 2026-04-07
- Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies 15 upvotes, #19 of 2026-04-07
- Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers 18 upvotes, #15 of 2026-04-03
- Make Geometry Matter for Spatial Reasoning 32 upvotes, #8 of 2026-03-31
- Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models 28 upvotes, #9 of 2026-03-17
- ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer 24 upvotes, #11 of 2026-03-17
- Can Vision-Language Models Solve the Shell Game? 39 upvotes, #2 of 2026-03-16
- In-Context Reinforcement Learning for Tool Use in Large Language Models 39 upvotes, #5 of 2026-03-12
- EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding 1 upvotes, #18 of 2026-03-05
- LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency 1 upvotes, #29 of 2026-02-25
- dVoting: Fast Voting for dLLMs 20 upvotes, #14 of 2026-02-13
- Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model 20 upvotes, #14 of 2026-02-11
- Factorized Learning for Temporally Grounded Video-Language Models 6 upvotes, #16 of 2026-01-01
- SpotEdit: Selective Region Editing in Diffusion Transformers 37 upvotes, #8 of 2025-12-30
- SlideTailor: Personalized Presentation Slide Generation for Scientific Papers 8 upvotes, #12 of 2025-12-29
- WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion 29 upvotes, #5 of 2025-12-23
- PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing 58 upvotes, #4 of 2025-12-05
- Vision Bridge Transformer at Scale 43 upvotes, #4 of 2025-12-01
- In-Video Instructions: Visual Signals as Generative Control 28 upvotes, #7 of 2025-11-25
- SAM2S: Segment Anything in Surgical Videos via Semantic Long-term Tracking 7 upvotes, #17 of 2025-11-21
- SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization 15 upvotes, #12 of 2025-11-11
- Diffusion Language Models are Super Data Learners 110 upvotes, #1 of 2025-11-06
- EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation 2 upvotes, #23 of 2025-10-31
- FML-bench: A Benchmark for Automatic ML Research Agents Highlighting the Importance of Exploration Breadth 7 upvotes, #29 of 2025-10-17
- RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems 1 upvotes, #47 of 2025-10-17
- MixReasoning: Switching Modes to Think 21 upvotes, #8 of 2025-10-08
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use 156 upvotes, #2 of 2025-10-01
- dParallel: Learnable Parallel Decoding for dLLMs 18 upvotes, #15 of 2025-10-01
- Efficient Reasoning Models: A Survey 18 upvotes, #8 of 2025-04-16
- DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting 3 upvotes, #27 of 2025-04-02
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.