kaipeng
kaipeng on Hugging Face Daily Papers: 53 papers, 16 in the top 3 of their day, 1,897 upvotes.
- Marionette: Predicting World States, Rendering Geometry, Painting Appearance 33 upvotes, #6 of 2026-08-17
- Alaya-EVOKE: From Linear-Scaling Supervision to Endless World 132 upvotes, #1 of 2026-08-14
- MASS: Multiplayer World Models with Authoritative Shared State 16 upvotes, #25 of 2026-08-07
- HelloWorld: Enabling Socially Interactive Characters in Video World Models 38 upvotes, #6 of 2026-08-06
- GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience 15 upvotes, #18 of 2026-08-05
- ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow 20 upvotes, #19 of 2026-07-31
- AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report 57 upvotes, #6 of 2026-07-22
- Generative World Renderer at the Speed of Play 80 upvotes, #3 of 2026-07-22
- From Pixels to States: Rethinking Interactive World Models as Game Engines 35 upvotes, #7 of 2026-07-17
- Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos 13 upvotes, #17 of 2026-07-16
- AlayaWorld: Long-Horizon and Playable Video World Generation 87 upvotes, #2 of 2026-07-08
- AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents 60 upvotes, #2 of 2026-07-03
- JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines 3 upvotes, #30 of 2026-06-19
- YoCausal: How Far is Video Generation from World Model? A Causality Perspective 51 upvotes, #7 of 2026-05-29
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? 169 upvotes, #3 of 2026-05-22
- WorldMark: A Unified Benchmark Suite for Interactive Video World Models 36 upvotes, #2 of 2026-04-24
- Generative World Renderer 101 upvotes, #3 of 2026-04-03
- PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference 50 upvotes, #3 of 2026-03-30
- WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG 90 upvotes, #2 of 2026-03-25
- PyVision-RL: Forging Open Agentic Vision Models via RL 29 upvotes, #3 of 2026-02-25
- LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces 12 upvotes, #10 of 2026-02-25
- World Craft: Agentic Framework to Create Visualizable Worlds via Text 20 upvotes, #9 of 2026-01-28
- MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences 11 upvotes, #12 of 2026-01-26
- Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models 4 upvotes, #20 of 2026-01-15
- Yume-1.5: A Text-Controlled Interactive World Generation Model 57 upvotes, #3 of 2025-12-30
- SVBench: Evaluation of Video Generation Models on Social Reasoning 7 upvotes, #13 of 2025-12-29
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning 15 upvotes, #16 of 2025-11-04
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling 100 upvotes, #1 of 2025-09-16
- Symbolic Graphics Programming with Large Language Models 44 upvotes, #3 of 2025-09-08
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 170 upvotes, #1 of 2025-08-26
- InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles 2 upvotes, #14 of 2025-08-25
- Neural-Driven Image Editing 25 upvotes, #10 of 2025-07-14
- PyVision: Agentic Vision with Dynamic Tooling 28 upvotes, #8 of 2025-07-11
- SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model 4 upvotes, #49 of 2025-05-30
- IA-T2I: Internet-Augmented Text-to-Image Generation 15 upvotes, #16 of 2025-05-22
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models 4 upvotes, #24 of 2025-04-15
- CLS-RL: Image Classification with Rule-Based Reinforcement Learning 9 upvotes, #28 of 2025-03-21
- Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction 6 upvotes, #36 of 2025-03-21
- MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification 9 upvotes, #16 of 2025-03-19
- PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models 5 upvotes, #22 of 2025-03-19
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges 8 upvotes, #14 of 2025-03-17
- Neighboring Autoregressive Modeling for Efficient Visual Generation 8 upvotes, #14 of 2025-03-17
- ARMOR v0.1: Empowering Autoregressive Multimodal Understanding Model with Interleaved Multimodal Generation via Asymmetric Synergy 8 upvotes, #14 of 2025-03-17
- Enhance-A-Video: Better Generated Video for Free 18 upvotes, #10 of 2025-02-12
- ZipAR: Accelerating Autoregressive Image Generation through Spatial Locality 7 upvotes, #24 of 2024-12-06
- ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification and KV Cache Compression 11 upvotes, #12 of 2024-10-17
- T3M: Text Guided 3D Human Motion Synthesis from Speech 8 upvotes, #7 of 2024-08-26
- EfficientQAT: Efficient Quantization-Aware Training for Large Language Models 5 upvotes, #14 of 2024-07-17
- Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
- GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices 21 upvotes, #8 of 2024-06-17
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models 20 upvotes, #2 of 2023-08-28
- Tiny LVLM-eHub: Early Multimodal Experiments with Bard 11 upvotes, #10 of 2023-08-08
- Meta-Transformer: A Unified Framework for Multimodal Learning 45 upvotes, #2 of 2023-07-21
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.