kaipeng

kaipeng on Hugging Face Daily Papers: 53 papers, 16 in the top 3 of their day, 1,897 upvotes.

  1. Marionette: Predicting World States, Rendering Geometry, Painting Appearance 33 upvotes, #6 of 2026-08-17
  2. Alaya-EVOKE: From Linear-Scaling Supervision to Endless World 132 upvotes, #1 of 2026-08-14
  3. MASS: Multiplayer World Models with Authoritative Shared State 16 upvotes, #25 of 2026-08-07
  4. HelloWorld: Enabling Socially Interactive Characters in Video World Models 38 upvotes, #6 of 2026-08-06
  5. GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience 15 upvotes, #18 of 2026-08-05
  6. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow 20 upvotes, #19 of 2026-07-31
  7. AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report 57 upvotes, #6 of 2026-07-22
  8. Generative World Renderer at the Speed of Play 80 upvotes, #3 of 2026-07-22
  9. From Pixels to States: Rethinking Interactive World Models as Game Engines 35 upvotes, #7 of 2026-07-17
  10. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos 13 upvotes, #17 of 2026-07-16
  11. AlayaWorld: Long-Horizon and Playable Video World Generation 87 upvotes, #2 of 2026-07-08
  12. AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents 60 upvotes, #2 of 2026-07-03
  13. JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines 3 upvotes, #30 of 2026-06-19
  14. YoCausal: How Far is Video Generation from World Model? A Causality Perspective 51 upvotes, #7 of 2026-05-29
  15. Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? 169 upvotes, #3 of 2026-05-22
  16. WorldMark: A Unified Benchmark Suite for Interactive Video World Models 36 upvotes, #2 of 2026-04-24
  17. Generative World Renderer 101 upvotes, #3 of 2026-04-03
  18. PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference 50 upvotes, #3 of 2026-03-30
  19. WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG 90 upvotes, #2 of 2026-03-25
  20. PyVision-RL: Forging Open Agentic Vision Models via RL 29 upvotes, #3 of 2026-02-25
  21. LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces 12 upvotes, #10 of 2026-02-25
  22. World Craft: Agentic Framework to Create Visualizable Worlds via Text 20 upvotes, #9 of 2026-01-28
  23. MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences 11 upvotes, #12 of 2026-01-26
  24. Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models 4 upvotes, #20 of 2026-01-15
  25. Yume-1.5: A Text-Controlled Interactive World Generation Model 57 upvotes, #3 of 2025-12-30
  26. SVBench: Evaluation of Video Generation Models on Social Reasoning 7 upvotes, #13 of 2025-12-29
  27. TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning 15 upvotes, #16 of 2025-11-04
  28. OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling 100 upvotes, #1 of 2025-09-16
  29. Symbolic Graphics Programming with Large Language Models 44 upvotes, #3 of 2025-09-08
  30. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency 170 upvotes, #1 of 2025-08-26
  31. InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles 2 upvotes, #14 of 2025-08-25
  32. Neural-Driven Image Editing 25 upvotes, #10 of 2025-07-14
  33. PyVision: Agentic Vision with Dynamic Tooling 28 upvotes, #8 of 2025-07-11
  34. SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model 4 upvotes, #49 of 2025-05-30
  35. IA-T2I: Internet-Augmented Text-to-Image Generation 15 upvotes, #16 of 2025-05-22
  36. MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models 4 upvotes, #24 of 2025-04-15
  37. CLS-RL: Image Classification with Rule-Based Reinforcement Learning 9 upvotes, #28 of 2025-03-21
  38. Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction 6 upvotes, #36 of 2025-03-21
  39. MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification 9 upvotes, #16 of 2025-03-19
  40. PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models 5 upvotes, #22 of 2025-03-19
  41. ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges 8 upvotes, #14 of 2025-03-17
  42. Neighboring Autoregressive Modeling for Efficient Visual Generation 8 upvotes, #14 of 2025-03-17
  43. ARMOR v0.1: Empowering Autoregressive Multimodal Understanding Model with Interleaved Multimodal Generation via Asymmetric Synergy 8 upvotes, #14 of 2025-03-17
  44. Enhance-A-Video: Better Generated Video for Free 18 upvotes, #10 of 2025-02-12
  45. ZipAR: Accelerating Autoregressive Image Generation through Spatial Locality 7 upvotes, #24 of 2024-12-06
  46. ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification and KV Cache Compression 11 upvotes, #12 of 2024-10-17
  47. T3M: Text Guided 3D Human Motion Synthesis from Speech 8 upvotes, #7 of 2024-08-26
  48. EfficientQAT: Efficient Quantization-Aware Training for Large Language Models 5 upvotes, #14 of 2024-07-17
  49. Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
  50. GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices 21 upvotes, #8 of 2024-06-17
  51. OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models 20 upvotes, #2 of 2023-08-28
  52. Tiny LVLM-eHub: Early Multimodal Experiments with Bard 11 upvotes, #10 of 2023-08-08
  53. Meta-Transformer: A Unified Framework for Multimodal Learning 45 upvotes, #2 of 2023-07-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.