ByteDance

ByteDance on Hugging Face Daily Papers: 73 papers, 17 in the top 3 of their day, 3 paper of the day.

  1. Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE 6 upvotes, #65 of 2026-09-30
  2. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces 72 upvotes, #9 of 2026-09-29
  3. Disentangling Representation Evolution in Transformers through Directional Decomposition 9 upvotes, #17 of 2026-09-16
  4. AgenticGen: Reward-Guided Agentic Video Generation for Advertising 9 upvotes, #23 of 2026-09-10
  5. DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents 91 upvotes, #3 of 2026-08-31
  6. Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction 3 upvotes, #27 of 2026-08-17
  7. SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 129 upvotes, #4 of 2026-08-11
  8. Douyin Multimodal Embedding Model Technical Report 14 upvotes, #10 of 2026-08-10
  9. When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation 23 upvotes, #12 of 2026-08-06
  10. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks 155 upvotes, #2 of 2026-08-04
  11. CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization 83 upvotes, #3 of 2026-07-30
  12. TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation 7 upvotes, #21 of 2026-07-24
  13. SWE-Pruner Pro: The Coder LLM Already Knows What to Prune 77 upvotes, #5 of 2026-07-21
  14. UniVR: Thinking in Visual Space for Unified Visual Reasoning 32 upvotes, #10 of 2026-07-17
  15. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation 19 upvotes, #13 of 2026-07-16
  16. Dockerless: Environment-Free Program Verifier for Coding Agents 108 upvotes, #1 of 2026-07-01
  17. SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing 6 upvotes, #34 of 2026-06-30
  18. FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation 1 upvotes, #24 of 2026-06-24
  19. PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models 64 upvotes, #2 of 2026-06-22
  20. ActWorld: From Explorable to Interactive World Model via Action-Aware Memory 8 upvotes, #21 of 2026-06-17
  21. Towards One-to-Many Temporal Grounding 7 upvotes, #22 of 2026-06-05
  22. SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue 56 upvotes, #5 of 2026-06-01
  23. MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation 25 upvotes, #12 of 2026-05-27
  24. Bernini: Latent Semantic Planning for Video Diffusion 12 upvotes, #24 of 2026-05-22
  25. Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning 7 upvotes, #35 of 2026-05-15
  26. Let ViT Speak: Generative Language-Image Pre-training 32 upvotes, #3 of 2026-05-04
  27. OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation 69 upvotes, #5 of 2026-04-14
  28. DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing 18 upvotes, #20 of 2026-03-31
  29. HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images 28 upvotes, #6 of 2026-03-06
  30. Timer-S1: A Billion-Scale Time Series Foundation Model with Serial Scaling 16 upvotes, #11 of 2026-03-06
  31. Heterogeneous Agent Collaborative Reinforcement Learning 170 upvotes, #1 of 2026-03-05
  32. Helios: Real Real-Time Long Video Generation Model 159 upvotes, #2 of 2026-03-05
  33. DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation 38 upvotes, #4 of 2026-02-26
  34. Does Your Reasoning Model Implicitly Know When to Stop Thinking? 253 upvotes, #1 of 2026-02-23
  35. UniWeTok: An Unified Binary Tokenizer with Codebook Size 2^{128} for Unified Multimodal Large Language Model 12 upvotes, #12 of 2026-02-17
  36. BitDance: Scaling Autoregressive Generative Models with Binary Tokens 48 upvotes, #3 of 2026-02-17
  37. MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs 59 upvotes, #3 of 2026-02-16
  38. Thinking with Drafting: Optical Decompression via Logical Reconstruction 32 upvotes, #9 of 2026-02-13
  39. NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control 43 upvotes, #7 of 2026-02-13
  40. FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space 18 upvotes, #24 of 2026-02-03
  41. DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context Learning 12 upvotes, #16 of 2026-02-02
  42. SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents 87 upvotes, #2 of 2026-01-26
  43. SAMTok: Representing Any Mask with Two Words 41 upvotes, #9 of 2026-01-23
  44. OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer 44 upvotes, #6 of 2026-01-21
  45. FlowAct-R1: Towards Interactive Humanoid Video Generation 33 upvotes, #10 of 2026-01-16
  46. The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning 48 upvotes, #4 of 2026-01-12
  47. ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing 6 upvotes, #12 of 2026-01-08
  48. DreamStyle: A Unified Framework for Video Stylization 22 upvotes, #10 of 2026-01-07
  49. NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation 56 upvotes, #3 of 2026-01-06
  50. DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer 49 upvotes, #5 of 2026-01-06
  51. VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation 32 upvotes, #6 of 2026-01-06
  52. DreamOmni3: Scribble-based Editing and Generation 14 upvotes, #4 of 2025-12-31
  53. Bridging Your Imagination with Audio-Video Generation via a Unified Director 5 upvotes, #24 of 2025-12-30
  54. DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation 32 upvotes, #3 of 2025-12-25
  55. StoryMem: Multi-shot Long Video Storytelling with Memory 17 upvotes, #11 of 2025-12-23
  56. MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation 64 upvotes, #6 of 2025-11-18
  57. TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning 2 upvotes, #13 of 2025-11-12
  58. PairUni: Pairwise Training for Unified Multimodal Language Models 13 upvotes, #16 of 2025-10-30
  59. Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents 50 upvotes, #5 of 2025-10-29
  60. Video-As-Prompt: Unified Semantic Control for Video Generation 44 upvotes, #2 of 2025-10-27
  61. Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence 52 upvotes, #3 of 2025-10-24
  62. MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation 38 upvotes, #8 of 2025-10-22
  63. Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 35 upvotes, #9 of 2025-10-22
  64. SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model 10 upvotes, #21 of 2025-10-15
  65. Lynx: Towards High-Fidelity Personalized Video Generation 12 upvotes, #7 of 2025-09-22
  66. ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks 13 upvotes, #14 of 2025-08-27
  67. Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology 44 upvotes, #3 of 2025-07-11
  68. CyberV: Cybernetics for Test-time Scaling in Video Understanding 4 upvotes, #32 of 2025-06-10
  69. MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query 3 upvotes, #35 of 2025-06-04
  70. MAGREF: Masked Guidance for Any-Reference Video Generation 9 upvotes, #33 of 2025-05-30
  71. The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer 15 upvotes, #10 of 2025-04-16
  72. Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding 28 upvotes, #6 of 2025-04-16
  73. Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos 40 upvotes, #4 of 2025-01-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.