ByteDance
ByteDance on Hugging Face Daily Papers: 73 papers, 17 in the top 3 of their day, 3 paper of the day.
- Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE 6 upvotes, #65 of 2026-09-30
- TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces 72 upvotes, #9 of 2026-09-29
- Disentangling Representation Evolution in Transformers through Directional Decomposition 9 upvotes, #17 of 2026-09-16
- AgenticGen: Reward-Guided Agentic Video Generation for Advertising 9 upvotes, #23 of 2026-09-10
- DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents 91 upvotes, #3 of 2026-08-31
- Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction 3 upvotes, #27 of 2026-08-17
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring 129 upvotes, #4 of 2026-08-11
- Douyin Multimodal Embedding Model Technical Report 14 upvotes, #10 of 2026-08-10
- When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation 23 upvotes, #12 of 2026-08-06
- SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks 155 upvotes, #2 of 2026-08-04
- CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization 83 upvotes, #3 of 2026-07-30
- TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation 7 upvotes, #21 of 2026-07-24
- SWE-Pruner Pro: The Coder LLM Already Knows What to Prune 77 upvotes, #5 of 2026-07-21
- UniVR: Thinking in Visual Space for Unified Visual Reasoning 32 upvotes, #10 of 2026-07-17
- ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation 19 upvotes, #13 of 2026-07-16
- Dockerless: Environment-Free Program Verifier for Coding Agents 108 upvotes, #1 of 2026-07-01
- SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing 6 upvotes, #34 of 2026-06-30
- FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation 1 upvotes, #24 of 2026-06-24
- PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models 64 upvotes, #2 of 2026-06-22
- ActWorld: From Explorable to Interactive World Model via Action-Aware Memory 8 upvotes, #21 of 2026-06-17
- Towards One-to-Many Temporal Grounding 7 upvotes, #22 of 2026-06-05
- SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue 56 upvotes, #5 of 2026-06-01
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation 25 upvotes, #12 of 2026-05-27
- Bernini: Latent Semantic Planning for Video Diffusion 12 upvotes, #24 of 2026-05-22
- Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning 7 upvotes, #35 of 2026-05-15
- Let ViT Speak: Generative Language-Image Pre-training 32 upvotes, #3 of 2026-05-04
- OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation 69 upvotes, #5 of 2026-04-14
- DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing 18 upvotes, #20 of 2026-03-31
- HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images 28 upvotes, #6 of 2026-03-06
- Timer-S1: A Billion-Scale Time Series Foundation Model with Serial Scaling 16 upvotes, #11 of 2026-03-06
- Heterogeneous Agent Collaborative Reinforcement Learning 170 upvotes, #1 of 2026-03-05
- Helios: Real Real-Time Long Video Generation Model 159 upvotes, #2 of 2026-03-05
- DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation 38 upvotes, #4 of 2026-02-26
- Does Your Reasoning Model Implicitly Know When to Stop Thinking? 253 upvotes, #1 of 2026-02-23
- UniWeTok: An Unified Binary Tokenizer with Codebook Size 2^{128} for Unified Multimodal Large Language Model 12 upvotes, #12 of 2026-02-17
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens 48 upvotes, #3 of 2026-02-17
- MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs 59 upvotes, #3 of 2026-02-16
- Thinking with Drafting: Optical Decompression via Logical Reconstruction 32 upvotes, #9 of 2026-02-13
- NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control 43 upvotes, #7 of 2026-02-13
- FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space 18 upvotes, #24 of 2026-02-03
- DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context Learning 12 upvotes, #16 of 2026-02-02
- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents 87 upvotes, #2 of 2026-01-26
- SAMTok: Representing Any Mask with Two Words 41 upvotes, #9 of 2026-01-23
- OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer 44 upvotes, #6 of 2026-01-21
- FlowAct-R1: Towards Interactive Humanoid Video Generation 33 upvotes, #10 of 2026-01-16
- The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning 48 upvotes, #4 of 2026-01-12
- ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing 6 upvotes, #12 of 2026-01-08
- DreamStyle: A Unified Framework for Video Stylization 22 upvotes, #10 of 2026-01-07
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation 56 upvotes, #3 of 2026-01-06
- DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer 49 upvotes, #5 of 2026-01-06
- VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation 32 upvotes, #6 of 2026-01-06
- DreamOmni3: Scribble-based Editing and Generation 14 upvotes, #4 of 2025-12-31
- Bridging Your Imagination with Audio-Video Generation via a Unified Director 5 upvotes, #24 of 2025-12-30
- DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation 32 upvotes, #3 of 2025-12-25
- StoryMem: Multi-shot Long Video Storytelling with Memory 17 upvotes, #11 of 2025-12-23
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation 64 upvotes, #6 of 2025-11-18
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning 2 upvotes, #13 of 2025-11-12
- PairUni: Pairwise Training for Unified Multimodal Language Models 13 upvotes, #16 of 2025-10-30
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents 50 upvotes, #5 of 2025-10-29
- Video-As-Prompt: Unified Semantic Control for Video Generation 44 upvotes, #2 of 2025-10-27
- Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence 52 upvotes, #3 of 2025-10-24
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation 38 upvotes, #8 of 2025-10-22
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 35 upvotes, #9 of 2025-10-22
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model 10 upvotes, #21 of 2025-10-15
- Lynx: Towards High-Fidelity Personalized Video Generation 12 upvotes, #7 of 2025-09-22
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks 13 upvotes, #14 of 2025-08-27
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology 44 upvotes, #3 of 2025-07-11
- CyberV: Cybernetics for Test-time Scaling in Video Understanding 4 upvotes, #32 of 2025-06-10
- MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query 3 upvotes, #35 of 2025-06-04
- MAGREF: Masked Guidance for Any-Reference Video Generation 9 upvotes, #33 of 2025-05-30
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer 15 upvotes, #10 of 2025-04-16
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding 28 upvotes, #6 of 2025-04-16
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos 40 upvotes, #4 of 2025-01-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.