Bohan Zeng
Bohan Zeng on Hugging Face Daily Papers: 29 papers, 6 in the top 3 of their day, 1,520 upvotes.
- Flux-OPD: On-Policy Distillation with Evolving Contexts 44 upvotes, #10 of 2026-07-31
- DataPrep-Bench: Benchmarking LLMs as Training Data Preparators 55 upvotes, #1 of 2026-07-27
- KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation 42 upvotes, #6 of 2026-07-17
- SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 42 upvotes, #8 of 2026-06-09
- LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
- LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning 46 upvotes, #8 of 2026-05-22
- Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos 22 upvotes, #12 of 2026-05-20
- VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction 26 upvotes, #15 of 2026-05-15
- OpenWorldLib: A Unified Codebase and Definition of Advanced World Models 200 upvotes, #2 of 2026-04-07
- DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models 345 upvotes, #1 of 2026-04-03
- OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models 46 upvotes, #5 of 2026-02-05
- Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers 27 upvotes, #13 of 2026-02-05
- Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks 46 upvotes, #7 of 2026-02-04
- CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation 30 upvotes, #12 of 2026-01-16
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI 193 upvotes, #1 of 2025-12-23
- VABench: A Comprehensive Benchmark for Audio-Video Generation 7 upvotes, #19 of 2025-12-18
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling 40 upvotes, #4 of 2025-12-17
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder 35 upvotes, #3 of 2025-12-15
- Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation 42 upvotes, #3 of 2025-12-12
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks 10 upvotes, #19 of 2025-10-30
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning 19 upvotes, #13 of 2025-10-20
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
- Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge 6 upvotes, #11 of 2025-09-17
- MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios 38 upvotes, #12 of 2025-05-28
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model 30 upvotes, #7 of 2025-04-15
- WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes 16 upvotes, #11 of 2025-03-18
- Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks 10 upvotes, #5 of 2025-01-30
- Semantic Score Distillation Sampling for Compositional Text-to-3D Generation 12 upvotes, #10 of 2024-10-14
- Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis 11 upvotes, #22 of 2024-10-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.