Bohan Zeng

Bohan Zeng on Hugging Face Daily Papers: 29 papers, 6 in the top 3 of their day, 1,520 upvotes.

  1. Flux-OPD: On-Policy Distillation with Evolving Contexts 44 upvotes, #10 of 2026-07-31
  2. DataPrep-Bench: Benchmarking LLMs as Training Data Preparators 55 upvotes, #1 of 2026-07-27
  3. KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation 42 upvotes, #6 of 2026-07-17
  4. SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 42 upvotes, #8 of 2026-06-09
  5. LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
  6. LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning 46 upvotes, #8 of 2026-05-22
  7. Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos 22 upvotes, #12 of 2026-05-20
  8. VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction 26 upvotes, #15 of 2026-05-15
  9. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models 200 upvotes, #2 of 2026-04-07
  10. DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models 345 upvotes, #1 of 2026-04-03
  11. OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models 46 upvotes, #5 of 2026-02-05
  12. Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers 27 upvotes, #13 of 2026-02-05
  13. Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks 46 upvotes, #7 of 2026-02-04
  14. CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation 30 upvotes, #12 of 2026-01-16
  15. DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI 193 upvotes, #1 of 2025-12-23
  16. VABench: A Comprehensive Benchmark for Audio-Video Generation 7 upvotes, #19 of 2025-12-18
  17. Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling 40 upvotes, #4 of 2025-12-17
  18. SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder 35 upvotes, #3 of 2025-12-15
  19. Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation 42 upvotes, #3 of 2025-12-12
  20. Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks 10 upvotes, #19 of 2025-10-30
  21. MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning 19 upvotes, #13 of 2025-10-20
  22. RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
  23. Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge 6 upvotes, #11 of 2025-09-17
  24. MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios 38 upvotes, #12 of 2025-05-28
  25. Mavors: Multi-granularity Video Representation for Multimodal Large Language Model 30 upvotes, #7 of 2025-04-15
  26. WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes 16 upvotes, #11 of 2025-03-18
  27. Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks 10 upvotes, #5 of 2025-01-30
  28. Semantic Score Distillation Sampling for Compositional Text-to-3D Generation 12 upvotes, #10 of 2024-10-14
  29. Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis 11 upvotes, #22 of 2024-10-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.