Daily Papers of 2025-03-14

  1. Transformers without Normalization 128 upvotes, #1 of 2025-03-14
  2. CoSTAast: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing 70 upvotes, #2 of 2025-03-14
  3. Charting and Navigating Hugging Face's Model Atlas 67 upvotes, #3 of 2025-03-14
  4. World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning 45 upvotes, #4 of 2025-03-14
  5. GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing 45 upvotes, #4 of 2025-03-14
  6. Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models 34 upvotes, #6 of 2025-03-14
  7. VisualPRM: An Effective Process Reward Model for Multimodal Reasoning 31 upvotes, #7 of 2025-03-14
  8. CoRe^2: Collect, Reflect and Refine to Generate Better and Faster 29 upvotes, #8 of 2025-03-14
  9. 4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models 28 upvotes, #9 of 2025-03-14
  10. OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting 26 upvotes, #10 of 2025-03-14
  11. Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond 25 upvotes, #11 of 2025-03-14
  12. SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation 24 upvotes, #12 of 2025-03-14
  13. VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 20 upvotes, #13 of 2025-03-14
  14. Shifting Long-Context LLMs Research from Input to Output 19 upvotes, #14 of 2025-03-14
  15. New Trends for Modern Machine Translation with Large Reasoning Models 19 upvotes, #14 of 2025-03-14
  16. GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding 18 upvotes, #16 of 2025-03-14
  17. DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation 17 upvotes, #17 of 2025-03-14
  18. Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k 16 upvotes, #18 of 2025-03-14
  19. R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization 16 upvotes, #18 of 2025-03-14
  20. Distilling Diversity and Control in Diffusion Models 14 upvotes, #20 of 2025-03-14
  21. Long Context Tuning for Video Generation 13 upvotes, #21 of 2025-03-14
  22. Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo 12 upvotes, #22 of 2025-03-14
  23. Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark 11 upvotes, #23 of 2025-03-14
  24. On the Limitations of Vision-Language Models in Understanding Image Transforms 10 upvotes, #24 of 2025-03-14
  25. CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance 10 upvotes, #24 of 2025-03-14
  26. ConsisLoRA: Enhancing Content and Style Consistency for LoRA-based Style Transfer 8 upvotes, #26 of 2025-03-14
  27. Discovering Influential Neuron Path in Vision Transformers 6 upvotes, #27 of 2025-03-14
  28. Quantization for OpenAI's Whisper Models: A Comparative Analysis 6 upvotes, #27 of 2025-03-14
  29. Piece it Together: Part-Based Concepting with IP-Priors 6 upvotes, #27 of 2025-03-14
  30. Autoregressive Image Generation with Randomized Parallel Decoding 6 upvotes, #27 of 2025-03-14
  31. UniGoal: Towards Universal Zero-shot Goal-oriented Navigation 6 upvotes, #27 of 2025-03-14
  32. "Silent Is Not Actually Silent": An Investigation of Toxicity on Bug Report Discussion 4 upvotes, #32 of 2025-03-14
  33. MinorBench: A hand-built benchmark for content-based risks for children 4 upvotes, #32 of 2025-03-14
  34. TruthPrInt: Mitigating LVLM Object Hallucination Via Latent Truthful-Guided Pre-Intervention 4 upvotes, #32 of 2025-03-14
  35. The Curse of Conditions: Analyzing and Improving Optimal Transport for Conditional Flow-Based Generation 3 upvotes, #35 of 2025-03-14
  36. PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM 2 upvotes, #36 of 2025-03-14
  37. PerCoV2: Improved Ultra-Low Bit-Rate Perceptual Image Compression with Implicit Hierarchical Masked Image Modeling 2 upvotes, #36 of 2025-03-14
  38. A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 2 upvotes, #36 of 2025-03-14
  39. Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective 2 upvotes, #36 of 2025-03-14

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.