Daily Papers of 2025-09-12

  1. VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model 189 upvotes, #1 of 2025-09-12
  2. HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning 117 upvotes, #2 of 2025-09-12
  3. SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning 73 upvotes, #3 of 2025-09-12
  4. MachineLearningLM: Continued Pretraining Language Models on Millions of Synthetic Tabular Prediction Tasks Scales In-Context ML 60 upvotes, #4 of 2025-09-12
  5. EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs 56 upvotes, #5 of 2025-09-12
  6. Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis 47 upvotes, #6 of 2025-09-12
  7. Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents 42 upvotes, #7 of 2025-09-12
  8. FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark 39 upvotes, #8 of 2025-09-12
  9. Can Understanding and Generation Truly Benefit Together -- or Just Coexist? 32 upvotes, #9 of 2025-09-12
  10. SpatialVID: A Large-Scale Video Dataset with Spatial Annotations 28 upvotes, #10 of 2025-09-12
  11. AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs 21 upvotes, #11 of 2025-09-12
  12. mmBERT: A Modern Multilingual Encoder with Annealed Language Learning 12 upvotes, #12 of 2025-09-12
  13. Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes 10 upvotes, #13 of 2025-09-12
  14. Visual Programmability: A Guide for Code-as-Thought in Chart Understanding 9 upvotes, #14 of 2025-09-12
  15. Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval 7 upvotes, #15 of 2025-09-12
  16. LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering 6 upvotes, #16 of 2025-09-12
  17. 2D Gaussian Splatting with Semantic Alignment for Image Inpainting 5 upvotes, #17 of 2025-09-12
  18. All You Need Is A Fuzzing Brain: An LLM-Powered System for Automated Vulnerability Detection and Patching 4 upvotes, #18 of 2025-09-12
  19. The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward 3 upvotes, #19 of 2025-09-12
  20. Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis 3 upvotes, #19 of 2025-09-12
  21. OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning 3 upvotes, #19 of 2025-09-12
  22. ObjectReact: Learning Object-Relative Control for Visual Navigation 3 upvotes, #19 of 2025-09-12
  23. Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation 2 upvotes, #23 of 2025-09-12
  24. Cross-Domain Evaluation of Transformer-Based Vulnerability Detection on Open & Industry Data 2 upvotes, #23 of 2025-09-12
  25. Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated 1 upvotes, #25 of 2025-09-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.