Daily Papers of 2025-09-12
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model 189 upvotes, #1 of 2025-09-12
- HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning 117 upvotes, #2 of 2025-09-12
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning 73 upvotes, #3 of 2025-09-12
- MachineLearningLM: Continued Pretraining Language Models on Millions of Synthetic Tabular Prediction Tasks Scales In-Context ML 60 upvotes, #4 of 2025-09-12
- EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs 56 upvotes, #5 of 2025-09-12
- Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis 47 upvotes, #6 of 2025-09-12
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents 42 upvotes, #7 of 2025-09-12
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark 39 upvotes, #8 of 2025-09-12
- Can Understanding and Generation Truly Benefit Together -- or Just Coexist? 32 upvotes, #9 of 2025-09-12
- SpatialVID: A Large-Scale Video Dataset with Spatial Annotations 28 upvotes, #10 of 2025-09-12
- AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs 21 upvotes, #11 of 2025-09-12
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning 12 upvotes, #12 of 2025-09-12
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes 10 upvotes, #13 of 2025-09-12
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding 9 upvotes, #14 of 2025-09-12
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval 7 upvotes, #15 of 2025-09-12
- LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering 6 upvotes, #16 of 2025-09-12
- 2D Gaussian Splatting with Semantic Alignment for Image Inpainting 5 upvotes, #17 of 2025-09-12
- All You Need Is A Fuzzing Brain: An LLM-Powered System for Automated Vulnerability Detection and Patching 4 upvotes, #18 of 2025-09-12
- The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward 3 upvotes, #19 of 2025-09-12
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis 3 upvotes, #19 of 2025-09-12
- OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning 3 upvotes, #19 of 2025-09-12
- ObjectReact: Learning Object-Relative Control for Visual Navigation 3 upvotes, #19 of 2025-09-12
- Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation 2 upvotes, #23 of 2025-09-12
- Cross-Domain Evaluation of Transformer-Based Vulnerability Detection on Open & Industry Data 2 upvotes, #23 of 2025-09-12
- Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated 1 upvotes, #25 of 2025-09-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.