Daily Papers of 2025-03-24

  1. When Less is Enough: Adaptive Token Reduction for Efficient Image Representation 70 upvotes, #1 of 2025-03-24
  2. MAPS: A Multi-Agent Framework Based on Big Seven Personality and Socratic Guidance for Multimodal Scientific Problem Solving 52 upvotes, #2 of 2025-03-24
  3. A Comprehensive Survey on Long Context Language Modeling 47 upvotes, #3 of 2025-03-24
  4. MARS: A Multi-Agent Framework Incorporating Socratic Guidance for Automated Prompt Optimization 42 upvotes, #4 of 2025-03-24
  5. RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints 39 upvotes, #5 of 2025-03-24
  6. Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation 34 upvotes, #6 of 2025-03-24
  7. Modifying Large Language Model Post-Training for Diverse Creative Writing 33 upvotes, #7 of 2025-03-24
  8. TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian Splatting 22 upvotes, #8 of 2025-03-24
  9. OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement 20 upvotes, #9 of 2025-03-24
  10. Single Image Iterative Subject-driven Generation and Editing 13 upvotes, #10 of 2025-03-24
  11. MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems 13 upvotes, #10 of 2025-03-24
  12. Enabling Versatile Controls for Video Diffusion Models 13 upvotes, #10 of 2025-03-24
  13. ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering 11 upvotes, #13 of 2025-03-24
  14. From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration 9 upvotes, #14 of 2025-03-24
  15. Can Large Vision Language Models Read Maps Like a Human? 9 upvotes, #14 of 2025-03-24
  16. FastCuRL: Curriculum Reinforcement Learning with Progressive Context Extension for Efficient Training R1-like Reasoning Models 8 upvotes, #16 of 2025-03-24
  17. Implicit Bias-Like Patterns in Reasoning Models 7 upvotes, #17 of 2025-03-24
  18. PVChat: Personalized Video Chat with One-Shot Learning 7 upvotes, #17 of 2025-03-24
  19. GAEA: A Geolocation Aware Conversational Model 6 upvotes, #19 of 2025-03-24
  20. When Preferences Diverge: Aligning Diffusion Models with Minority-Aware Adaptive DPO 6 upvotes, #19 of 2025-03-24
  21. Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model 5 upvotes, #21 of 2025-03-24
  22. FFaceNeRF: Few-shot Face Editing in Neural Radiance Fields 5 upvotes, #21 of 2025-03-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.