Daily Papers of 2025-09-15

  1. The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs 33 upvotes, #1 of 2025-09-15
  2. InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis 30 upvotes, #2 of 2025-09-15
  3. Virtual Agent Economies 26 upvotes, #3 of 2025-09-15
  4. X-Part: high fidelity and structure coherent shape decomposition 25 upvotes, #4 of 2025-09-15
  5. IntrEx: A Dataset for Modeling Engagement in Educational Conversations 24 upvotes, #5 of 2025-09-15
  6. HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering 24 upvotes, #5 of 2025-09-15
  7. Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models 21 upvotes, #7 of 2025-09-15
  8. MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools 15 upvotes, #8 of 2025-09-15
  9. Inpainting-Guided Policy Optimization for Diffusion Large Language Models 15 upvotes, #8 of 2025-09-15
  10. FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies 13 upvotes, #10 of 2025-09-15
  11. World Modeling with Probabilistic Structure Integration 13 upvotes, #10 of 2025-09-15
  12. LoFT: Parameter-Efficient Fine-Tuning for Long-tailed Semi-Supervised Learning in Open-World Scenarios 13 upvotes, #10 of 2025-09-15
  13. QuantAgent: Price-Driven Multi-Agent LLMs for High-Frequency Trading 13 upvotes, #10 of 2025-09-15
  14. Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation 11 upvotes, #14 of 2025-09-15
  15. VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions 10 upvotes, #15 of 2025-09-15
  16. CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models 4 upvotes, #16 of 2025-09-15
  17. Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts 4 upvotes, #16 of 2025-09-15
  18. Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images 4 upvotes, #16 of 2025-09-15
  19. Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation 2 upvotes, #19 of 2025-09-15
  20. DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning 2 upvotes, #19 of 2025-09-15
  21. CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China 1 upvotes, #21 of 2025-09-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.