Daily Papers of 2025-09-15
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs 33 upvotes, #1 of 2025-09-15
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis 30 upvotes, #2 of 2025-09-15
- Virtual Agent Economies 26 upvotes, #3 of 2025-09-15
- X-Part: high fidelity and structure coherent shape decomposition 25 upvotes, #4 of 2025-09-15
- IntrEx: A Dataset for Modeling Engagement in Educational Conversations 24 upvotes, #5 of 2025-09-15
- HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering 24 upvotes, #5 of 2025-09-15
- Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models 21 upvotes, #7 of 2025-09-15
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools 15 upvotes, #8 of 2025-09-15
- Inpainting-Guided Policy Optimization for Diffusion Large Language Models 15 upvotes, #8 of 2025-09-15
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies 13 upvotes, #10 of 2025-09-15
- World Modeling with Probabilistic Structure Integration 13 upvotes, #10 of 2025-09-15
- LoFT: Parameter-Efficient Fine-Tuning for Long-tailed Semi-Supervised Learning in Open-World Scenarios 13 upvotes, #10 of 2025-09-15
- QuantAgent: Price-Driven Multi-Agent LLMs for High-Frequency Trading 13 upvotes, #10 of 2025-09-15
- Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation 11 upvotes, #14 of 2025-09-15
- VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions 10 upvotes, #15 of 2025-09-15
- CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models 4 upvotes, #16 of 2025-09-15
- Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts 4 upvotes, #16 of 2025-09-15
- Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images 4 upvotes, #16 of 2025-09-15
- Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation 2 upvotes, #19 of 2025-09-15
- DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning 2 upvotes, #19 of 2025-09-15
- CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China 1 upvotes, #21 of 2025-09-15
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.