Daily Papers of 2025-06-17
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention 233 upvotes, #1 of 2025-06-17
- Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning 65 upvotes, #2 of 2025-06-17
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents 56 upvotes, #3 of 2025-06-17
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency 46 upvotes, #4 of 2025-06-17
- Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression 46 upvotes, #4 of 2025-06-17
- DoTA-RAG: Dynamic of Thought Aggregation RAG 46 upvotes, #4 of 2025-06-17
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning 42 upvotes, #7 of 2025-06-17
- Discrete Diffusion in Large Language and Multimodal Models: A Survey 40 upvotes, #8 of 2025-06-17
- Essential-Web v1.0: 24T tokens of organized web data 37 upvotes, #9 of 2025-06-17
- TaskCraft: Automated Generation of Agentic Tasks 30 upvotes, #10 of 2025-06-17
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation 28 upvotes, #11 of 2025-06-17
- Test3R: Learning to Reconstruct 3D at Test Time 26 upvotes, #12 of 2025-06-17
- AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy 23 upvotes, #13 of 2025-06-17
- PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization 21 upvotes, #14 of 2025-06-17
- VGR: Visual Grounded Reasoning 19 upvotes, #15 of 2025-06-17
- Language Surgery in Multilingual Large Language Models 16 upvotes, #16 of 2025-06-17
- From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding 15 upvotes, #17 of 2025-06-17
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models 11 upvotes, #18 of 2025-06-17
- AI Agent Behavioral Science 10 upvotes, #19 of 2025-06-17
- Provably Learning from Language Feedback 8 upvotes, #20 of 2025-06-17
- A Technical Study into Small Reasoning Language Models 8 upvotes, #20 of 2025-06-17
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering 7 upvotes, #22 of 2025-06-17
- LETS Forecast: Learning Embedology for Time Series Forecasting 5 upvotes, #23 of 2025-06-17
- Supernova Event Dataset: Interpreting Large Language Model's Personality through Critical Event Analysis 5 upvotes, #23 of 2025-06-17
- SeqPE: Transformer with Sequential Position Encoding 5 upvotes, #23 of 2025-06-17
- SRLAgent: Enhancing Self-Regulated Learning Skills through Gamification and LLM Assistance 4 upvotes, #26 of 2025-06-17
- Steering LLM Thinking with Budget Guidance 4 upvotes, #26 of 2025-06-17
- Incorporating Domain Knowledge into Materials Tokenization 3 upvotes, #28 of 2025-06-17
- Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index 3 upvotes, #28 of 2025-06-17
- EgoPrivacy: What Your First-Person Camera Says About You? 3 upvotes, #28 of 2025-06-17
- QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety 3 upvotes, #28 of 2025-06-17
- Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts 3 upvotes, #28 of 2025-06-17
- MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos 3 upvotes, #28 of 2025-06-17
- Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition 3 upvotes, #28 of 2025-06-17
- DiffusionBlocks: Blockwise Training for Generative Models via Score-Based Diffusion 3 upvotes, #28 of 2025-06-17
- BOW: Bottlenecked Next Word Exploration 2 upvotes, #36 of 2025-06-17
- Hatevolution: What Static Benchmarks Don't Tell Us 1 upvotes, #37 of 2025-06-17
- Personalizable Long-Context Symbolic Music Infilling with MIDI-RWKV 1 upvotes, #37 of 2025-06-17
- Ai-Facilitated Analysis of Abstracts and Conclusions: Flagging Unsubstantiated Claims and Ambiguous Pronouns 1 upvotes, #37 of 2025-06-17
- Uncertainty-Aware Remaining Lifespan Prediction from Images 1 upvotes, #37 of 2025-06-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.