Daily Papers of 2025-06-26
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation 63 upvotes, #1 of 2025-06-26
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 55 upvotes, #2 of 2025-06-26
- Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models 43 upvotes, #3 of 2025-06-26
- OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling 42 upvotes, #4 of 2025-06-26
- Inverse-and-Edit: Effective and Fast Image Editing by Cycle Consistency Models 41 upvotes, #5 of 2025-06-26
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning 21 upvotes, #6 of 2025-06-26
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation 16 upvotes, #7 of 2025-06-26
- HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling 15 upvotes, #8 of 2025-06-26
- Use Property-Based Testing to Bridge LLM Code Generation and Validation 10 upvotes, #9 of 2025-06-26
- Thought Anchors: Which LLM Reasoning Steps Matter? 9 upvotes, #10 of 2025-06-26
- When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs 8 upvotes, #11 of 2025-06-26
- Is There a Case for Conversation Optimized Tokenizers in Large Language Models? 7 upvotes, #12 of 2025-06-26
- GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching 7 upvotes, #12 of 2025-06-26
- ReCode: Updating Code API Knowledge with Reinforcement Learning 7 upvotes, #12 of 2025-06-26
- FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation 5 upvotes, #15 of 2025-06-26
- MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications 5 upvotes, #15 of 2025-06-26
- Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content 3 upvotes, #17 of 2025-06-26
- The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs 2 upvotes, #18 of 2025-06-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.