Daily Papers of 2025-06-26

  1. ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation 63 upvotes, #1 of 2025-06-26
  2. FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 55 upvotes, #2 of 2025-06-26
  3. Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models 43 upvotes, #3 of 2025-06-26
  4. OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling 42 upvotes, #4 of 2025-06-26
  5. Inverse-and-Edit: Effective and Fast Image Editing by Cycle Consistency Models 41 upvotes, #5 of 2025-06-26
  6. DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning 21 upvotes, #6 of 2025-06-26
  7. RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation 16 upvotes, #7 of 2025-06-26
  8. HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling 15 upvotes, #8 of 2025-06-26
  9. Use Property-Based Testing to Bridge LLM Code Generation and Validation 10 upvotes, #9 of 2025-06-26
  10. Thought Anchors: Which LLM Reasoning Steps Matter? 9 upvotes, #10 of 2025-06-26
  11. When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs 8 upvotes, #11 of 2025-06-26
  12. Is There a Case for Conversation Optimized Tokenizers in Large Language Models? 7 upvotes, #12 of 2025-06-26
  13. GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching 7 upvotes, #12 of 2025-06-26
  14. ReCode: Updating Code API Knowledge with Reinforcement Learning 7 upvotes, #12 of 2025-06-26
  15. FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation 5 upvotes, #15 of 2025-06-26
  16. MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications 5 upvotes, #15 of 2025-06-26
  17. Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content 3 upvotes, #17 of 2025-06-26
  18. The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs 2 upvotes, #18 of 2025-06-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.