Daily Papers of 2025-04-14
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model 119 upvotes, #1 of 2025-04-14
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation 47 upvotes, #2 of 2025-04-14
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft 38 upvotes, #3 of 2025-04-14
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model 30 upvotes, #4 of 2025-04-14
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning 25 upvotes, #5 of 2025-04-14
- PixelFlow: Pixel-Space Generative Models with Flow 17 upvotes, #6 of 2025-04-14
- ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration 17 upvotes, #6 of 2025-04-14
- CoRAG: Collaborative Retrieval-Augmented Generation 10 upvotes, #8 of 2025-04-14
- Do PhD-level LLMs Truly Grasp Elementary Addition? Probing Rule Learning vs. Memorization in Large Language Models 10 upvotes, #8 of 2025-04-14
- FlexIP: Dynamic Control of Preservation and Personality for Customized Image Generation 10 upvotes, #8 of 2025-04-14
- ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance 9 upvotes, #11 of 2025-04-14
- Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images 9 upvotes, #11 of 2025-04-14
- Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs 8 upvotes, #13 of 2025-04-14
- In-2-4D: Inbetweening from Two Single-View Images to 4D Generation 8 upvotes, #13 of 2025-04-14
- BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing 6 upvotes, #15 of 2025-04-14
- UKBOB: One Billion MRI Labeled Masks for Generalizable 3D Medical Image Segmentation 6 upvotes, #15 of 2025-04-14
- Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization 6 upvotes, #15 of 2025-04-14
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning 5 upvotes, #18 of 2025-04-14
- Latent Diffusion Autoencoders: Toward Efficient and Meaningful Unsupervised Representation Learning in Medical Imaging 5 upvotes, #18 of 2025-04-14
- InteractVLM: 3D Interaction Reasoning from 2D Foundational Models 4 upvotes, #20 of 2025-04-14
- SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs 4 upvotes, #20 of 2025-04-14
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.