Daily Papers of 2025-04-04
- Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems 230 upvotes, #1 of 2025-04-04
- ZClip: Adaptive Spike Mitigation for LLM Pre-Training 74 upvotes, #2 of 2025-04-04
- Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing 67 upvotes, #3 of 2025-04-04
- GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation 54 upvotes, #4 of 2025-04-04
- Inference-Time Scaling for Generalist Reward Modeling 50 upvotes, #5 of 2025-04-04
- JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization 49 upvotes, #6 of 2025-04-04
- WikiVideo: Article Generation from Multiple Videos 36 upvotes, #7 of 2025-04-04
- Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation 36 upvotes, #7 of 2025-04-04
- SkyReels-A2: Compose Anything in Video Diffusion Transformers 34 upvotes, #9 of 2025-04-04
- Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme 30 upvotes, #10 of 2025-04-04
- Scaling Analysis of Interleaved Speech-Text Language Models 27 upvotes, #11 of 2025-04-04
- ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers 21 upvotes, #12 of 2025-04-04
- FreSca: Unveiling the Scaling Space in Diffusion Models 17 upvotes, #13 of 2025-04-04
- Efficient Model Selection for Time Series Forecasting via LLMs 16 upvotes, #14 of 2025-04-04
- Scaling Laws in Scientific Discovery with AI and Robot Scientists 12 upvotes, #15 of 2025-04-04
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning 12 upvotes, #15 of 2025-04-04
- NeuralGS: Bridging Neural Fields and 3D Gaussian Splatting for Compact 3D Representations 11 upvotes, #17 of 2025-04-04
- Interpreting Emergent Planning in Model-Free Reinforcement Learning 11 upvotes, #17 of 2025-04-04
- OpenCodeReasoning: Advancing Data Distillation for Competitive Coding 11 upvotes, #17 of 2025-04-04
- Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages 10 upvotes, #20 of 2025-04-04
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models 10 upvotes, #20 of 2025-04-04
- Instruction-Guided Autoregressive Neural Network Parameter Generation 6 upvotes, #22 of 2025-04-04
- Scene-Centric Unsupervised Panoptic Segmentation 5 upvotes, #23 of 2025-04-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.