Daily Papers of 2025-04-23
- Kuwain 1.5B: An Arabic SLM via Language Injection 113 upvotes, #1 of 2025-04-23
- TTRL: Test-Time Reinforcement Learning 96 upvotes, #2 of 2025-04-23
- The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks 61 upvotes, #3 of 2025-04-23
- Describe Anything: Detailed Localized Image and Video Captioning 58 upvotes, #4 of 2025-04-23
- Learning Adaptive Parallel Reasoning with Language Models 42 upvotes, #5 of 2025-04-23
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale 32 upvotes, #6 of 2025-04-23
- BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation 26 upvotes, #7 of 2025-04-23
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs 22 upvotes, #8 of 2025-04-23
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities 20 upvotes, #9 of 2025-04-23
- Efficient Pretraining Length Scaling 19 upvotes, #10 of 2025-04-23
- Personalized Text-to-Image Generation with Auto-Regressive Models 18 upvotes, #11 of 2025-04-23
- WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents 18 upvotes, #11 of 2025-04-23
- CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning 17 upvotes, #13 of 2025-04-23
- Vidi: Large Multimodal Models for Video Understanding and Editing 15 upvotes, #14 of 2025-04-23
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning 15 upvotes, #14 of 2025-04-23
- RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild 9 upvotes, #16 of 2025-04-23
- Progent: Programmable Privilege Control for LLM Agents 7 upvotes, #17 of 2025-04-23
- MR. Video: "MapReduce" is the Principle for Long Video Understanding 6 upvotes, #18 of 2025-04-23
- CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting 5 upvotes, #19 of 2025-04-23
- IPBench: Benchmarking the Knowledge of Large Language Models in Intellectual Property 4 upvotes, #20 of 2025-04-23
- DiffVox: A Differentiable Model for Capturing and Analysing Professional Effects Distributions 2 upvotes, #21 of 2025-04-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.