Daily Papers of 2026-05-04
- UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors 80 upvotes, #1 of 2026-05-04
- Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction 38 upvotes, #2 of 2026-05-04
- Let ViT Speak: Generative Language-Image Pre-training 32 upvotes, #3 of 2026-05-04
- Map2World: Segment Map Conditioned Text to 3D World Generation 24 upvotes, #4 of 2026-05-04
- From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills 21 upvotes, #5 of 2026-05-04
- Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance 17 upvotes, #6 of 2026-05-04
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning 16 upvotes, #7 of 2026-05-04
- Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions 15 upvotes, #8 of 2026-05-04
- Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies 12 upvotes, #9 of 2026-05-04
- Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models 11 upvotes, #10 of 2026-05-04
- End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer 11 upvotes, #10 of 2026-05-04
- When Do Diffusion Models learn to Generate Multiple Objects? 8 upvotes, #12 of 2026-05-04
- Trees to Flows and Back: Unifying Decision Trees and Diffusion Models 7 upvotes, #13 of 2026-05-04
- Soft Anisotropic Diagrams for Differentiable Image Representation 5 upvotes, #14 of 2026-05-04
- MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks 5 upvotes, #14 of 2026-05-04
- AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval 3 upvotes, #16 of 2026-05-04
- Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling 3 upvotes, #16 of 2026-05-04
- Online Self-Calibration Against Hallucination in Vision-Language Models 3 upvotes, #16 of 2026-05-04
- Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization 3 upvotes, #16 of 2026-05-04
- Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring 3 upvotes, #16 of 2026-05-04
- LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation 2 upvotes, #21 of 2026-05-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.