Daily Papers of 2026-05-04

  1. UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors 80 upvotes, #1 of 2026-05-04
  2. Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction 38 upvotes, #2 of 2026-05-04
  3. Let ViT Speak: Generative Language-Image Pre-training 32 upvotes, #3 of 2026-05-04
  4. Map2World: Segment Map Conditioned Text to 3D World Generation 24 upvotes, #4 of 2026-05-04
  5. From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills 21 upvotes, #5 of 2026-05-04
  6. Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance 17 upvotes, #6 of 2026-05-04
  7. Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning 16 upvotes, #7 of 2026-05-04
  8. Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions 15 upvotes, #8 of 2026-05-04
  9. Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies 12 upvotes, #9 of 2026-05-04
  10. Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models 11 upvotes, #10 of 2026-05-04
  11. End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer 11 upvotes, #10 of 2026-05-04
  12. When Do Diffusion Models learn to Generate Multiple Objects? 8 upvotes, #12 of 2026-05-04
  13. Trees to Flows and Back: Unifying Decision Trees and Diffusion Models 7 upvotes, #13 of 2026-05-04
  14. Soft Anisotropic Diagrams for Differentiable Image Representation 5 upvotes, #14 of 2026-05-04
  15. MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks 5 upvotes, #14 of 2026-05-04
  16. AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval 3 upvotes, #16 of 2026-05-04
  17. Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling 3 upvotes, #16 of 2026-05-04
  18. Online Self-Calibration Against Hallucination in Vision-Language Models 3 upvotes, #16 of 2026-05-04
  19. Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization 3 upvotes, #16 of 2026-05-04
  20. Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring 3 upvotes, #16 of 2026-05-04
  21. LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation 2 upvotes, #21 of 2026-05-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.