Daily Papers of 2026-04-01

  1. CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence 336 upvotes, #1 of 2026-04-01
  2. FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization 325 upvotes, #2 of 2026-04-01
  3. LongCat-Next: Lexicalizing Modalities as Discrete Tokens 137 upvotes, #3 of 2026-04-01
  4. GEMS: Agent-Native Multimodal Generation with Memory and Skills 84 upvotes, #4 of 2026-04-01
  5. Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells 75 upvotes, #5 of 2026-04-01
  6. Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development 67 upvotes, #6 of 2026-04-01
  7. All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models 65 upvotes, #7 of 2026-04-01
  8. VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward 59 upvotes, #8 of 2026-04-01
  9. CutClaw: Agentic Hours-Long Video Editing via Music Synchronization 47 upvotes, #9 of 2026-04-01
  10. Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis 46 upvotes, #10 of 2026-04-01
  11. daVinci-LLM:Towards the Science of Pretraining 33 upvotes, #11 of 2026-04-01
  12. Extend3D: Town-Scale 3D Generation 25 upvotes, #12 of 2026-04-01
  13. Think Anywhere in Code Generation 25 upvotes, #12 of 2026-04-01
  14. MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models 21 upvotes, #14 of 2026-04-01
  15. RawGen: Learning Camera Raw Image Generation 20 upvotes, #15 of 2026-04-01
  16. Dynin-Omni: Omnimodal Unified Large Diffusion Language Model 18 upvotes, #16 of 2026-04-01
  17. FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration 16 upvotes, #17 of 2026-04-01
  18. Learn2Fold: Structured Origami Generation with World Model Planning 16 upvotes, #17 of 2026-04-01
  19. Meta-Harness: End-to-End Optimization of Model Harnesses 15 upvotes, #19 of 2026-04-01
  20. Falcon Perception 13 upvotes, #20 of 2026-04-01
  21. MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation 13 upvotes, #20 of 2026-04-01
  22. The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning 12 upvotes, #22 of 2026-04-01
  23. BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation 11 upvotes, #23 of 2026-04-01
  24. OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training 9 upvotes, #24 of 2026-04-01
  25. WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation 9 upvotes, #24 of 2026-04-01
  26. AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing 8 upvotes, #26 of 2026-04-01
  27. PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models 7 upvotes, #27 of 2026-04-01
  28. Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models 6 upvotes, #28 of 2026-04-01
  29. VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing 6 upvotes, #28 of 2026-04-01
  30. OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation 6 upvotes, #28 of 2026-04-01
  31. When Documents Disagree: Measuring Institutional Variation in Transplant Guidance with Retrieval-Augmented Language Models 5 upvotes, #31 of 2026-04-01
  32. ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions 5 upvotes, #31 of 2026-04-01
  33. CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions 5 upvotes, #31 of 2026-04-01
  34. TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets 5 upvotes, #31 of 2026-04-01
  35. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal 5 upvotes, #31 of 2026-04-01
  36. Tabular LLMs for Interpretable Few-Shot Alzheimer's Disease Prediction with Multimodal Biomedical Data 4 upvotes, #36 of 2026-04-01
  37. How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation 4 upvotes, #36 of 2026-04-01
  38. Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos 4 upvotes, #36 of 2026-04-01
  39. MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model 4 upvotes, #36 of 2026-04-01
  40. TrajectoryMover: Generative Movement of Object Trajectories in Videos 3 upvotes, #40 of 2026-04-01
  41. SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering 3 upvotes, #40 of 2026-04-01
  42. It Takes Two: A Duet of Periodicity and Directionality for Burst Flicker Removal 2 upvotes, #42 of 2026-04-01
  43. Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR 2 upvotes, #42 of 2026-04-01

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.