Daily Papers of 2026-04-13

  1. WildDet3D: Scaling Promptable 3D Detection in the Wild 239 upvotes, #1 of 2026-04-13
  2. FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios 94 upvotes, #2 of 2026-04-13
  3. EXAONE 4.5 Technical Report 63 upvotes, #3 of 2026-04-13
  4. Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory 46 upvotes, #4 of 2026-04-13
  5. RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details 41 upvotes, #5 of 2026-04-13
  6. Multi-User Large Language Model Agents 26 upvotes, #6 of 2026-04-13
  7. ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion 22 upvotes, #7 of 2026-04-13
  8. ELT: Elastic Looped Transformers for Visual Generation 19 upvotes, #8 of 2026-04-13
  9. AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents 17 upvotes, #9 of 2026-04-13
  10. Envisioning the Future, One Step at a Time 12 upvotes, #10 of 2026-04-13
  11. Backdoor Attacks on Decentralised Post-Training 11 upvotes, #11 of 2026-04-13
  12. Structured Causal Video Reasoning via Multi-Objective Alignment 11 upvotes, #11 of 2026-04-13
  13. p1: Better Prompt Optimization with Fewer Prompts 9 upvotes, #13 of 2026-04-13
  14. ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery 9 upvotes, #13 of 2026-04-13
  15. VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images 8 upvotes, #15 of 2026-04-13
  16. Large Language Models Align with the Human Brain during Creative Thinking 6 upvotes, #16 of 2026-04-13
  17. Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video 6 upvotes, #16 of 2026-04-13
  18. Process Reward Agents for Steering Knowledge-Intensive Reasoning 6 upvotes, #16 of 2026-04-13
  19. Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism 6 upvotes, #16 of 2026-04-13
  20. Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance 5 upvotes, #20 of 2026-04-13
  21. Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models 5 upvotes, #20 of 2026-04-13
  22. Robust Reasoning Benchmark 4 upvotes, #22 of 2026-04-13
  23. On Semiotic-Grounded Interpretive Evaluation of Generative Art 4 upvotes, #22 of 2026-04-13
  24. EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers 4 upvotes, #22 of 2026-04-13
  25. MixFlow: Mixed Source Distributions Improve Rectified Flows 4 upvotes, #22 of 2026-04-13
  26. Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling 3 upvotes, #26 of 2026-04-13
  27. AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation 3 upvotes, #26 of 2026-04-13
  28. Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization 2 upvotes, #28 of 2026-04-13
  29. CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation 2 upvotes, #28 of 2026-04-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.