Daily Papers of 2026-04-13
- WildDet3D: Scaling Promptable 3D Detection in the Wild 239 upvotes, #1 of 2026-04-13
- FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios 94 upvotes, #2 of 2026-04-13
- EXAONE 4.5 Technical Report 63 upvotes, #3 of 2026-04-13
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory 46 upvotes, #4 of 2026-04-13
- RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details 41 upvotes, #5 of 2026-04-13
- Multi-User Large Language Model Agents 26 upvotes, #6 of 2026-04-13
- ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion 22 upvotes, #7 of 2026-04-13
- ELT: Elastic Looped Transformers for Visual Generation 19 upvotes, #8 of 2026-04-13
- AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents 17 upvotes, #9 of 2026-04-13
- Envisioning the Future, One Step at a Time 12 upvotes, #10 of 2026-04-13
- Backdoor Attacks on Decentralised Post-Training 11 upvotes, #11 of 2026-04-13
- Structured Causal Video Reasoning via Multi-Objective Alignment 11 upvotes, #11 of 2026-04-13
- p1: Better Prompt Optimization with Fewer Prompts 9 upvotes, #13 of 2026-04-13
- ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery 9 upvotes, #13 of 2026-04-13
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images 8 upvotes, #15 of 2026-04-13
- Large Language Models Align with the Human Brain during Creative Thinking 6 upvotes, #16 of 2026-04-13
- Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video 6 upvotes, #16 of 2026-04-13
- Process Reward Agents for Steering Knowledge-Intensive Reasoning 6 upvotes, #16 of 2026-04-13
- Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism 6 upvotes, #16 of 2026-04-13
- Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance 5 upvotes, #20 of 2026-04-13
- Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models 5 upvotes, #20 of 2026-04-13
- Robust Reasoning Benchmark 4 upvotes, #22 of 2026-04-13
- On Semiotic-Grounded Interpretive Evaluation of Generative Art 4 upvotes, #22 of 2026-04-13
- EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers 4 upvotes, #22 of 2026-04-13
- MixFlow: Mixed Source Distributions Improve Rectified Flows 4 upvotes, #22 of 2026-04-13
- Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling 3 upvotes, #26 of 2026-04-13
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation 3 upvotes, #26 of 2026-04-13
- Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization 2 upvotes, #28 of 2026-04-13
- CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation 2 upvotes, #28 of 2026-04-13
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.