Daily Papers of 2026-03-12

  1. Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning 204 upvotes, #1 of 2026-03-12
  2. OpenClaw-RL: Train Any Agent Simply by Talking 130 upvotes, #2 of 2026-03-12
  3. Flash-KMeans: Fast and Memory-Efficient Exact K-Means 78 upvotes, #3 of 2026-03-12
  4. LLM2Vec-Gen: Generative Embeddings from Large Language Models 40 upvotes, #4 of 2026-03-12
  5. In-Context Reinforcement Learning for Tool Use in Large Language Models 39 upvotes, #5 of 2026-03-12
  6. MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents 28 upvotes, #6 of 2026-03-12
  7. ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuning 25 upvotes, #7 of 2026-03-12
  8. ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA 19 upvotes, #8 of 2026-03-12
  9. Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams 17 upvotes, #9 of 2026-03-12
  10. SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing 15 upvotes, #10 of 2026-03-12
  11. CodePercept: Code-Grounded Visual STEM Perception for MLLMs 13 upvotes, #11 of 2026-03-12
  12. RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback 12 upvotes, #12 of 2026-03-12
  13. Lost in Backpropagation: The LM Head is a Gradient Bottleneck 11 upvotes, #13 of 2026-03-12
  14. Prism-Δ: Differential Subspace Steering for Prompt Highlighting in Large Language Models 11 upvotes, #13 of 2026-03-12
  15. V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts 10 upvotes, #15 of 2026-03-12
  16. Just-in-Time: Training-Free Spatial Acceleration for Diffusion Transformers 7 upvotes, #16 of 2026-03-12
  17. RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation 6 upvotes, #17 of 2026-03-12
  18. According to Me: Long-Term Personalized Referential Memory QA 5 upvotes, #18 of 2026-03-12
  19. Hindsight Credit Assignment for Long-Horizon LLM Agents 5 upvotes, #18 of 2026-03-12
  20. Meissa: Multi-modal Medical Agentic Intelligence 5 upvotes, #18 of 2026-03-12
  21. CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR 5 upvotes, #18 of 2026-03-12
  22. UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations 4 upvotes, #22 of 2026-03-12
  23. COMIC: Agentic Sketch Comedy Generation 4 upvotes, #22 of 2026-03-12
  24. Causal Concept Graphs in LLM Latent Space for Stepwise Reasoning 3 upvotes, #24 of 2026-03-12
  25. V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation 3 upvotes, #24 of 2026-03-12
  26. Any to Full: Prompting Depth Anything for Depth Completion in One Stage 2 upvotes, #26 of 2026-03-12
  27. Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models 2 upvotes, #26 of 2026-03-12
  28. EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation 1 upvotes, #28 of 2026-03-12
  29. StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving 1 upvotes, #29 of 2026-03-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.