Daily Papers of 2026-03-12
- Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning 204 upvotes, #1 of 2026-03-12
- OpenClaw-RL: Train Any Agent Simply by Talking 130 upvotes, #2 of 2026-03-12
- Flash-KMeans: Fast and Memory-Efficient Exact K-Means 78 upvotes, #3 of 2026-03-12
- LLM2Vec-Gen: Generative Embeddings from Large Language Models 40 upvotes, #4 of 2026-03-12
- In-Context Reinforcement Learning for Tool Use in Large Language Models 39 upvotes, #5 of 2026-03-12
- MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents 28 upvotes, #6 of 2026-03-12
- ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuning 25 upvotes, #7 of 2026-03-12
- ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA 19 upvotes, #8 of 2026-03-12
- Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams 17 upvotes, #9 of 2026-03-12
- SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing 15 upvotes, #10 of 2026-03-12
- CodePercept: Code-Grounded Visual STEM Perception for MLLMs 13 upvotes, #11 of 2026-03-12
- RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback 12 upvotes, #12 of 2026-03-12
- Lost in Backpropagation: The LM Head is a Gradient Bottleneck 11 upvotes, #13 of 2026-03-12
- Prism-Δ: Differential Subspace Steering for Prompt Highlighting in Large Language Models 11 upvotes, #13 of 2026-03-12
- V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts 10 upvotes, #15 of 2026-03-12
- Just-in-Time: Training-Free Spatial Acceleration for Diffusion Transformers 7 upvotes, #16 of 2026-03-12
- RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation 6 upvotes, #17 of 2026-03-12
- According to Me: Long-Term Personalized Referential Memory QA 5 upvotes, #18 of 2026-03-12
- Hindsight Credit Assignment for Long-Horizon LLM Agents 5 upvotes, #18 of 2026-03-12
- Meissa: Multi-modal Medical Agentic Intelligence 5 upvotes, #18 of 2026-03-12
- CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR 5 upvotes, #18 of 2026-03-12
- UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations 4 upvotes, #22 of 2026-03-12
- COMIC: Agentic Sketch Comedy Generation 4 upvotes, #22 of 2026-03-12
- Causal Concept Graphs in LLM Latent Space for Stepwise Reasoning 3 upvotes, #24 of 2026-03-12
- V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation 3 upvotes, #24 of 2026-03-12
- Any to Full: Prompting Depth Anything for Depth Completion in One Stage 2 upvotes, #26 of 2026-03-12
- Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models 2 upvotes, #26 of 2026-03-12
- EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation 1 upvotes, #28 of 2026-03-12
- StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving 1 upvotes, #29 of 2026-03-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.