Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
Alexander, Mariia Trofimova, Sergei Polezhaev, Ibragim, Maksim Nekrashevich, Anton, Simon Karasik, Sergey Abramov, Andrei Andriushchenko, Filipp Fisin, Sergei, Boris Yangel
Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning: 48 upvotes on Hugging Face Daily Papers, #5 of 38 papers on 2025-08-07. Day-by-day upvote history. It lost 11 votes when the Hub removed votes in bulk.
Research on applications of Reinforcement Learning (RL) to Large Language Models (LLMs) has mostly been focused on single-turn problems, such as mathematical reasoning or single-shot code generation. While these problems can be viewed as token-level multi-turn MDPs, this view corresponds to a degenerate case of multi-turn interaction where the environment provides no feedback. This contrasts with many real-world domains, such as software engineering (SWE), which require rich multi-turn interactions with a stateful environment that responds to each action with a non-trivial observation. To bridge this gap, we demonstrate the successful application of RL to this general regime. Using a modified Decoupled Advantage Policy Optimization (DAPO) algorithm, we train an agent based on Qwen2.5-72B-Instruct to solve real-world software engineering tasks. Our approach increases the agent's success rate on the SWE-bench Verified benchmark from a 20% rejection fine-tuned baseline to 39%, without relying on any teacher models. On SWE-rebench, our agent matches or outperforms leading open-weight models such as DeepSeek-V3-0324 and Qwen3-235B-A22B using an identical scaffolding, offering a viable path toward building more capable autonomous agents for complex real-world problems based on open models.
Paper page on Hugging Face · arXiv
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.