Learning Cooperative Multi-Agent Policies with Partial Reward Decoupling

Benjamin Freed, Aditya Kapoor, Ian Abraham, Jeff Schneider, Howie Choset
<span title="">2021</span> <i title="Institute of Electrical and Electronics Engineers (IEEE)"> <a target="_blank" rel="noopener" href="https://fatcat.wiki/container/g32gcto57vduveq2eb7bes5y7a" style="color: black;">IEEE Robotics and Automation Letters</a> </i> &nbsp;
One of the preeminent obstacles to scaling multi-agent reinforcement learning to large numbers of agents is assigning credit to individual agents' actions. In this paper, we address this credit assignment problem with an approach that we call partial reward decoupling (PRD), which attempts to decompose large cooperative multi-agent RL problems into decoupled subproblems involving subsets of agents, thereby simplifying credit assignment. We empirically demonstrate that decomposing the RL problem
more &raquo; ... using PRD in an actor-critic algorithm results in lower variance policy gradient estimates, which improves data efficiency, learning stability, and asymptotic performance across a wide array of multi-agent RL tasks, compared to various other actor-critic approaches. Additionally, we relate our approach to counterfactual multi-agent policy gradient (COMA), a state-of-the-art MARL algorithm, and empirically show that our approach outperforms COMA by making better use of information in agents' reward streams, and by enabling recent advances in advantage estimation to be used.
<span class="external-identifiers"> <a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1109/lra.2021.3135930">doi:10.1109/lra.2021.3135930</a> <a target="_blank" rel="external noopener" href="https://fatcat.wiki/release/jcvm5imgifcjphqikov74t7xu4">fatcat:jcvm5imgifcjphqikov74t7xu4</a> </span>
<a target="_blank" rel="noopener" href="https://web.archive.org/web/20220104192740/https://arxiv.org/pdf/2112.12740v1.pdf" title="fulltext PDF download [not primary version]" data-goatcounter-click="serp-fulltext" data-goatcounter-title="serp-fulltext"> <button class="ui simple right pointing dropdown compact black labeled icon button serp-button"> <i class="icon ia-icon"></i> Web Archive [PDF] <span style="color: #f43e3e;">&#10033;</span> <div class="menu fulltext-thumbnail"> <img src="https://blobs.fatcat.wiki/thumbnail/pdf/9a/f1/9af16c96e83cfce38af4a15de0d88f30aa409f49.180px.jpg" alt="fulltext thumbnail" loading="lazy"> </div> </button> </a> <a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1109/lra.2021.3135930"> <button class="ui left aligned compact blue labeled icon button serp-button"> <i class="external alternate icon"></i> ieee.com </button> </a>