NeurIPS 2026/MARS
Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning
MAPPO and MASPO use clipping and quadratic penalties, respectively, to control updates in cooperative RL. Clipping can discard gradients from outlier samples, while the quadratic penalty can still allow action-probability collapse. MARS replaces these mechanisms with a smooth barrier symmetric under reciprocal likelihood changes. Across 47 tasks in eight environments, it matches or exceeds both methods in aggregate performance.
01 / Motivation
Two flaws in MAPPO and MASPO.
In cooperative policy-gradient RL, each agent’s update is weighted by a joint advantage estimate: team returns relative to a centralized critic’s baseline. Simultaneously changing teammate policies can make these estimates noisy. A negative advantage may reflect temporary miscoordination rather than an action that should be eliminated from an agent’s policy.
MAPPO’s hard clipping removes gradients once a sample’s policy likelihood ratio crosses a boundary in the improving direction. This weakens recovery from policy drift across minibatch updates. MASPO retains gradients with a quadratic penalty, but that penalty remains finite as the ratio approaches zero. Large negative advantages can therefore drive a sampled action’s likelihood toward zero.
02 / Method
A symmetric barrier for policy updates.
MARS uses a smooth penalty on the ratio of action likelihoods under the new and old policies. The penalty is invariant under ratio inversion: reciprocal changes, such as doubling and halving an action’s likelihood, incur the same cost. MAPPO and MASPO instead regulate ratios using additive distance from 1.
The penalty diverges as the ratio approaches zero, creating a geometric barrier against probability collapse. It also preserves gradient information for outlier samples, avoiding the flat regions introduced by clipping. Training uses a centralized critic; each agent’s policy executes from its own observations.
03 / Contributions
What this work adds.
- 01
A policy objective with a symmetric barrier
MARS replaces additive ratio control with a smooth penalty that is symmetric under inversion and diverges as the ratio approaches zero.
- 02
An analysis of clipping and collapse
A per-sample analysis shows that the objective avoids clipping’s flat-gradient regions and makes probability-ratio extinction non-optimal.
- 03
New benchmarks and a 47-task evaluation
The evaluation covers 47 tasks across eight environments, including the new JAX benchmarks PaxMen and AeroJAX. Ablations isolate penalty geometry from the choice of update boundaries.
04 / Evidence
47 tasks across eight environments.
The evaluation spans simulated aerial combat in AeroJAX, navigation in JaxNav, warehouse coordination in RWARE, exploration in PaxMen, and other cooperative tasks. MARS matches or exceeds MAPPO and MASPO in aggregate performance within each environment.
Trust-region ablations separate the penalty’s geometry from the freedom to choose independent expansion and contraction targets. The results point to the symmetric barrier as the main source of ratio stability; flexible boundaries primarily improve learning speed.
Ratio stability during training
These AeroJAX 8v8 ablations compare MARS variants with asymmetric MAPPO and MASPO, which allow independently tuned expansion and contraction limits. The minimum- and maximum-ratio diagnostics isolate whether boundary flexibility alone can stabilize updates.
- MARS
- Asymmetric MAPPO
- Asymmetric MASPO
- Multiplicative Symmetric MARS
- Additive Symmetric MARS
Task scores are normalized before aggregation within each environment. Curves average ten independent runs, with 95% confidence intervals. Insets estimate the aggregate probability that MARS outperforms MAPPO or MASPO.
Scope and limitations
The analysis concerns the per-sample objective in ratio space; it does not establish global convergence or monotonic improvement for neural multi-agent policy updates. Performance claims apply to the cooperative benchmarks and training protocol in the paper.
05 / Citation
Cite this paper.
@misc{wijesundara2026rethinking,
title = {Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning},
author = {Chulabhaya Wijesundara and Andrea Baisero and Zhongheng Li and Gregory Castañón and Alan Carlin and Christopher Amato},
year = {2026},
eprint = {2605.09212},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2605.09212}
}This citation uses the arXiv manuscript. A NeurIPS proceedings citation will be added when available.