← All papers

NeurIPS 2026/MARS

Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning

Chulabhaya Wijesundara·Andrea Baisero·Zhongheng Li·Gregory Castañón·Alan Carlin·Christopher Amato

The short version

MAPPO and MASPO use clipping and quadratic penalties, respectively, to control updates in cooperative RL. Clipping can discard gradients from outlier samples, while the quadratic penalty can still allow action-probability collapse. MARS replaces these mechanisms with a smooth barrier symmetric under reciprocal likelihood changes. Across 47 tasks in eight environments, it matches or exceeds both methods in aggregate performance.

01 / Motivation

Two flaws in MAPPO and MASPO.

In cooperative policy-gradient RL, each agent’s update is weighted by a joint advantage estimate: team returns relative to a centralized critic’s baseline. Simultaneously changing teammate policies can make these estimates noisy. A negative advantage may reflect temporary miscoordination rather than an action that should be eliminated from an agent’s policy.

MAPPO’s hard clipping removes gradients once a sample’s policy likelihood ratio crosses a boundary in the improving direction. This weakens recovery from policy drift across minibatch updates. MASPO retains gradients with a quadratic penalty, but that penalty remains finite as the ratio approaches zero. Large negative advantages can therefore drive a sampled action’s likelihood toward zero.

02 / Method

A symmetric barrier for policy updates.

MARS uses a smooth penalty on the ratio of action likelihoods under the new and old policies. The penalty is invariant under ratio inversion: reciprocal changes, such as doubling and halving an action’s likelihood, incur the same cost. MAPPO and MASPO instead regulate ratios using additive distance from 1.

The penalty diverges as the ratio approaches zero, creating a geometric barrier against probability collapse. It also preserves gradient information for outlier samples, avoiding the flat regions introduced by clipping. Training uses a centralized critic; each agent’s policy executes from its own observations.

Three policy-objective plots compare MAPPO’s flat clipping regions, MASPO’s finite quadratic penalty, and MARS’s barrier near a probability ratio of zero.
Policy objectives plotted against the likelihood ratio. MAPPO’s clipping creates flat-gradient regions. MASPO’s quadratic penalty remains finite as the ratio approaches zero. MARS preserves gradients for ratio outliers while imposing an unbounded barrier near zero. Original figure (PDF)

03 / Contributions

What this work adds.

  1. 01

    A policy objective with a symmetric barrier

    MARS replaces additive ratio control with a smooth penalty that is symmetric under inversion and diverges as the ratio approaches zero.

  2. 02

    An analysis of clipping and collapse

    A per-sample analysis shows that the objective avoids clipping’s flat-gradient regions and makes probability-ratio extinction non-optimal.

  3. 03

    New benchmarks and a 47-task evaluation

    The evaluation covers 47 tasks across eight environments, including the new JAX benchmarks PaxMen and AeroJAX. Ablations isolate penalty geometry from the choice of update boundaries.

04 / Evidence

47 tasks across eight environments.

The evaluation spans simulated aerial combat in AeroJAX, navigation in JaxNav, warehouse coordination in RWARE, exploration in PaxMen, and other cooperative tasks. MARS matches or exceeds MAPPO and MASPO in aggregate performance within each environment.

Trust-region ablations separate the penalty’s geometry from the freedom to choose independent expansion and contraction targets. The results point to the symmetric barrier as the main source of ratio stability; flexible boundaries primarily improve learning speed.

Ratio stability during training

These AeroJAX 8v8 ablations compare MARS variants with asymmetric MAPPO and MASPO, which allow independently tuned expansion and contraction limits. The minimum- and maximum-ratio diagnostics isolate whether boundary flexibility alone can stabilize updates.

  • MARS
  • Asymmetric MAPPO
  • Asymmetric MASPO
  • Multiplicative Symmetric MARS
  • Additive Symmetric MARS
In AeroJAX, minimum action-probability ratios fall most strongly for asymmetric MASPO and later for asymmetric MAPPO. All three MARS variants stay higher.
Asymmetric MASPO (green) shows the largest decline in minimum likelihood ratios; asymmetric MAPPO (blue) also drifts downward. MARS variants retain higher, more stable minima, indicating less severe suppression of sampled actions. Original figure (PDF)
In AeroJAX, maximum action-probability ratios spike for asymmetric MASPO, while the other methods remain much closer to 1.
Asymmetric MASPO (green) develops sharp maximum-ratio spikes. MARS variants remain closer to 1, limiting large relative increases in action likelihood. Original figure (PDF)
Scope and limitations

The analysis concerns the per-sample objective in ratio space; it does not establish global convergence or monotonic improvement for neural multi-agent policy updates. Performance claims apply to the cooperative benchmarks and training protocol in the paper.

Read the full paper

05 / Citation

Cite this paper.

BibTeX
Download .bib
@misc{wijesundara2026rethinking,
  title = {Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning},
  author = {Chulabhaya Wijesundara and Andrea Baisero and Zhongheng Li and Gregory Castañón and Alan Carlin and Christopher Amato},
  year = {2026},
  eprint = {2605.09212},
  archivePrefix = {arXiv},
  primaryClass = {cs.LG},
  url = {https://arxiv.org/abs/2605.09212}
}

This citation uses the arXiv manuscript. A NeurIPS proceedings citation will be added when available.