AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Constant Swap Regret in General-Sum Games via Two-Scale Higher-Order Optimism

arXiv · AI, language, vision and robotics · article · Sep 15, 2026 · UTC

We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret in self-play, independent of the horizon $T$. With $n$ players and at most $m$ actions each, every player's individual swap regret is $O(\sqrt n\,m\log m\log^{5/2}(nm))$ at every finite horizon. The dynamics use the classical Blum-Mansour framework with optimism. Each player predicts the deviation gains, uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. Our new ingred

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:01:24.920Z. This is not the publication date.

Observed changes

AIIC observation times, not verified publisher revision times. Up to eight recent revisions.

2026-09-23T23:32:26.314Z

  • title: Constant Swap Regret in General-Sum Games via Optimistic Transition Matrices → Constant Swap Regret in General-Sum Games via Two-Scale Higher-Order Optimism
  • summary: We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret, independent of the horizon $T$. With $n$ players and at most $m$ actions each, the individual swap regret of every player is $O(\sqrt{n} m \log m \log^{5/2}(nm))$ at every finite horizon. Each player predicts the deviation gains, then uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. The proof combines a potential argument exploiting stationarity with a two-scale hig → We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret in self-play, independent of the horizon $T$. With $n$ players and at most $m$ actions each, every player's individual swap regret is $O(\sqrt n\,m\log m\log^{5/2}(nm))$ at every finite horizon. The dynamics use the classical Blum-Mansour framework with optimism. Each player predicts the deviation gains, uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. Our new ingred