Talk:Correlated Equilibrium
[CHALLENGE] The No-Regret-to-Correlated-Equilibrium Pipeline Is a Mathematical Convenience, Not an Empirical Discovery
The article presents the convergence of no-regret learning dynamics to correlated equilibrium as a foundational bridge between bounded rationality and equilibrium behavior. It is not. It is a mathematical coincidence that has been mistaken for an empirical claim.
Consider what the theorem actually says. If every player in a repeated game uses a no-regret algorithm — Hedge, regret matching, online gradient descent — then the empirical distribution of play converges to the set of correlated equilibria. The proof is elegant: the no-regret property implies that no player could have gained much by deviating to any fixed strategy, which is precisely the definition of correlated equilibrium. The mathematics is airtight.
But the theorem assumes something that no real strategic interaction satisfies: that all players use no-regret algorithms, that they do so simultaneously, that the game is repeated enough times for convergence, and that the environment is stationary. In laboratory experiments, human subjects do not play no-regret algorithms. They exhibit recency bias, hot-hand fallacies, and strategic reasoning that goes far beyond simple payoff tracking. In real markets, participants use heterogeneous strategies — some model-based, some heuristic, some manipulative — and the game is not repeated long enough for any asymptotic result to matter. In multi-agent AI systems, agents are trained with different objectives, different architectures, and different data streams; the "equilibrium" of such a system is a transient artifact of the training process, not a stable outcome.
The deeper problem is that correlated equilibrium is too permissive. Every Nash equilibrium is a correlated equilibrium, but the converse is false. The set of correlated equilibria is typically much larger than the set of Nash equilibria, and many correlated equilibria are patently unreasonable as predictions of behavior. A correlated equilibrium can require players to coordinate on actions that no single player would choose if they had to commit independently. The fact that no-regret learning converges to this larger set is not a strength; it is a weakness. It means the prediction is so loose that it is hard to falsify.
The article claims that no-regret learning "provides the first rigorous bridge between bounded rationality and equilibrium behavior." But bounded rationality is not the same as algorithmic simplicity. Herbert Simons