Preprint

Preprint: Simulation reports abrupt cooperation collapse as defection becomes more tempting

A lattice-based model links reputation-weighted learning with sharp shifts, competing outcomes and the growth or extinction of cooperative clusters.

A computer simulation reports an abrupt break in cooperation as its temptation parameter rose: the model could evolve toward full cooperation when b was 0.42 or lower, but above that range it shifted abruptly from full cooperation to full defection.

The result came from synthetic agents arranged on a periodic square lattice, a grid that wraps around at the edges. Each agent interacted with four nearest neighbours in repeated prisoner's dilemma games. The work is an arXiv preprint, so the finding is limited to the simulated setup rather than being a direct observation of people or animals.

Reputation changes the learning signal

The agents used reinforcement learning, a trial-and-error approach in which action preferences are updated from rewards. In the model, cooperation raised an agent's reputation and defection lowered it; local reputation was used to set how much weight the learning reward gave to the agent's own payoff and the average payoff of its neighbours.

Near the transition, the model could settle into two different long-term outcomes. At b = 0.47, random initial conditions were associated with either a cooperative outcome or a defective one. In reported trajectories, the cooperator share ended near zero in some runs and around 0.8 in others; some trajectories first entered a near-fully defective state before recovering after a long delay.

At b = 0.47, the steady-state reputation distribution had two distinct peaks, separating persistent cooperative and defective groups.

Why the outcomes split

Snapshots suggested a spatial explanation for the split. Cooperative domains grew through nucleation—the formation and expansion of local patches—then merged and spread. When nucleation failed, clusters shrank and vanished, and the system became fully defective.

The learned Q-tables, or tables of action preferences for different states, also differed between outcomes. In high-cooperation runs, the average preference favoured cooperation in states 2 through 5; in low-cooperation runs, it favoured defection in every state.

The result depends on the setup

The outcome also depended on the learning settings. Cooperation was concentrated in the region of a small learning rate α and a large discount factor γ; with γ fixed at 0.9, increasing α was associated with movement from high cooperation toward full defection.

The authors also used a pair-approximation calculation, a simplified theory that follows neighbouring pairs while omitting higher-order spatial correlations and differences among Q-values. It qualitatively reproduced the discontinuous transition and hysteresis—different outcomes tied to the system's path—and predicted a simulation threshold near b ≈ 0.45.

A result for a model, not a universal rule

The simulations used a default 100 × 100 lattice, and reported observables were averaged after a transient across at least 20 independent realizations. No confidence intervals or formal statistical uncertainty estimates were reported, and outcomes near the transition depended on initial conditions.

Those limits matter: the threshold and cluster pathway are properties of this stylized lattice model. The study does not establish that the same behaviour would appear under another network, payoff structure or learning rule, or in real human or animal groups.

The paper reports that code for generating its key results is available on GitHub.

Paper data and sources

Original title: Emergence of cooperation: A reputation-modulated reinforcement learning
Authors: Chenyang Zhao, Jiqiang Zhang, Li Chen, Yong Zou
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.