A benchmark of language-model agent societies found that matching a human average at the beginning of a game often did not carry through to later rounds. Eight of 11 computable models matched the human anchor for public-goods contributions in round one, but none matched it by round 10. Elsewhere, one of 12 models matched the ultimatum-offer benchmark, while none matched dictator giving or cooperation in the repeated prisoner's dilemma. The results point to partial resemblance, not a general human-like social pattern.
The work is an arXiv preprint dated 28 August 2026 and is framed as a methods study. It tested 12 open-weight models in five human-anchored environments covering cooperation, public-goods contribution, bargaining, strategic reasoning and convention formation. The protocol produced 9,115 model runs and 150 classical-baseline runs, using one consumer graphics card. The human side of the comparison came through behavioral anchors rather than a newly recruited participant sample. The audit asked whether agents could match human levels and directions, withstand changes in the apparatus, respond to payoffs rather than memorized scripts, and meet its own certification standard.
To judge a match, the audit used pre-fixed equivalence margins rather than treating the absence of a clear difference as agreement. Two one-sided tests were run at an alpha level of 0.05. A model counted as equivalent only when the full 90% confidence interval for its difference from the human anchor stayed inside the margin set in advance. In ordinary terms, the model had to be close enough across the range of uncertainty, not merely avoid a statistically significant mismatch.
Averages faded across rounds
That distinction mattered over time. Only one of 12 models reproduced the human pattern of public-goods contributions declining across rounds. In the repeated prisoner's dilemma, three of 12 reproduced the corresponding decline. Convention formation, meaning settling on a shared choice or label, was the outlier: 10 of 12 models formed a convention. An early average could therefore look familiar even when the path through the game did not.
The punishment arm added another warning against reading a social signal as a human mechanism. Final-round contributions fell in four of nine computable models, by as much as 0.53 of the endowment. Agents nevertheless bought punishment in a median 97% of runs, spending 2.5 points per agent per round. Spending and contributions were negatively associated, with a Spearman correlation of -0.84. That is an association within simulated runs, not evidence that punishment caused the decline or would do so in human groups. Two models never bought punishment, and their contributions did not move.
The models also differed sharply before any perturbation. Baseline cooperation ranged from 0.00 to 1.00, with a between-model standard deviation of 0.36. A measure of between-model inconsistency, I-squared, reached 100% for four of the five primary outcomes and 90.8% in the 11 to 20 money-request game. The paper notes that I-squared can be uninformative at these sample sizes when within-model variation is small, but the spread still limits claims about language-model agents as a single class.
The presentation was part of the result
The audit tested whether behavior survived changes in the apparatus, or the way the same interaction was presented. Each cell changed exactly one factor. Representation-level cells re-rendered information already present, while design-level cells changed what agents were given. Design changes shifted behavior in 56 of 111 computable contrasts, compared with eight of 71 representation-level contrasts. Content sensitivity was more common than sensitivity to form, although the form effects could be large.
Swapping the order of action labels was associated with a 58-point cooperation difference for Qwen3-14B and a 40-point difference for Llama-3.1. In another case, putting persona information into a table changed Phi-3 convention formation from 0.90 to 0.00. These were model-specific effects from representation-only changes, making the format part of the observed behavior rather than a neutral wrapper around the game.
A mixed test of incentives
A fixed offer schedule provided a direct test of whether models responded to incentives rather than repeating a memorized script. Only one model placed its acceptance threshold where the outside option required, meaning where the payoff from rejecting an offer made sense. Two moved partway, two moved in the wrong direction, and three did not acquire a threshold. R1-Distill rejected offers from 10 to 37 in more than 95% of games, accepted 40 and above, and reached a payoff-consistency score of 0.98. The authors treat this as evidence of situation responsiveness in one reasoning-trained model, not as a property of the roster.
The naming-game result became more complicated when each agent received a shuffled pool of names. Seven of eight models still converged in every seed, but took 1.5 to 2.6 times longer. Qwen2.5-7B converged in two of three seeds and took 14 to 24 times as long as under the fixed order. The authors interpret the surviving convergence as evidence of a shared prior over labels and negotiation after the list-position cue was removed. They also say the mechanism did not match the human one, so convergence alone was not enough to establish human-like convention formation.
A ceiling, not a verdict
On the instrument's certification ladder, no finding reached Tier 3. Convention formation alone reached Tier 2, while every other finding remained at Tier 1. The study used a timestamped public deposit rather than a frozen OSF registration, so the ladder is best read as the stated framework of this audit, not an established field standard. The authors describe the overall result as a narrow validation ceiling and say only exploratory claims are currently supportable.
The conclusions are bounded by the sample. It covers a 12-model open-weight roster and does not establish how frontier models or larger reasoning systems would behave. R1-Distill completed only part of the battery, and the human comparison relied on anchors from particular populations rather than a broad test of individual or cross-cultural distributions. The preprint reports no external funding or competing interests. Its practical message is modest: matching a human-looking starting average is not enough to establish a human-like trajectory, institutional response or underlying mechanism.
Paper data and sources
Original title: Benchmarking large language model agent societies against human behavioural distributions
Authors: Raad Bin Tareaf
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text