A computer simulation of networked language models found that competing persuaders were associated with more unsettled belief paths than a lone affirmative source, while the relationship between network density and upward movement reversed between the two settings. The result highlights how influence can look inside an LLM network, but it is not evidence that people would respond the same way: the work followed simulated agents in a closed environment.
Inside the simulation
At the center of the study was a controlled testbed in which simulated persuaders and persuadees repeatedly interacted around one declarative proposition. The analysis tracked each persuadee’s belief trajectory round by round rather than relying only on a final response. The full factorial corpus crossed two persuader settings, five graphs, four language-model models, 55 English public-policy statements and two random seeds, yielding 4,400 independent runs over a fixed 10-round horizon.
In the singlePR setting, one affirmative persuader was pinned to the positive end of the belief scale. In dualPR, that affirmative agent, PR1, faced an opposing agent, PR2, pinned to the negative end. After every round, persuadees answered a seven-point multiple-choice belief probe based on token probabilities. The system also logged exposure, rationale, action, belief-check and summary events, giving researchers a record of both the score and the route by which messages arrived.
Different settings, different paths
On average, the single-persuader runs moved topic-level beliefs away from each model’s no-context prior, its preference without network context, and toward the affirmative persuader. In the competitive setting, terminal means stayed closer to that prior. Topic ordering remained broadly stable across backbones, suggesting that competition was associated with a difference in how far trajectories traveled more than in which topics ranked higher or lower.
The network pattern was less straightforward. The researchers counted an upward trajectory as a belief increase of at least 0.10 over 10 rounds. In dense graph G3, 22% to 52% of trajectories met that threshold under singlePR, compared with 69% to 83% in sparser graphs. Under dualPR, dense G3 reached 49% to 68%, while sparser graphs ranged from 37% to 55%. In this selected set of graphs, the association between density and upward movement therefore changed direction when competition was present.
The shape of the movement changed as well. In GPT-4o runs, the biggest singlePR category was converted_pro, an affirmative-side conversion, at 53.4%. Under dualPR, converted_con, an opposing-side conversion, reached 17.1%; the share ending in the con band rose from 4.5% to 25.5%, and back-and-forth tug_of_war paths accounted for 57.8% of neutral-start trajectories. Gemini-2.5-Pro recorded an even higher dualPR tug_of_war share of 69.6%.
Influence was direct and indirect
To examine direct contact alongside relays through other persuadees, the researchers ran ordinary least squares regressions, a way of estimating the relationship between exposure counts and next-round belief change. The unit of analysis was one persuadee in one round, and the models used heteroskedasticity-robust standard errors. In dualPR, direct PR1 exposure had positive coefficients from +0.0025 to +0.0055 across the four backbones; direct PR2 exposure had negative coefficients from -0.0069 to -0.0037. With about 50 direct exposures per side over ten rounds, the reported cumulative associations were +0.16 to +0.26 for PR1 and -0.22 to -0.32 for PR2.
Peer-mediated exposure showed a smaller and more specification-sensitive signal. PR2’s peer-mediated coefficient was significantly negative on all four backbones, while PR1’s was significantly positive on two of four. The cumulative peer-PR2 association fell between -0.015 and -0.053. Because exposure was generated by the simulation rather than randomly assigned, these results describe associations with movement, not proof that any channel caused it.
The words were not a complete record
The message logs added another wrinkle. The study used Cialdini labels, tags for six rhetorical principles, to classify persuasive tactics, but only 72.7% of the labels planned by persuaders appeared in their generated text. In dualPR, PR1 and PR2 differed by 14% to 27% in text mix for some principles, even though their action-type gaps were no more than 7%. Comments accounted for 49% to 60% of persuader actions. Plans, wording and action choices were therefore not interchangeable signals.
Persuadee language also rarely announced a change in direction. By rounds 2 and 3, 94.0% of persuadee messages shared a Cialdini label with a principle used earlier by a persuader. Hedging was classified as low in 73% and high in 10%, while explicit reversals appeared in 0.18% and softening in 12.4%. That combination meant that persuadees often echoed persuasive material without plainly stating the stance shift indicated by the probe.
The language labels were useful, but imperfect. In a human validation check of 600 judgments, the classifier was judged correct 87.3% of the time. It scored 86.3% on GPT-generated messages and 88.3% on Gemini-generated messages, and each of the six principles exceeded 81%. Those figures make the labels a diagnostic within the simulation, not an infallible readout of what an agent believed.
What the result cannot settle
There is also a structural limit on the starting point. Initial persuadee beliefs came from Personalized PageRank, a network-ranking method, which entangled each agent’s network location with its initial stance. The paper leaves random or uniform prior assignment for a future ablation. Until that is tested, the density comparisons cannot be treated as clean, independent tests of network structure.
More broadly, the work covers a closed simulation of four English-language LLM backbones, a fixed action vocabulary and no external tools. It does not validate these dynamics against humans, place agents on real platforms, or involve human subjects or personal data. The trajectory shares and strategy measures are descriptive for this run set, and the study does not establish generalization beyond its models, English policy statements, fixed action space and selected graphs.
For evaluations of LLM-agent networks, the study’s evidence points to a narrow lesson: final text and direct output may miss peer relays or changes visible only in a round-by-round probe. The work is an arXiv version 1 preprint dated 25 August 2026, and its findings remain confined to the simulated setting.
Paper data and sources
Original title: Belief Cascades Drive Persuasion in LLM Agent Networks
Authors: Haoyi Qiu, Genglin Liu, Pranav Narayanan Venkit et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text