Preprint

AI models miss subtle power cues in French and Egyptian Arabic dialogue

Preprint: A benchmark of movie scenes finds models closer to human labels on explicit cues than on relational, culturally grounded judgments.

Language models were closer to humans on obvious social cues than on the less visible forces that shape a conversation, according to a benchmark using French and Egyptian Arabic movie dialogue. Across the test, models tracked demographic and contextual details more closely than emotions, speaker dynamics and socially interpretive features. In Egyptian Arabic, text-model agreement on religion ranged from 0.19 to 0.53, while GPT-5.1 matched human labels for French social-status difference at only 0.12 to 0.15.

The figures are agreement comparisons between model outputs and human annotations. The dataset has only two languages, and the authors explicitly do not claim that particular power types are associated with those cultures.

A benchmark built from dialogue

The team built a three-part framework: a theoretically grounded schema, a native-speaker annotation process refined through pilot studies, and a custom interface for cross-lingual analysis. It asks how social power is expressed through demographic and interactional features across cultures, and whether language models can identify those dynamics.

The initial corpus contained 15,836 annotated instances from 100 scenes in French and Egyptian Arabic movies. The French subset had 38 scenes, 824 utterances and 7,998 annotated items. The Arabic subset had 62 scenes, 1,027 utterances and 7,838 annotated items. The total is a count of feature annotations attached to scenes, not a count of separate conversations.

Two independent native-speaker annotators labeled all the movies. For categorical and ordinal features, the study used Gwet's AC2, an agreement measure for categories or ordered ratings. For multi-label features, it used mean Jaccard similarity, which compares the overlap between two sets of labels.

Where social meaning becomes harder

Human agreement was strongest for observable attributes and weaker for relational or perspective-dependent ones. In the Arabic-Egyptian subset, agreement was 0.97 for relationship category, 0.93 for socio-economic class, 0.90 for social class and 0.83 for familiarity. Agreement was lower for power difference, social-status asymmetry and intention alignment.

The same divide appeared in the model comparison. Models were closer to human agreement on demographic and contextual attributes than on emotions, speaker dynamics and socially interpretive variables. In practice, the systems had more trouble with social meaning than with information stated more plainly.

Power labels themselves were often contested. The most disputed was Referent/Charismatic power, the study's term for a form of personal influence that can be confused with other kinds of authority. Disagreement about whether it was present occurred in 23% of such cases, compared with 9% for cases split between an unidentified source of power and Legitimate authority. Disagreements involving two explicit labels made up 13% of cases; Expert-versus-Referent/Charismatic and Legitimate-versus-Referent patterns each accounted for 4%.

Models showed a related problem in distinguishing authority. Across systems, Coercive and Reward-based power were most often folded into Legitimate/Legal authority. Referent/Charismatic and Expert power were less stable, while open-weight or multimodal systems more often produced NA labels.

The visual test needs caution

GPT-5.1 and Gemini-3.1-Pro covered 100% of the items, versus 92.4% for Gemma-3-27B and 89.9% for Qwen-2.5-14B. Both multimodal models covered 55.7%. Their agreement scores were conditional on successful annotations, so they do not represent every attempted case.

The paper notes that safety-aligned refusals and structural output failures constrained the multimodal evaluation, reducing effective coverage and complicating interpretation of the scores. For that reason, the study does not establish a general benefit from visual input. It also covers only two languages, limiting any broad claim about cultural settings; the authors explicitly do not say that particular power types are associated with either culture.

The study also counts a narrow pool of human judgments: two independent native-speaker annotators handled all the movies, and agreement was lower on power difference, social-status asymmetry and intention alignment. The authors say the schema, prompts and annotated data are released so future systems can be evaluated under identical conditions.

Publication and disclosure

The document is an arXiv preprint dated 28 Aug 2026. Its front matter lists Mohamed bin Zayed University of Artificial Intelligence as the authors' affiliation, while the supplied text reports no funding source or conflict-of-interest statement.

Paper data and sources

Original title: The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues
Authors: Farah Atif, Sougata Saha, Monojit Choudhury
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.