Three open-weight language models did not show equal playing strength when the same games were presented through different language interfaces, an arXiv preprint reports. Across the models, English was consistently a strong interface, Hebrew was weak, and Qwen3-4B showed the sharpest language hierarchy.
The test kept the game itself constant. Two instances of the same model played translated versions of one environment while board states, cards, numerical information, available actions and rules remained fixed. The study asked whether equal interactive skill would produce equal playing strength and randomly distributed wins and losses across languages, with knowledge access and general benchmark performance considered separately.
The primary evaluation covered three open-weight models, eight languages and six games spanning spatial reasoning, imperfect information, resource allocation and repeated interaction. Across 18 model–game runs, it recorded 518,400 self-play games; each run contained 28,800 games across all eight language interfaces. The experiments used inference only, with no model training or hyperparameter search.
To compare languages, the researchers used a role-pooled win–loss margin, combining results according to the role each language occupied. Same-language diagonal games served as a comparison, with sides assigned at random.
The size of the gap depended on the game
The size of the difference depended on the game. Colonel Blotto showed the largest language gap across all three models, while Kuhn Poker was consistently among the least sensitive.
The researchers then changed only the language used for the model’s intermediate reasoning, leaving the environment interface unchanged. In TicTacToe, the reported margin reached +0.20, with 89.4% of the gap recovered. In SimpleTak, it improved from -0.21 to +0.05, corresponding to 60.5% recovery. Kuhn Poker showed limited and non-monotonic recovery, so the pattern was not uniform across games.
The differences showed up in decisions
The supporting analyses found that the differences could also appear in the type of defeat. In Gemma’s TicTacToe results, English defeats were distributed across rows, columns and diagonals at 28.5%, 34.8% and 36.7%, respectively. Under Arabic, the corresponding shares were 20.3%, 34.4% and 45.3%, shifting the conditional pattern toward diagonal defeats. These figures describe the mix of defeats, not an overall probability of losing.
Kuhn Poker produced model- and language-dependent strategy profiles. Gemma was most stable in strategically clear states, Qwen was more aggressive and more language-sensitive, and Ministral was less reliable. For Ministral, the reported BadFoldK rate ranged from 10.4% to 23.7%.
A separate Nim analysis highlighted a gap between stating an optimal strategy and carrying it out. For Qwen3-4B, the model mentioned the optimal strategy 150,005 times in English and 159,675 times in French, but made the optimal first move in 80.8% and 24.6% of cases, respectively. The study treats strategy mentions as a proxy for accessible knowledge, not a direct measure of internal knowledge.
Useful clues, but a narrow test
Language-level playing margins were positively associated with scores on two static multilingual benchmarks. The correlations ranged from 0.73 to 0.92 for Global MMLU and from 0.71 to 0.79 for Belebele. A FineWeb-2 estimate of web-text availability showed a positive, strong within-model relationship with mean language margin, with an average correlation of 0.79. These are associations, not evidence that benchmark scores or web-text availability caused the game results; the web-text figure is an estimate, not the models’ actual training distribution.
That evidence remains narrow. It covers three open-weight models, eight languages and six games, with supplementary comparisons against the two static benchmarks and a web-text proxy. It does not establish broader multilingual performance beyond the tested setup.
The study’s practical message is that multilingual evaluation should examine whether behavior stays consistent across language interfaces, rather than rely only on static accuracy. The document is an arXiv version-1 preprint dated 26 August 2026 and states that the relevant code and data resources are publicly available through TextArena.
Paper data and sources
Original title: Skill Issue: Are Skills Language-Invariant in LLMs?
Authors: Bobby Cheng, Adam Gaber, Zhengyuan Liu et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text