An arXiv preprint presents a way to analyze AI systems that make decisions for many people as social-choice mechanisms, and its computational illustrations show a trade-off between protecting people from low modeled welfare and maximizing total welfare. In the Community Alignment analysis, a harm-bounded rule kept welfare above a chosen floor, but the mean-welfare cost grew as that floor was tightened.
The manuscript is a v1 arXiv preprint dated 25 August 2026. Its central mathematical move is to represent welfare through an impact vector. Under linear-utility assumptions, utilitarian alignment becomes a linear optimization over a convex feasible impact set. In the data exercise, the authors fit individual preferences as linear weights on model features.
The map behind the rules
The feasible impact closure has the geometry of an origin-symmetric zonotope, a shape balanced around zero, while attainable impacts make up its relative interior. In practical terms, the framework distinguishes the full set of impacts that can be approached from the interior outcomes that can be attained under the model. That geometry supplies a common space for comparing alternative alignment rules.
One formal result covers voting-by-issues mechanisms, which decide different aspects of an outcome separately. The paper states that every mechanism in this class is strategyproof, meaning an individual cannot improve the result by misreporting preferences, and unanimous, meaning that a shared preference is followed. A random dictatorship can be compiled into a single model parameter for an interior target impact; for other targets, it can only be approached as a limiting sequence. Under the theorem's non-degenerate welfare-weight condition, every random dictatorship has unbounded worst-case distortion, meaning the analysis finds no finite upper bound.
Where the rules can bend
RLHF, or reinforcement learning from human feedback, produces another warning in the framework. When the training and deployment queries have the same distribution, an anonymous RLHF stationary point, a point where the training objective is stationary, is impact-equivalent to randomly labeling each query. The resulting impact depends on the training data only through the frequency of each label on each query.
At finite temperature, however, the mechanism is not strategyproof in the stated setting. For every agent with a positive welfare weight and a generic true preference, reporting that preference at any scale greater than 1 strictly improves on truthful reporting, and no optimal report exists.
Linear welfare constraints do not change the basic optimization form: whenever the constrained set is nonempty, the problem remains a linear optimization over the feasible impact set. In practical terms, the framework can impose a minimum welfare requirement and compare the resulting cost to total welfare. On Community Alignment, utilitarian, RLHF and strategyproof mechanisms showed wide welfare variation, with some agents harmed, while harm-bounded alignment stayed above the selected floor at a growing mean-welfare cost as the floor tightened.
What the illustrations found
To illustrate how the framework behaves on human preference data, the authors fitted participant-level linear preferences with Bradley-Terry-Luce models to pairwise choices. They varied a choice-precision parameter, projected Community Alignment responses onto a 15-dimensional interpretable basis, and computed the impact zonotope for the observed queries exactly.
The examples drew on kidney allocation, 412 Food Rescue, Moral Machine and Community Alignment. Food Rescue included 19 participants with 45 pairwise responses each. The Community Alignment analysis retained 2,387 annotators after filtering for at least 20 responses.
Across the datasets, preference directions were more often aligned in the kidney and food-rescue domains, while they disagreed in Moral Machine and Community Alignment. The strategyproof mechanism lost a fraction of optimal welfare, but at high choice precision it also had the least welfare variation among the mechanisms compared.
A conditional result
These conclusions depend on the framework's design. It assumes preferences are linear in an interpretable model feature space, even though interpreting such features for LLM alignment is difficult. It also relies on query-distribution assumptions that may fail in practice, and its strategyproofness results require individual selection probabilities to be broadcast in advance.
The experiments ran locally. For Community Alignment feature embeddings, the authors rented an A100; the embeddings took approximately five minutes. They also provide a code repository.
Paper data and sources
Original title: Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
Authors: Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text