An inference-time technique for generative vision-language models recorded the lowest reported average bias-change score in all four systems tested, while MMStar accuracy stayed within 0.6 percentage points of the unsteered baseline. The result is reported in an arXiv preprint evaluating GGSS, a method aimed at reducing measured race- and gender-related bias while keeping general multimodal performance close to baseline.
GGSS works during inference, while a model is processing an input and generating an answer. It has two stages: first, it discovers a spherical counterfactual bias subspace; then it applies gated, token-level geodesic steering and restores the signal's norm. In ordinary terms, it looks for a direction in the model's internal signal associated with the tested bias and selectively adjusts it along a curved path.
What the researchers tested
The evaluation covered four generative VLMs from three model families: Pixtral-12B, LLaVA-1.6-Vicuna-7B, LLaVA-1.6-Mistral-7B and Qwen3-VL-4B-Instruct. The group included a 12-billion-parameter model, two seven-billion-parameter models and a four-billion-parameter model.
To discover the bias subspace, the researchers used 480 real-photograph, face-only counterfactual images. The set covered six occupations, eight source identities, five perceived races and two genders. The comparison also included ten adapted baselines, prompt mitigation and four structural ablations.
For each method-model pair, the researchers tested five steering strengths: alpha values of 0.25, 0.5, 0.75, 1.0 and 1.5. They then selected one under the best-avg-alpha rule. Uncertainty was assessed with 95% bootstrap confidence intervals based on 10,000 resamples, paired sign-flip permutation tests and exact McNemar tests for MMStar.
The largest changes were on bias measures
Across the reported bias probes, GGSS had the most favorable average percentage change on every model. The figures were -55% versus -37% on Pixtral, -90% versus -86% on LLaVA-Vicuna, -80% versus -78% on LLaVA-Mistral, and -60% versus -49% on Qwen3, with the first number in each pair belonging to GGSS and the second to the strongest external result. Because these are percentage changes, the more negative figure represents the larger reported reduction.
Individual probes showed the same direction. In the Nurse/Doctor gender test, the bias score moved from 0.600 to 0.025, a reported 96% reduction. The multiple-choice race-bias measure, reported on a JSD x 10^3 scale, fell from 8.70 to 1.36, or 84%. The 2AFC race-bias score fell from 0.243 to 0.096, or 61%.
Paired tests reported significant reductions on three of the four backbones. The LLaVA-Vicuna Nurse/Doctor comparison had p < 10^-4, indicating a particularly strong statistical difference in that paired test.
Capability stayed close to baseline
On the capability side, MMStar accuracy remained within plus or minus 0.6 percentage points of the unsteered baseline across the reported evaluations. That directly addressed the study's central trade-off: whether measured demographic bias could be reduced while general multimodal performance stayed close to its starting point.
The wider comparison showed that near-baseline capability was not automatic for every correction. Under INLP (sph.), Pixtral's MMStar score changed by -6.1 percentage points, with p < 10^-4. The near-baseline result was therefore specific to the tested GGSS setup rather than a result shared by every method examined.
The component tests pointed to both of GGSS's main pieces. In a representative ablation, the full method's JSD was 2.41, compared with 5.61 when the gate was removed and 4.53 when the gate remained but Slerp was removed. The full combination had the lowest score among those three settings.
Geometry also mattered in the reported Pixtral-12B checks. Spherical steering was better than Euclidean steering under both geometry toggles, with p-values of 0.031 and 0.015. The Euclidean variant had a -6.3 percentage-point MMStar change; in a degenerate over-steering cell, 72 of 80 responses were unparseable and bias overshot the unsteered level by nine times.
A cross-fitted check of the steering strength gave a similar signal: the selected alpha matched the paper's choice on six of eight identity folds, and the held-out reduction was negative in every fold on every model.
The evidence has a clear boundary
The findings are therefore best read within the tested setup. The evaluation covered four backbones from three model families and the specific race, gender and MMStar comparisons described in the study. It supports a conclusion about those model outputs and probes, not a universal verdict on every vision-language model or every form of demographic bias.
The paper is an arXiv preprint, version 1, dated 26 August 2026. It reports releasing the GGSS code along with configuration files, steering checkpoints and aggregate evaluation outputs. Funding information is not reported in the supplied text.
Paper data and sources
Original title: GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models
Authors: Yiqun Sun, Junyu Chen, Pengfei Wei, Lawrence B. Hsieh
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text