Varying the opening words of an AI model's refusals was associated with a refusal signal spread across more internal directions and with smaller changes after researchers removed one estimated direction. That pattern appeared in both frozen-model tests and a controlled fine-tuning exercise in an analytical and empirical case study of refusal-ablation robustness in OLMo-2-0425-1B-Instruct. The finding is an association, not evidence that varied openings caused the model to become harder to attack.
To make the geometry readable, the study used stable rank, a measure that indicates whether a matrix's signal is concentrated in a few dominant directions or distributed across more of them. It tracked refusal and harmfulness rates, defining refusal delta as the change in refusal rate between the pre-intervention and post-intervention outputs. The attack estimated a refusal vector from the difference between harmful and harmless activation means, then removed projections onto that vector.
How the test was arranged
The primary model was OLMo-2-0425-1B-Instruct, an instruction-tuned OLMo 2 1B model with 16 transformer blocks, 16 attention heads and a hidden dimension of 2048. The checkpoint analysis used 60 AdvBench prompts. In the frozen diversity experiment, the study held 80 harmful WildJailbreak prompts fixed and paired them with refusal phrases drawn from one, two, four, six, eight, 10, 12, 14 or 16 distinct first-token buckets.
That design changed the variety of how a refusal began while keeping the prompts fixed. In a separate controlled fine-tuning analysis, the prompt set was fixed and the study used 40 new-topic CAMEL chemistry prompts. It varied balanced refusal completions across one, four, eight, 12 or 16 first-token buckets, then selected checkpoints reaching or exceeding a refusal rate of 0.6.
The internal geometry shifted with diversity
Before comparing attack outcomes, the study compared refusal residuals with gradient-induced activation updates. Those updates were defined as the negative of the realized activation change after a small gradient step. Across the three geometry diagnostics, the results showed signed mean-direction alignment, partial and layer-dependent overlap between principal subspaces, and closely matched stable ranks between refusal residuals and gradient-induced updates.
At layer 7, the leading right-singular axis of the gradient-update matrix had a mean absolute cosine similarity of 0.56 with the refusal direction and 0.96 with the mean update vector. In practical terms, the strongest update direction was moderately aligned with refusal but almost aligned with the average update.
Greater refusal-start diversity was associated with higher stable ranks in both target space and the raw gradients. A layer-15-to-layer-7 transfer factor averaged 2.270, with a sample standard deviation of 0.175 and a range from 1.925 to 2.469. The end-to-end ratio from the refusal residuals to the layer-7 gradients averaged 0.942, with a standard deviation of 0.486 and a range from 0.504 to 1.942, making it the more variable comparison.
For the most diverse frozen condition, with 16 refusal starts, a rank-20 shared map from the layer-15 gradients to the layer-7 gradients captured 99.8% of source-gradient energy across 80 examples. Yet its relative Frobenius error was 0.594, its relative spectral error was 0.316, its mean row-wise cosine similarity was 0.806, and its transfer-factor fit error was 0.343. The fit was measured in-sample, so the energy figure was not an independent predictive validation.
The attack became less effective as rank rose
In fixed-model subset analyses, higher refusal-residual stable rank was associated with smaller harmfulness deltas and slightly smaller refusal deltas under the difference-in-means single-vector ablation. The attack was adapted to the target set, and no formal inferential test was reported. The pattern therefore describes how this estimator behaved on the tested prompts; it does not establish that stable rank caused the smaller changes.
The controlled fine-tuning result pointed in the same direction: more diverse refusal starts were associated with higher refusal-residual stable rank and smaller refusal deltas under ablation. Because the exercise used synthetic fine-tuning on non-harmful chemistry prompts, the study did not test whether the relationship generalizes more broadly.
Released-checkpoint comparisons added a different piece of evidence. The base checkpoint was not attack-effective under this estimator, a functional refusal direction was present by supervised fine-tuning, and later checkpoints after SFT showed larger refusal deltas than the SFT checkpoint. Those comparisons are descriptive because the checkpoints differed in objectives, data, optimization histories and update counts.
A narrow result, for now
The evidence has a narrow reach. It is primarily a case study of one small OLMo model family using relatively small English-language prompt sets, and the white-box attack estimated and applied its direction on the same prompt set. That means the study does not establish that an estimated refusal direction from one set will transfer to unseen prompts, or that first-token diversity causally improves robustness.
Within those boundaries, the findings support further testing of refusal-start diversity as a training-data design variable, while the study's own scope limits any claim of a general defense.
The supplied front matter identifies the work as arXiv version 1 dated 26 Aug 2026.
Paper data and sources
Original title: Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks
Authors: Andrey Labunets
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text