Preprint

Vision-Language Encoders Shift Which Clues Matter During Search

Preprint: In 800 controlled scenes, frozen OpenCLIP and SigLIP often changed which visual clue was most useful as evidence accumulated.

A new arXiv preprint reports that two frozen vision-language encoders can change which visual clue appears most useful after another clue has been acquired. The preferred next clue also changed with the semantic target being sought, even when the scene itself stayed fixed.

The main held-out confirmation used 800 new scenes, exactly 200 in each of four regimes. It evaluated frozen OpenCLIP and SigLIP without calibration or fine-tuning, comparing oracle-positive regimes designed to induce ordering changes with oracle-negative controls.

Rather than treating evidence as having a single fixed value, the analysis treated the target as one of 25 candidates and applied a 25-way softmax to frozen image-text logits. It then measured conditional marginal utility, meaning how much adding an evidence type changed the target's log-probability score.

When the ranking reversed

At the study's fixed margin threshold of 0.05, a robust reversal counted when one of the three evidence-pair comparisons had a margin above that threshold both before and after the third attribute was used as history. In the oracle-positive regimes, the reversal rate was 84.2% versus 8.2% in oracle-negative regimes for OpenCLIP using additive accumulation, a difference of 76.0 percentage points. With jointly rendered direct views, the rates were 96.5% versus 15.5%, a gap of 81.0 points. SigLIP showed 96.0% versus 22.5% with additive accumulation and 96.0% versus 23.2% with direct views, for gaps of 73.5 and 72.7 points.

The reported 95% confidence intervals for the four positive-minus-negative gaps all stayed above zero: [71.5, 80.2], [77.0, 85.0], [69.0, 78.0], and [68.0, 77.2] percentage points. Those intervals came from 10,000 percentile bootstrap repetitions, with scenes treated as the resampling clusters.

The reversals were not spread evenly across every pair. Comparisons of color with texture after shape ranged from 41.0% to 52.8%, and comparisons of shape with texture after color ranged from 41.5% to 53.2%. Color versus shape after texture was lower, ranging from 5.8% to 17.0%.

The pattern survived several checks

The pattern survived more than one way of representing acquired evidence. Reversal labels, conditional utilities and preferred next actions agreed strongly between additive singleton-logit sums and jointly rendered direct views for both backbones. Direct rendering changed some individual decisions, but it preserved the same regime-level structure, and every frozen wording template in the 800-scene confirmation retained a positive regime contrast.

The target mattered as well. In a separate earlier confirmation using 400 scenes and 20,670 matched observations, query-specific oracle-action agreement ranged from 92.06% to 94.84%, compared with 61.44% for a query-independent baseline. That was an improvement of 30.62 to 33.40 percentage points, and all scene-bootstrap 95% confidence intervals excluded zero. In practical terms within this test, changing the semantic target changed which evidence action was preferred in otherwise fixed scenes.

Under the query-scene derangement control, oracle balanced accuracy was 50.25% and 52.63% for OpenCLIP's additive and direct modes, and 48.88% and 51.0% for SigLIP. The deranged positive-minus-negative contrasts were 0.5, 5.25, -2.25 and 2.0 percentage points, leaving oracle alignment and regime separation near chance-scale levels. The supplied analysis does not report confidence intervals for those deranged contrasts.

A second-step gain after the first choice

Researchers also examined what happened if the first evidence choice was held constant. In this post-confirmation exploratory comparison, the policies matched the first acquisition and differed only at step 2, so the comparison focused on reranking the remaining evidence after the first acquisition.

That reranking produced positive step-2 target-log-probability gains of 0.623 to 0.730 across accumulation modes, 0.534 to 0.586 across wording conditions, and 0.587 to 0.701 across backbones. Action agreement in those comparisons ranged from 88.5% to 94.5%, 88.4% to 92.1%, and 85.4% to 89.9%, respectively, and every individual 95% confidence interval was positive.

That reranking result was post-confirmation and exploratory. The prospectively specified adaptive-versus-best-fixed comparison showed an overall advantage, but its mechanism-relevant interactions did not isolate the value of replanning.

The boundaries of the result

One selectivity control showed that both backbones achieved 100% isolated color accuracy and 100% isolated shape accuracy, while isolated texture accuracy was 74.86%. The strong-leakage flag did not trigger, but that result did not establish that removed attributes decoded at chance.

The paper is an arXiv preprint, version 1 dated 28 Aug 2026. Within the evidence reported here, the result concerns frozen model outputs in a controlled held-out test, while the reranking comparison remains exploratory.

Paper data and sources

Original title: Conditional Visual Evidence Utility: State-Dependent Rank Reversals in Frozen Vision-Language Encoders
Authors: Yunxuan Fang, Xinhe Wang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.