Preprint

Study Finds Sparse Method Can Shrink Multimodal Search Costs

Preprint: PUMA often matched or exceeded dense retrieval across five benchmark datasets while using less storage and speeding exact search on large candidate pools.

At the reported operating points, a post-hoc method kept pace with dense retrieval on most of the five benchmark tasks while using far less storage. On the larger candidate pools, exact sparse scoring was also much faster than exact dense scoring. With Qwen3-VL-Embedding-2B, PUMA’s nDCG@10 score—the study’s top-10 retrieval measure—rose from 0.4235 to 0.4356 on CIRR and from 0.0792 to 0.1144 on Fashion200K; both gains had p ≤ 0.001. It was statistically indistinguishable from dense retrieval on FashionIQ and MSCOCO, but lower on VisualNews.

A retrofit for existing embeddings

PUMA is built as a retrofit for an existing embedder. It applies a post-hoc sparse autoencoder—a model run after the original embedder—to frozen universal multimodal embeddings, producing compact retrieval codes while preserving the dense dot-product geometry, or pattern of similarity scores, during pretraining. The recipe then adds cross-modal alignment and retrieval fine-tuning without retraining the backbone.

The evaluation covered five M-BEIR datasets spanning composed image retrieval and text-to-image retrieval. The study compared PUMA with full dense retrieval and compression controls, including Raw TopK, matched-memory PCA and encoder-only TopK. Dense and sparse variants used identical prompts, cached embeddings and candidate pools; uncertainty was assessed with 95% bootstrap intervals and 1,000-trial paired approximate-randomization tests on nDCG@10.

Controls point to the full recipe

On the Qwen-2B CIRR test, PUMA also outscored the simpler compressors: its nDCG@10 was 0.4356, versus 0.3769 for Raw TopK, 0.4153 for matched-memory PCA and 0.3631 for encoder-only TopK. Across 15 model-dataset settings, a trained dense autoencoder scored below PUMA in 12; it was higher only on MSCOCO for Qwen-2B and Qwen-8B and on RZen-FashionIQ.

Under the paper’s FP32 storage accounting, PUMA used eight times less storage than dense embeddings with Qwen-2B and up to 16 times less with Qwen-8B. Exact sparse search was also much faster on larger candidate pools: VisualNews with Qwen-8B took 423.8 seconds for dense scoring and 17.3 seconds for sparse scoring, or 24.5 times faster. On Fashion200K, the corresponding times were 13.8 seconds and 0.77 seconds, or 17.9 times faster. Dense search remained faster on the smaller MSCOCO pool.

Where sparsity breaks down

The advantage was not universal. On Qwen3-VL-Embedding-8B, PUMA was about 0.027 nDCG@10 higher than dense on VisualNews and Fashion200K, with p ≤ 0.001, stayed close on MSCOCO and FashionIQ, and was below dense on CIRR. On RZenEmbed, used as a cross-family check, it was higher than dense on FashionIQ and Fashion200K but below dense on CIRR, VisualNews and MSCOCO. The Fashion200K difference was statistically reliable at p ≤ 0.001; the FashionIQ result was positive but uncertain at p=0.075.

The paper’s diagnostics found two different ways sparsity can fail. One is insufficient support before TopK, the step that selects a fixed number of features: on Qwen-8B CIRR, there were 166.5 positive pre-TopK activations on average, and the target of 160 features was reached exactly in 81.8% of examples. The other is retrieval-misaligned active support. On RZenEmbed’s VisualNews set, there were 465.3 positive activations on average and k=160 was reached on every example, yet PUMA still remained below dense retrieval.

An ultra-sparse sweep made the quality-efficiency trade-off visible. At k=16, PUMA retained 74% of dense nDCG@10 on CIRR and 69% on Fashion200K. At k=48, it reached 96% of dense nDCG@10 on CIRR and exceeded dense on Fashion200K, with nDCG@10 of 0.110 versus 0.079. These were inference-time support settings applied to existing checkpoints, so the comparison shows how the operating point changes the benchmark result.

The result changes with the backbone

Leave-one-out tests suggested that several parts of the recipe mattered. Every second-stage loss was described as load-bearing; AuxK had the largest observed loss on CIRR, while the second-stage contrastive blend was most important on Fashion200K. Turning the third stage off was associated with a decline of 0.019 in nDCG@10 on CIRR and 0.009 on Fashion200K.

A benchmark result with clear boundaries

The evidence supports a narrower conclusion than a universal performance claim: in the tested configurations, PUMA recorded lower storage use and faster exact search while keeping retrieval results competitive on several tasks. The main experiments focus on the Qwen family, with RZenEmbed serving as a cross-family check. The diagnosed support bottlenecks are not solved, and post-hoc sparsification inherits representational properties, domain imbalance and possible bias from the frozen backbone. Search timings come from the reported hardware and exact-scoring implementation and may not transfer to other deployments.

Paper data and sources

Original title: PUMA: Post-Hoc Sparsification of Universal Multimodal Embeddings for Efficient Retrieval
Authors: Matteo Attimonelli, Alessandro De Bellis, Franco Maria Nardini et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.