Preprint

Underwater image scores can rise as 3D surfaces deteriorate

An arXiv preprint comparing five systems found image scores diverged from geometry measures in murky water and on an operational survey.

The central finding was a split between image-quality scores and geometric accuracy. On the SOTRUE turbidity sequences, the vanilla baseline's image score rose in murkier water while its stereo-referenced surface-depth error became far larger.

The arXiv preprint evaluated five methods: vanilla 3DGS, a UIE-to-3DGS baseline, WaterSplatting, SeaSplat and UW-GS. It asked which available underwater Gaussian-splatting methods solve 3D reconstruction and which mainly produce visually plausible renders, testing them across four labeled underwater regimes.

In murkier water, image and geometry scores diverged

SOTRUE supplied six measured turbidity levels along an identical servo-driven trajectory, with motor-encoder poses and turbidity as the only sequence variable. For vanilla 3DGS, PSNR, an image-similarity score, was 32.0 decibels in clear water and 35.9 decibels at 7 NTU. The corresponding stereo-referenced surface-depth errors were 99 millimetres and 848 millimetres. The paper describes the error as levelling off near 850 millimetres.

Classical COLMAP, using SIFT features, registered 99.5 per cent of S2 frames in clear water, 1.0 per cent at 7 NTU and 0.0 per cent at 12 NTU. That result applies only to the classical SIFT-based pipeline tested here; learned feature matchers were not tested.

A geometry-free static-image control also scored higher as turbidity worsened, rising from 24.1 to 32.0 dB. Over the same range, the baseline's margin over that control narrowed from 7.9 to 3.4 dB. In this comparison, PSNR became less discriminating about the reconstruction.

The survey test reversed the picture

On S1, all four in-medium systems were within 1.3 dB of one another on image score, but WaterSplatting, labeled M2, had a floater mass of 0.077. Floater mass was a reference-free endpoint in the evaluation, and the near-tie in image score did not separate the systems on that measure.

On S4, SeaSplat, or M3, led the photometric benchmark at 24.7 dB but was last on the geometry measures, including a 369 mm Chamfer error. The UIE-to-3DGS baseline, WaterSplatting and vanilla 3DGS recorded Chamfer errors of 45, 46 and 58 mm, respectively. The study treated those values as tied within its 12 mm alignment residual. SeaSplat took 353 minutes to train, compared with 28 minutes for its peers.

The measurements varied with light, coverage and transfer

In S3, with an artificial light moving with the camera, vanilla 3DGS had the highest reported appearance score at 27.0 dB, ahead of SeaSplat at 25.9 dB and WaterSplatting at 23.2 dB. Its floater mass was 0.116, versus 0.014 for SeaSplat and 0.008 for WaterSplatting. WaterSplatting used 27 k Gaussians, while the other systems used about 1.9 M.

The view-overlap comparison showed a large gap in the vanilla baseline's score. With every 16th frame retained instead of the full sequence, PSNR was 16.5 dB rather than 27.0 dB, a 10.5 dB difference. The measured view-overlap gap was larger than the gaps between methods in that test.

An S1 SeaSplat ablation compared the full system with versions in which tested components were absent. The row without backscatter recorded 19.8 dB PSNR, versus 30.2 dB for the full version, a 10.4 dB difference. The row without the tested depth prior recorded PSNR 2.8 dB below full, while floater mass was 0.25 instead of 0.37, or 31 per cent less. In the reported comparisons, the tested components showed an appearance-geometry tradeoff.

A cross-site transfer test applied medium networks fitted on S1 to the S3 model. The resulting PSNR was 17.0 dB, 9.0 dB below the 25.9 dB within-site S3 result.

A result tied to the test conditions

Every system used shared per-scene poses and sparse initialization. Each cell used a single run of 30 k optimization iterations on one 24 GB GPU. The evaluation combined held-out-view photometric metrics with stereo- or photogrammetry-based geometry measures and a reference-free floater-mass metric.

Across the four labeled regimes, the measured outcomes varied with turbidity, illumination, view overlap and cross-site medium transfer. The result is a condition-sensitive comparison, not a single method ranking.

The practical message for underwater reconstruction is to assess appearance and geometry together under the conditions of use. The tests showed that close image scores could coexist with different floater mass, near-zero classical registration at high turbidity, or a large surface error alongside a leading photometric score.

This document is an arXiv preprint and reports the release of scene builds, per-run configurations and evaluation code through a GitHub repository. The main reported funding source was Innovation Fund Denmark through DeepODO, a project focused on deep visual odometry for underwater intervention drones.

Paper data and sources

Original title: Gaussian Splatting Underwater: A Controlled Cross-Regime Study
Authors: Olaya Álvarez-Tuñón, Stella Graßhof
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.