Preprint

RF Simulation Finds Learned Labels Hold Together in Heavy Noise

Preprint: A controlled 3-D simulation found semantic labels remained coherent deeper into noise, while fading-input tests mislabelled 18% of held-out novel objects as known.

A new arXiv preprint reports that, in controlled radio-frequency (RF) imaging simulations, learned outputs kept object categories coherent at signal levels where naive intensity reconstructions were already noise-like. Semantic labels stayed coherent to about −40 dB, while known-object recall remained near-perfect to about −50 dB for models using fading inputs. The work uses a controlled simulation rather than measured radio data, and the document is an arXiv preprint dated 24 August 2026.

Inside the simulation

The simulated setup used three UPA terminals outside the field of view and six transmitter-receiver pairs. For each imaging method, a separate 3-D U-Net turned the six per-pair reconstructions into class probabilities for each voxel, a small cube in the reconstructed scene. The most-probable voxel labels were then clustered and fitted with oriented boxes. That let the study assess semantic reconstruction—what an object was—as well as 3-D detection—where it was.

The comparison operated in a deliberately under-determined regime: each view had 16,384 measurements for 65,536 unknowns, a fourfold imbalance. The deterministic setting used one snapshot, while the fading setting used eight. BP and LASSO were paired with the single-snapshot case, and incoherent BP and group-LASSO with the multiple-snapshot case; naive intensity fusion was included as a reference.

The experiment used 1,000 training scenes, 200 validation scenes and 200 test scenes. Each scene contained six to nine non-overlapping objects, and each was reconstructed at 16 levels, including clean conditions and values from −70 to 0 dB in 5-dB steps. The study used analytic primitives, additive white Gaussian noise and idealized fading instead of measured channels.

The classical trade-off

On naive intensity fusion, sparse solvers were sharpest when conditions were clean. The study’s surface-distance measure, Esdf, was about 0.43 metres for LASSO and 0.50 metres for group-LASSO. Below roughly −40 dB, the sparse methods deteriorated steeply, reaching 1.3 to 1.4 metres by −60 dB. Deterministic BP changed more gradually, from 0.59 to 0.97 metres, and overtook the sparse methods at low signal levels.

The semantic results followed the broad robustness ordering: group-LASSO had the highest clean mean intersection-over-union (mIoU), an overlap score for predicted and target regions, at 0.53. BP with fading inputs degraded most gently, and semantic labels stayed coherent to about −40 dB—after classical intensity reconstruction had dissolved into noise in the comparison.

Labels, boxes and unfamiliar objects

Detection results also favored the fading-input setting at low signal levels. Fading-input models found known objects almost perfectly in clean scenes; group-LASSO was best with about 0.3 false positives per scene. Known-object recall stayed near-perfect to about −50 dB, while deterministic inputs were weaker. Because the boxes came from clustering the model’s most-probable voxel labels, this result covers object-level detection as well as per-voxel labeling.

The open-set test examined what happened when the object classes in testing were unfamiliar. During training, outlier exposure mapped nine exposed-novel classes to one unknown label, while the test pool contained five disjoint held-out classes. With fading inputs, 18% of held-out novel objects were mislabeled as known, compared with roughly one-third to one-half for deterministic inputs. Object-level open-set AUROC was described as high for fading inputs, although the supplied results did not give a numeric AUROC.

Without an explicit unknown class, 68% of held-out-class voxels were assigned to known classes, and post-hoc rejection produced an AUROC of only 0.66 to 0.69. The study presents the explicit-unknown and closed-set results as a comparison of model setups, not a causal estimate of what the unknown label alone changed.

What the evidence does not establish

The simulation used analytic primitives, a Born y=Ac model, additive white Gaussian noise and idealized fading; it included no material scattering or measured data. Each input had one training run. Although the ordering was consistent across all five reported metrics, no error bars were provided to quantify variation.

Within this controlled simulated setting, the results support a narrower conclusion: learned semantic fusion remained usable farther into noise than naive intensity fusion, and fading-input models were the strongest setting for known-object detection and held-out-object rejection. Whether that pattern transfers to measured RF channels or real environments remains unestablished.

Paper data and sources

Original title: Semantic Reconstruction and 3-D Detection via Learned Multi-Pair Fusion in RF Imaging
Authors: Amir Rezaei, Wen-Xin Pan, Giuseppe Caire
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.