An arXiv preprint comparing three ceiling-mounted sensing approaches—FMCW, IR-UWB and Wi-Fi—found that their ranking depended on what the system was asked to recognize. IR-UWB led when the task involved fine-grained activity recognition across held-out participants. FMCW led when the evaluation held out a bed position to test performance across room layouts. On the coarser sleep-disruption task, all three produced high scores, with only a narrow spread under the layout holdout.
The paper is an algorithmic methods benchmark, not a clinical study. It compared synchronized ceiling-mounted recordings from the three sensing systems, using the same activity labels, a common machine-learning architecture and shared evaluation protocols. The aim was to see how the systems handled both detailed in-bedroom activity recognition and broader sleep-monitoring categories.
The result changed with the test
The first task asked the models to distinguish among 10 activity classes. The researchers reported macro F1, a score that combines precision and recall while accounting for class imbalance, so performance on less common labels is not hidden by the most frequent ones. In leave-one-person-out testing, which held out each participant in turn, IR-UWB reached 89.0% macro F1. FMCW scored 83.4%, and Wi-Fi scored 79.0%.
The ordering changed under leave-one-bed-position-out validation, the stricter test of whether a system can cope with a different layout. FMCW reached 83.8% macro F1, ahead of IR-UWB at 78.5% and Wi-Fi at 68.8%. In other words, the sensing approach that led across participants was not the one that led when the physical arrangement used for testing was less familiar.
The authors frame this pattern as a trade-off between activity discriminability—the ability to separate detailed actions—and environmental robustness, meaning the ability to retain performance when the layout changes. The figures support that interpretation within this benchmark, but they do not show that one individual hardware property caused the differences between systems.
The simpler sleep task narrowed the gap
The second evaluation used four broad sleep-monitoring classes rather than the 10-class activity task. After temporal majority voting, which smooths a sequence of window-level predictions by selecting the label that appears most often within each known activity, LOPO macro F1 reached 98.2% for IR-UWB, 96.1% for Wi-Fi and 95.1% for FMCW.
The scores remained high and closely grouped when the test held out a bed position. Under LOBPO, IR-UWB scored 94.2%, FMCW 93.4% and Wi-Fi 92.6%. The largest difference between the three systems was 1.6 percentage points, far smaller than the spread seen in the fine-grained activity task under the same layout-focused evaluation.
Voting itself improved the reported results for every modality. Relative to the pre-voting scores, the gain was 14.0 percentage points for Wi-Fi, 10.4 for IR-UWB and 8.4 for FMCW. But the analysis used known ground-truth activity boundaries when applying the vote. In a deployment where those boundaries were not already known, the improvement could be smaller.
A controlled comparison with deliberate holdouts
The recordings came from 20 participants: 14 male and six female. They completed prompted activity flows in six room layouts, producing up to 120 person-scenario recordings. All three sensing approaches were used in the synchronized comparison, allowing the evaluation to focus on how their measurements performed under matched activity executions and ceiling-mounted placement.
The researchers selected the model architecture and training settings by grid search using the training data, with the test fold kept unseen. They then fixed the selected architecture across the three modalities. That design was intended to keep the classifier and model-selection process common while preserving the modality-specific measurements being compared.
Three holdout strategies were used. Leave-one-person-out tested a participant absent from training; leave-one-scenario-out changed the held-out scenario; and leave-one-bed-position-out tested a bed position not represented in training. Results were reported as fold means with standard deviations. The analysis did not report confidence intervals or inferential hypothesis tests.
Useful for system design, not clinical claims
The deployment conclusion follows the same task-dependent pattern. The authors identify IR-UWB as the most cost- and power-efficient option and as natively supporting unified sensing and communication. Read alongside the recognition results, that makes IR-UWB the paper’s practical choice when cost, power and integrated connectivity dominate, while FMCW becomes more attractive when robustness to layout changes is the priority.
Those conclusions have a narrow evidence base. The work used one residential bedroom test environment rather than independent homes or clinical sites, and the sample was small and demographically imbalanced. The authors say the participant mix does not support broad conclusions about demographic robustness. The benchmark therefore speaks most directly to controlled system comparison, not to performance across the full range of homes or users.
The activity flows were externally prompted by text-to-speech instructions rather than being fully spontaneous. The authors also note that leave-one-scenario-out testing can be partly optimistic because individual bed or chair positions remain represented in other scenarios. These choices make the held-out bed-position result especially useful for reading the layout question, while leaving broader real-world behavior as an open test.
The comparison also does not isolate the independent contribution of hardware, signal representation or preprocessing. Range resolution, Doppler resolution, antenna diversity and spatial retention were not varied separately. In addition, the centered running-median clutter-suppression step was non-causal and could introduce up to five seconds of latency, which matters when considering real-time operation.
Most importantly, the coarse sleep task should not be read as clinical sleep-interruption detection. No clinical sleep reference standard, patient cohort or health outcome assessment was reported. The study does not establish a diagnosis, improved patient outcomes or generalization to clinical populations; it reports an algorithmic benchmark in a controlled bedroom environment.
The paper states that an open synchronized dataset contains FMCW, IR-UWB and Wi-Fi measurements from 20 participants across six room layouts. That makes the comparison available for further examination, although the unanswered questions remain whether the pattern will hold in independent homes, assisted-living settings, clinical populations, multiple-occupant rooms and behavior without known activity boundaries.
The manuscript is identified as arXiv:2608.20322v1 and dated 20 Aug 2026. It reports funding from the DistriMuSe project, the CORNET project DARTS and the NAV-ALERT project from Belgian Defense.
Paper data and sources
Original title: A comparison between ceiling-mounted FMCW, IR-UWB and Wi-Fi radar for in-bedroom human activity monitoring and sleep interruption detection
Authors: Anton Lambrecht, Reda El Hail, Xianjun Jiao et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text