Preprint

Overnight Bed Sensors Offer Modest Clues to Next-Day Agitation

A preprint study found that a full-night model using minute-level overnight signals showed modest ability to rank next-day agitation in a 65-patient hospital cohort.

An arXiv preprint suggests that sensors placed under a mattress may offer modest clues about whether a person with dementia will show recorded agitation the next day. The best-performing approach kept the overnight signal in minute-by-minute form and reached an AUROC of 0.692. AUROC is a measure of how well a model separates nights followed by agitation from nights without it. Its balanced accuracy was 0.658, so the result points to a usable signal in this dataset but not a dependable individual forecast.

A broad test in a hospital ward

The analysis covered 423 patient-nights from 65 patients in a specialized hospital dementia unit. Here, a patient-night means one person's overnight sensor record paired with the following day's agitation assessment. The recordings came from two contactless under-mattress systems, EMFIT and WSA. Ethics approval was recorded under reference S62882, and participants or legally authorized representatives gave consent as appropriate.

The outcome was broad by design. A night counted as positive when any domain on the Pittsburgh Agitation Scale scored above zero. That yielded 306 positive nights, or 72.3 percent of the 423 observations. The test therefore concerned any recorded agitation under that rule, rather than a separate measure limited to clinically significant agitation.

Keeping the night intact

To ask whether the pattern across the night mattered, researchers compared four representations of the sensor data. B0 reduced each night to conventional handcrafted summaries. B1 retained summaries for three time periods. D0 used full-night sequence modeling, preserving minute-level temporal structure, while M0 analyzed sliding windows through multiple-instance learning. The benchmark used source-specific preprocessing and pooled out-of-fold predictions, allowing the models to be compared on observations outside the fit used to produce them.

D0, the full-night sequence model, led on the headline discrimination measure, with an AUROC of 0.692 and a patient-level bootstrap interval of 0.622 to 0.759. At the fixed predicted-probability threshold of 0.5, its balanced accuracy was 0.658, with an interval of 0.598 to 0.718. On the area under the precision-recall curve, another measure used in the evaluation, it scored 0.849, with an interval of 0.778 to 0.907. Its specificity, or ability to recognize nights without the recorded outcome, was 0.675. The uncertainty intervals came from 1,000 patient-level bootstrap replicates.

M0, the other temporal approach, was nearly tied with D0 on balanced accuracy, at 0.656 versus 0.658. Its AUROC was 0.679, with an interval of 0.592 to 0.764. The simpler B1 and B0 approaches recorded AUROCs of 0.622 and 0.565. The point estimates favor preserving temporal detail over reducing the night to a single summary, but the study does not establish that D0 or M0 is definitively superior.

A signal that still needs testing

That distinction matters because the study is small at the level that matters most for generalization: 65 independent patients, not 423 independent people. It came from one specialized hospital dementia unit and used two sensor sources, so performance may vary with a different hospital, device, acquisition period or labeling protocol. The analysis did not assess inter-rater reliability, leaving open whether differences among staff affected the agitation labels. Representative model configurations and the M0 windowing setup were also selected using the same grouped folds used for evaluation, rather than a separate selection cohort or nested cross-validation. That can produce optimistic future performance estimates, and the bootstrap intervals do not include the resulting selection uncertainty.

Calibration was another unresolved issue. Reliability diagrams were used to examine how closely predicted probabilities matched the observed outcomes, but calibration remained limited. In practical terms, a model may help rank higher-risk nights without providing a reliable absolute percentage for one person. The findings therefore do not support treatment decisions or automated alerts for individual patients, and they do not demonstrate patient-specific day-to-day forecasting. Nor can they be assumed to transfer to other hospitals, devices or community populations.

The work remains an arXiv preprint dated 28 August 2026. Before individual-care use, the authors call for prospective calibration and external validation across hospitals and devices. They also treat the overnight activity, heart-rate and respiratory-rate patterns highlighted by the models as hypotheses for physiological validation, rather than established mechanisms. The primary dataset is not public because of privacy and ethical restrictions. De-identified data were available from the corresponding author on reasonable request, and experiment code would be made available on the same basis.

Paper data and sources

Original title: Under-Mattress Temporal Sensing for Next-Day Agitation Risk Scoring in Dementia Wards
Authors: Zhen Liu, Marta Bono, Robbe Decloedt et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.