Preprint

DeMMO leads wearable-data benchmark across three patient groups

This arXiv preprint reports the strongest held-out performance across four clinical outcomes, while selected mobility measures remain candidates for clinical validation.

An arXiv preprint reports that DeMMO, a machine-learning model using repeated wearable mobility measurements, performed best in held-out tests across Parkinson's disease, multiple sclerosis and PFF. It also had the best mean result for each of four clinical outcomes in a comparison with nine other methods. The finding is a benchmark result, not clinical validation: the mobility measures selected by the model remain candidates for later clinical validation rather than established biomarkers.

The data behind the comparison

The analysis used the Mobilise-D Clinical Validation Study, described as a 2,400-participant resource spanning Parkinson's disease, multiple sclerosis, chronic obstructive pulmonary disease and PFF. The resource involved monitoring on five occasions over two years. For this analysis, researchers used 24 weekly digital mobility outcome measures at five visits and predicted four clinical targets: two Parkinson's disease measures, one multiple sclerosis measure and a PFF impairment measure. COPD was excluded because the required outcomes were unavailable at the second and fourth visits.

After analytic filtering, the analysis included 574 Parkinson's disease participants for each of the two Parkinson's outcomes, 578 multiple sclerosis participants and 469 PFF participants. The retained participant-visit records numbered 2,076 for each Parkinson's outcome, 2,083 for multiple sclerosis and 1,312 for PFF. A record was kept only when it contained the clinical outcome and all 24 mobility measures. Missing values were not imputed, and participants did not need to contribute data at all five visits.

How the model was tested

DeMMO combines regression across repeated visits with temporal smoothness, grouped selection of mobility measures and automatic learning of relationships among outcomes. It retains separate coefficient matrices for each objective and can share information without requiring paired participants across disease groups. The model was compared with nine baselines spanning linear, longitudinal and deep-regression methods. Tuning settings were selected on validation data, and the held-out test set was reserved for final evaluation.

Participants, rather than individual visits, were assigned to the splits, so all visits from one participant stayed together. The allocation was 70% for training, 10% for validation and 20% for testing, repeated over five matched splits using seeds 42 through 46. Across those splits, DeMMO's overall normalized mean squared error, or nMSE, was 0.722, with a standard deviation of 0.032. Its weighted correlation, or wR, was 0.515, with a standard deviation of 0.030. Lower nMSE is better, while higher wR indicates closer agreement between predictions and observed values. The reported one-sided paired-test p-values against the best baseline were 0.0015 for nMSE and 0.0148 for wR.

The advantage held when the outcomes were considered separately. DeMMO had the best mean nMSE and wR for all four outcomes. The paper reports significant gains on both reported metrics for the Parkinson's disease outcome labeled H&Y, and on wR for the multiple sclerosis outcome labeled EDSS. But performance was not uniformly first at every visit: across 20 outcome-visit comparisons, DeMMO ranked first or second in 18 and was uniquely best in 12. No visit-level significance test was displayed.

Patterns the model found

DeMMO also learned signed relations among the four objectives. The strongest positive relation was between the two Parkinson's disease outcomes, with a weight of 0.49. Other positive weights were 0.47 between the multiple sclerosis and PFF outcomes, 0.36 between the Parkinson's MDS-UPDRS III and PFF outcomes, and 0.30 between the Parkinson's H&Y and multiple sclerosis outcomes. Negative weights were -0.32 between Parkinson's MDS-UPDRS III and multiple sclerosis and -0.27 between Parkinson's H&Y and PFF.

These signs describe whether the model's learned mobility mappings point in aligned or opposing directions. They are not direct associations between diseases or raw clinical scores, and the paper reports no uncertainty interval for individual relation weights.

A separate stability analysis used ten hyperparameter settings, with ten half-samples for each setting, producing 100 fits. It found mobility patterns that were generally stable across visits but differed by outcome. No single DMO had a mean selection probability above 0.4 across all outcomes. The pattern argues against a universal panel and leaves the selected measures as candidates for clinical validation.

What the result does not establish

The results show how DeMMO performed in this held-out comparison and which mobility patterns it selected in the included data. They do not amount to clinical validation of those measures. The evaluation used five matched participant-level splits, while the complete-case rule required an outcome and all 24 DMOs at a retained visit and did not use imputation. Participants also did not need all five visits. For now, the selected DMOs remain candidates for clinical validation.

The document is an arXiv version 1 preprint dated 25 August 2026.

Paper data and sources

Original title: DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning
Authors: Menghui Zhou, Zhipeng Yuan, Vitaveska Lanfranchi, Po Yang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.