An arXiv preprint reports that its model-selection analysis favored a two-component contaminated fit for 6,126 American Time Use Survey responses, a dataset with 49.116% of cells missing. The fit classified 16.389% of rows as outliers.
The same fit produced two model-derived groups containing 32% and 68% of the responses, with 20% to 30% of individuals within each group classified as atypical. The labels were not externally validated.
A model for parts of a whole
The paper proposes finite mixtures of contaminated Dirichlet distributions, a probability model for parts of a whole, to represent heterogeneous compositional data and mild outliers. In the time-use example, daily activity durations add up to 24 hours, or 1,440 minutes.
To deal with unobserved parts, the method assumes missing at random, or MAR, an assumption about how missing values arise, and derives a closed-form conditional distribution for the missing compositional parts.
The observed-data log-likelihood, meaning the fit calculated from values that are present, is formulated as a convex optimization problem. A tailored expectation-maximization, or EM, algorithm incorporates moments for missing components and tractable iterative updates, while selected mixture parameters are estimated through constrained Newton-Raphson steps.
The paper’s theoretical conclusion is that the likelihood estimates exist and are unique under MAR. It also says the model is identifiable up to permutation, meaning component labels can be reordered, when the number of components does not exceed the simplex dimension.
Performance varied with missingness
To test the method, the authors generated 1,000 datasets from a contaminated two-component Dirichlet mixture with missingness at random. The simulated datasets used either 100 or 1,000 observations.
Higher missingness was associated with lower clustering accuracy and adjusted Rand index, a measure of agreement between groupings. Exact metric values were presented in figures rather than reported in the text.
Outlier-detection accuracy stayed high even when 90% of cells were missing, but the true-positive rate fell while the false-positive rate remained low. In the simulations, sparse rows were more likely to be treated as typical than flagged as outliers, and the metrics were better overall with 1,000 observations than with 100.
Parameter recovery was less uniform. Mean estimates remained unbiased as missingness varied, although their root mean squared error, a measure of estimation error, was affected by the amount of missing information. Variability estimates tended to be overestimated with fully observed data, improved at moderate missingness, and became more biased again when missingness was high.
Model-selection criteria did not behave alike. Under the contaminated model, AIC, BIC and ICL were more likely to select the correct component count. More missingness favored fewer components, while AIC selected more components overall than BIC or ICL.
What the evidence leaves open
The simulation evidence comes with limited numerical detail: exact metric values, curves and uncertainty estimates are not reported in the prose, although the results are shown in figures.
The ATUS outlier and cluster labels are model-derived and have no external validation or reported inferential uncertainty. The missing-part distribution and uniqueness conclusion are conditional on MAR.
The document is an arXiv preprint, version 1, dated 24 August 2026.
Paper data and sources
Original title: Handling mild outliers and unobserved values in compositional datasets using finite mixtures of mean-parametrised Dirichlet models
Authors: Jason Pillay, Andriëtte Bekker, Cristina Tortora, Antonio Punzo
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text