A quantum generative model produced synthetic patient data that its authors judged better than three classical generators in a test built around scarce data from myelodysplastic syndrome, or MDS. The quantum system performed better on measures of how closely the synthetic records matched the source distribution, and the results improved as the training set grew.
The study was a seven-variable MDS proof of concept. Its results concern the quality of generated data and model comparisons, rather than clinical outcomes.
A narrow test of synthetic patient data
The input data came from SintraREV, described in the study as a randomized, double-blind phase 3 trial of low-dose lenalidomide versus placebo in non-transfusion-dependent patients with low-risk del(5q) MDS. The researchers arranged the information in two related databases. One held basic clinical and blood-related measurements; the other also included treatment assignment and survival outcome.
The researchers trained and tested the systems with subsets of 100, 150, 200, 300 and 400 patients. The quantum pipeline first estimated the target joint distribution of the variables using maximum entropy. It then encoded that distribution with an ancilla qubit, compressed it into a matrix product state, and mapped it to a deep quantum circuit for sampling on IBM hardware.
The quantum implementation ran on IBM’s 156-qubit ibm basquecountry Heron r2 processor and used error detection followed by post-selection. The comparison models were TVAE, CTGAN and CopulaGAN. The comparison was therefore limited to these baselines and the seven-variable MDS representation used in the proof of concept.
Synthetic records were harder to spot
For generalization, the researchers used mean absolute error and Jensen-Shannon divergence, two measures used to compare generated and reference data. On both measures, the quantum framework was reported to perform better than the classical baselines at every tested training-set size, with its performance improving as more patients were used for training.
A separate test used a Random Forest classifier to distinguish real records from synthetic ones. The balanced evaluation set contained 500 real and 500 synthetic samples, split 70/30 between training and testing, and the classifier used 100 trees. Accuracy close to 50% was treated as the point at which the two types of records were difficult to tell apart.
The quantum-generated records were described as difficult to distinguish from real data and as having the highest or near-highest coverage, a measure of how much of the reference data’s range the synthetic set represented. The reported 2 × 2 confusion matrices were near chance, while the classical models showed non-trivial biases. The authors therefore identified the quantum model as the best of the systems examined for both generalization and expressivity, meaning its ability to reproduce varied patterns in the data.
What the result still leaves open
In the supplied analysis, the comparisons are qualitative: it gives no numerical values for the error measures, classifier accuracy, area under the curve, coverage, confusion-matrix entries or error bars. The size and precision of the reported differences are therefore difficult to judge.
The study also describes a trade-off in the data representation. Finer discretization was reported to improve clinical expressivity, but it required more quantum resources and made the system more sensitive to noise. The use of error detection and post-selection was part of the quantum pipeline tested on hardware.
The reported comparison does not establish performance beyond this MDS representation, the tested baselines or the hardware implementation. The paper identifies its arXiv version 1 as dated 28 August 2026, making it a preprint rather than a reported peer-reviewed result in the supplied publication information.
The work was supported through the Basque Quantum strategy by the QSynthInSilico project. The supporting data and software code are available from the corresponding authors on reasonable request.
Paper data and sources
Original title: A quantum generative model for in silico clinical trials using scarce training datasets
Authors: Olatz Sanz Larrarte, Reza Dastbasteh, Roberto Sanchez-Navarro et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text