Preprint

Preprint tests a cross-checking score for ranking possible drivers

In computer-based tests, the method put known links near the top, but it cannot show that changing a driver would change an outcome.

A new preprint evaluates a computer-based system that combines evidence from 11 analytical methods to rank candidate driver–outcome pairs. In its main synthetic test, the five highest-ranked pairs all matched links built into the data, while 96% of the top 10 did so when results were averaged over 20 random seeds.

That result is about ranking, not intervention. The authors say the system’s Convergent Evidence Score does not establish that changing a driver would change an outcome.

A score built from different kinds of evidence

The system, called Multi-Method Causal Evidence Synthesis, or MCES, runs 11 analytical methods spanning eight mathematical traditions. They include association tests, regularised regression, distance-based analysis, mixed-effects models, machine learning, time-series analysis, information-flow measures, Bayesian-network structure learning and causal forests. The outputs are put on a common 0-to-1 scale and pooled into one score, with equal weights used by default.

The study asks whether this combined score can rank candidate drivers by convergent evidence for relevance to an outcome. Where the underlying structure was known, the researchers compared links built into the test data with pairs that were not part of that structure.

The main synthetic panel contained 23 units observed over 20 periods, with 95 candidate drivers, six outcomes and 18 true links. The evaluation also used real flow-cytometry measurements from 11 phosphoproteins across 853 observational cells, plus six standard Bayesian-network benchmarks based on 1,000 sampled observations.

Strong rankings, but no universal winner

On the primary synthetic panel, MCES recorded Precision@5 of 1.0 and Precision@10 of 0.96. In plain terms, its first five picks matched known links in that test, and nearly all of its first 10 did so.

The pooled score was competitive rather than consistently superior. In the reported scenario, its F1 score—a measure that balances correct picks against missed true links—was 0.686 with a standard deviation of 0.035, compared with 0.714 for the best individual method. The difference was −0.029.

A controlled test examined algebraic identity pairs, in which one variable can mirror another through a built-in mathematical relationship. Removing those pairs raised Precision@3 from 0.3333 to 1.0, but Precision@5 stayed at 0.600; no identity pair remained in the top five.

Held-out calibration also improved. The Brier score fell from 0.025 to 0.011, while expected calibration error fell from 0.101 to 0.004. The authors caution that this calibration was specific to the tested scenario and was not shown to transfer to a new domain.

Performance depended on the setup

On the Sachs benchmark, MCES achieved perfect top-five precision and 0.70 precision among its top 10. Because the data were cross-sectional, only seven methods were applicable.

On six standard Bayesian-network benchmarks, top-five precision was 1.0 on five networks. It was 0.6 on Hailfinder, a network with 56 nodes. These results used a driver/outcome partition derived from the known reference structure.

That partition is a key condition of the approach. When the researchers removed the declared driver and outcome roles in an orientation audit, top-five precision fell to between 0.2 and 0.6, and the ensemble F1 score fell to between 0.25 and 0.35. The test indicated that MCES does not independently recover which way a link runs.

A shortlist, not a causal answer

The reported false-positive figures are empirical results from the evaluated scenarios, not a formal guarantee. Benjamini–Hochberg adjustment reduced the mean per-method false-positive rate from 0.277 to 0.239, while the average rate of null pairs reaching the study’s Moderate-or-higher convergence threshold at CES ≥ 0.4 was 0.000, with a standard deviation of 0.001.

The authors frame MCES as a way to prioritise hypotheses when the right analytical method is uncertain. They say it should not replace an experiment or a well-specified causal design, and should not be used alone to decide that an intervention will change an outcome.

The work is identified as arXiv:2608.20187v1, dated 20 August 2026. Further validation would be needed before treating its convergence thresholds or scenario-specific calibration as portable to new settings.

Paper data and sources

Original title: Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational Data
Authors: Manish Gupta, Dipanjan De
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.