Preprint

New test targets false alarms in multi-arm clinical trials

Preprint proposes adjusted statistics for trials that balance covariates during randomization but may omit them from the final analysis.

A mismatch at the heart of the test

An arXiv preprint proposes a way to test treatment differences in multi-arm trials when randomization balances patients according to covariate profiles but the analysis model leaves some of those covariates out. It works within generalized linear models, or GLMs—models that connect treatment and other covariates with an outcome—and focuses on keeping treatment-effect inference valid under that mismatch.

The problem appears in the standard Wald test, a familiar model-based test statistic. When analysis covariates are omitted, the statistic need not have the standard multivariate normal behavior assumed for the test. Depending on the GLM and the amount of randomization imbalance, the resulting Type I error rate—the chance of a false positive when no treatment difference exists—can be too low or too high.

The proposed adjusted statistics are built to target a standard multivariate normal null distribution even when covariates have been omitted. For several treatment-versus-control comparisons in a multi-arm trial, the procedure uses the Simes test for the global null and Hochberg's step-up procedure for strong family-wise error control when valid independent or positively dependent p-values are available. Family-wise error tracks whether at least one of the parallel comparisons produces a false alarm.

The model changes the risk

The theory predicts different behavior across models. With bounded within-stratum imbalance, standard Wald testing is conservative for logistic regression, valid for canonical Poisson regression and inflated for the exponential model. Here, conservative means that the false-positive rate is lower than intended, while inflated means that it is higher.

The proposed procedure is therefore aimed at calibrating the test's null distribution, not at claiming better power in every setting. Its validity remains tied to the working model, the randomization design, the imbalance conditions and the variance assumptions used by the framework.

Simulations put numbers on the gap

The simulations examined three-arm trials with total sample sizes of 300 and 600 for Type I error, and 300 for power. Each scenario was repeated 5,000 times, using logistic, Poisson and exponential GLMs.

In the N = 300 logistic setting, adjusted family-wise Type I error was 0.0480 under STR-PB, 0.0433 under PS and 0.0534 under HuHuCAR. The corresponding standard Wald figures were 0.0190, 0.0140 and 0.0230. The adjusted values were closer to the intended level, while the Wald test was markedly conservative in that scenario.

Poisson results were closer between the two approaches. At N = 300, adjusted family-wise Type I error was 0.0555, 0.0521 and 0.0530 across STR-PB, PS and HuHuCAR, compared with 0.0525, 0.0570 and 0.0442 for the Wald test. The reported values showed broadly comparable error control in this model.

The sharpest contrast came with the exponential model. At N = 300, adjusted error rates were 0.0538, 0.0522 and 0.0568 for STR-PB, PS and HuHuCAR, whereas the Wald rates were 0.1554, 0.1655 and 0.1794. In the reported simulation, the adjustment substantially reduced the inflation seen with the standard test.

Power gains came with a qualification

Power—the proportion of simulated trials that detected a specified treatment effect—favored the adjustment in one reported logistic scenario. With δ2 = 0.15 and δ1 = 1.0, Treatment 1 power was 0.8148 for the adjusted test, versus 0.6404 for the Wald test.

The same analysis notes that the adjusted procedure may lose power when treatment effects are moderate or strong, so the reported advantage is not universal.

The cancer example used generated data

The paper also used nextMONARCH as an application. The source was an open-label randomized controlled phase 2 study with 234 patients allocated 1:1:1 to A+T, A-150 and A-200; the groups contained 78, 79 and 77 patients, respectively.

The application was not a reanalysis of the original patient-level records. Because the original odds ratios for the stratum factors were not provided, the redesigned exercise generated 1,000 individual-level records from reported objective response-rate values and specified stratum factors.

Within that generated dataset, adjusted power for A-150 was 0.7652 under STR-PB, 0.7657 under PS and 0.7660 under HuHuCAR. Wald power was 0.5438, 0.5492 and 0.5590 under the same schemes. Because the dataset was generated, these are method-comparison results rather than new patient-level treatment evidence.

Where the method stops

The framework does not directly cover continuous stratification variables, depends on the specified GLM and intended link function, and may reduce power for moderate or strong treatment effects.

The supplied document is an arXiv version 1 preprint dated 26 Aug 2026. Its numerical comparisons come from asymptotic theory, simulations and a generated-data application, so they describe the behavior of a statistical procedure under stated assumptions rather than clinical efficacy.

Paper data and sources

Original title: Valid test for multi-arm trials with generalized linear models under covariate-adaptive randomization
Authors: Guannan Zhai, Feifang Hu
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.