Preprint

Preprint: Reasoning model matches engineered controller in simulated plant tests

A bounded system stayed within hard constraints across its primary benchmark runs, but the study does not establish performance on a physical plant.

A high-effort Sonnet-5 reasoning model stayed within all of the study’s hard operating constraints in each of 39 primary benchmark episodes. The baseline controller recorded 15 trips in the same paired set of scenarios.

An engineered reference controller also met the endpoint in all 39 episodes. The model and reference had no discordant primary-endpoint pairs, and the 95% confidence interval for their difference ran from −0.09 to +0.09, meaning the study did not distinguish them on its main measure.

How the test worked

The model ran above decentralized regulatory control through a bounded, programmatically verified interface. It could not directly manipulate actuators, and the setup used no task-specific training or fine-tuning.

The benchmark used 13 scenarios and 3 paired noise seeds per scenario, producing 39 episodes for each configuration. Every configuration received the same measurement-noise realizations, so comparisons were paired run by run.

The scenarios covered three kinds of problem: no intervention, safety gaps and quality gaps. Basic decentralized control stayed within constraints in no-intervention cases, tripped in safety-gap cases and could not track quality-gap changes; the reference controller maintained safety-gap operation.

More than one model cleared the endpoint

The result was not confined to the primary configuration. DeepSeek-V4-Flash remained within constraints in 39 of 39 episodes, low-effort Sonnet-5 in 38, and GLM-5.2 in 37. At this sample size, none was distinguishable from the reference controller. Reported campaign costs ranged from about US$2.80 to US$90.48 for 39 episodes.

In the safety-gap episodes, the model’s diagnosis was stronger than the scripted supervisor’s: it identified the ground-truth tag in all 15 cases, while the rule table identified the cause in 2 of 15. The study did not report a separate uncertainty interval for this comparison.

The quality-gap results also differed according to the information provided. When commanded-target deviation was reported, all 9 quality episodes produced trim activation and the mean grade offset was 0.3 mol%. Without that information, trim activated in 6 of 9 episodes and mean offset was 4.8 mol%.

On 15 paired safety-gap episodes, mean peak pressure was 2,774 kPa for the model and 2,797 kPa for the reference. The model’s reactant feed was about 19% lower; the reported paired mean difference was 1.8 kscmh, with a 95% confidence interval of 1.5 to 2.3.

Where the evidence stops

The record was not perfect across model-supervised runs: 3 of 156 episodes ended in trips, each through a distinct mechanism. The reported failures involved reduced-reasoning diagnostic error, unnecessary intervention, or a revision away from a correct initial diagnosis.

The comparison between reasoning efforts was too small to establish a consistent difference. In one matched scenario-and-noise-seed pair, low-effort Sonnet-5 tripped while high-effort Sonnet-5 stabilized; across 39 pairs, that was one discordant pair and produced an exact McNemar p-value of 1.0.

The no-shadow and no-validator ablation runs each stayed within constraints in 39 of 39 episodes. The shadow veto never fired, while static checks rejected three proposals.

The authors interpret the findings as evidence that general-purpose reasoning models can manage abnormal situations when their authority is bounded and their outputs are checked. But the evaluation was simulation-only, used the Tennessee Eastman benchmark, and did not compare the system with human operators or provide a formal safety guarantee.

The open questions include whether the design transfers to another process, handles abnormalities without pre-engineered responses, or catches failures that emerge only after a sequence of individually admissible actions. The document is an arXiv version 1 preprint dated 20 August 2026.

Paper data and sources

Original title: Large reasoning models for abnormal situation management in safety-critical industrial processes
Authors: Khalid Alhazmi
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.