Preprint

Preprint: Controllers may need to trade short-term performance for learning

A review of four control strategies says exploration is essential when the parameters that matter most cannot be identified without probing.

A preprint review says controllers that operate while learning an unknown system face a built-in trade-off: they may need to probe the system for information even when that reduces short-term performance. It focuses on linear time-invariant systems—models whose governing relationships are treated as fixed—and on dual control, where actions both control the system and help reveal its unknown parameters.

The document is an arXiv v1 preprint dated 20 August 2026 and says it is to be published in Annual Review of Control, Robotics, and Autonomous Systems in 2027. It organizes the field around four directions: multi-armed bandits, self-tuning regulators, regret-rate minimization for linear-quadratic control, and minimax optimal dual control.

When learning is part of the job

A key point in the review is that the cost of learning depends on identifiability—whether the controller can distinguish among the parameters that describe the system. In linear-quadratic control, logarithmic regret is possible when parameters that cannot be identified are unimportant. But if non-identifiable parameters are essential under the optimal controller, exploration away from that optimum is required, and regret is higher.

In the bandit literature it surveys, the Bayesian formulation is presented as solved by the Gittins index, while an optimism-based policy is reported to achieve the optimal asymptotic regret rate in the stated frequentist setting. The review presents both as formal solutions to the problem of balancing information gathering with performance.

Four routes, different bets

Self-tuning regulators update parameter estimates recursively and use minimum-variance control. The review treats them as benchmarks, not as methods explicitly optimized for exploration–exploitation.

A certainty-equivalence policy acts on the assumption that its current parameter estimates are correct, while adding probes to gather information. The reviewed policy is reported to match the lower-bound asymptotic growth for general linear-quadratic problems, but its probing specification is not optimal for simpler self-tuning problems.

A worst-case approach

Minimax dual control takes a worst-case view, choosing a causal policy to minimize worst-case quadratic cost over unknown model parameters and external disturbances. Under the stated boundedness condition, the review reports that the limit of its value-iteration calculation—a repeated dynamic-programming update—equals the optimal value of the original problem.

The review also identifies a limit in a scalar case where the sign of a state-matrix parameter is uncertain. It reports a lower bound on ℓ2-gain, a measure of how strongly disturbances can be amplified into the state, and concludes that no discrete-time stabilizing adaptive controller is universal across all values of that scalar parameter.

Two model-based tests

The numerical section compares minimax, regret-rate certainty-equivalence and weighted recursive-least-squares certainty-equivalence controllers in two illustrative scenarios: a finite impulse response (FIR) system with white disturbances and an ARMAX system with colored disturbances.

In the Gaussian-white-noise simulation, all compared controllers were stabilizing despite non-minimum-phase behavior. Minimax explored only initially. The other controllers explored longer, and their parameter estimates converged faster, but at a short-term control-performance cost.

In the colored-noise simulation, all compared controllers were stabilizing. Minimax drove the state close to zero faster, while its parameter estimates did not converge exactly.

A map, not a universal verdict

The paper cautions that these numerical cases are illustrative and that more extensive analysis and simulations are needed before broad conclusions about applicability can be drawn.

Paper data and sources

Original title: Dual Control: On Exploration-Exploitation in Linear Systems
Authors: Tomas J. Meijer, Anders Rantzer
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.