Preprint

One Synthetic Toxic-Exposure Model Reports Different Policies by Timescale

A 2026 arXiv v1 Preprint reports different learned assignments in one toy toxic-exposure model and compares separate utility perspectives with a combined approach.

A 2026 arXiv v1 preprint dated 26 August reports different assignment policies and outcome distributions in one synthetic toxic-exposure scheduling model. Its numerical example compares SER, ESR and SF, with a further difference when three timescales are considered simultaneously. The result comes from a toy employer scenario.

The paper asks whether one decision problem can contain non-linear utility effects at different timescales and therefore require simultaneous timescale optimization. Here, utility is the scoring rule used to turn what happens in the model into something an algorithm can optimize. The paper distinguishes acute utility, applied at each timestep, from episodic utility, applied to returns or values. For SER, episodic utility was applied to the value; for ESR, it was applied to the return.

How the comparison was framed

The numerical example used three employees and three tasks. Their skill levels were 0.75, 0.60 and 0.2, while the task toxicity values were 70, 25 and 10. Employees initially started at their respective tasks.

Daily rewards were defined as the negative of the employees' toxic doses. Those doses depended on task toxicity, employee skill and stochastic noise, represented in the model as a normal-distribution term written epsilon ~ N(0, 10).

Four algorithmic routes were compared: EUPG with ESR, NLPPO with SER, SFDQN with SF/RSR, and MONES with the combined perspective. MONES multiplied all three utilities and divided acute utility by 30 to obtain an average acute utility.

Each simulated episode lasted 30 days and represented a monthly work schedule. Return distributions were compared over 50 policy rollouts, giving a qualitative view of how the different approaches behaved under repeated simulation.

The policies did not line up

In the toy problem, the compared utility perspectives had different learned policies, or assignment plans, and different outcome distributions. The example reported differences for SER, ESR and SF, with a further difference under the simultaneous three-timescale comparison. No quantitative effect estimates or inferential uncertainty were reported.

ESR used switching to keep the modeled toxicity penalty around a threshold of 400 for all employees. That strategy sometimes exceeded the daily intake limit. The paper describes this behavior qualitatively.

SER preferred rotating employees 1 and 2 to keep them under the threshold while sacrificing employee 3. The expected return for employees 1 and 2 was described as close to ESR's, but with increased variance. This, too, is a qualitative result from the single synthetic example.

RSR favored a switch-averse policy that kept employees on tasks corresponding to their skill levels. Put plainly, the reported policy was less inclined to change assignments when the existing task-skill match was considered suitable.

The combined method, MONES, also shuffled assignments but mainly put employee 3 on the most difficult task. Its optimization favored the utilities of the two most experienced employees. The result was another unequal allocation in the toy setting.

A warning limited to the model

The authors interpret the apparent sacrifice of the third employee as a safety issue in the toy employer scenario, not as evidence about real workplace exposure. They argue that single-criterion policies may not suffice when timescales are combined, and that this gap motivates more comprehensive algorithms.

The evidence remains narrow: only one abstract synthetic example was evaluated, and the comparison was qualitative. No quantitative effect estimates or inferential uncertainty were reported. The result does not establish that MONES, or any other criterion in the comparison, is generally superior beyond this model.

The experiment also omitted sub-chronic toxicity, which the discussion suggests could be added as a fourth timescale. That leaves open how the additional timescale would be incorporated and how a combined method could avoid sacrifice behavior while meeting acute and cumulative safety requirements.

For now, the preprint's message is specific to this model. It reports that different utility perspectives appeared alongside different policies and outcome distributions, but it does not identify a universally safe or best approach. The authors' proposed direction is to develop methods that handle several timescales together.

Paper data and sources

Original title: It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning
Authors: Liam P. H. Mertens, Lucas N. Alegre, Florent Delgrange et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.